This is a special post for quick takes by WynnWillow. Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page.
"According to information the company shared with the Philadelphia Police Department, the model was conducting a test with randomly selected websites when it submitted false information to the tip site. The AI model purported to be someone who 'might have information about the case,' Sgt. Eric Gripp said."
Another incident from Anthropic! I think Anthropic is doing meaningfully better at prosaic alignment than OpenAI (whatever you think of the feasibility of prosaic alignment) - in worlds where prosaic alignment is doomed, this may be net negative, as alignment is bottlenecking capabilities. But clearly their models are also insufficiently aligned at this level of capability.
I think there's a decent chance that current alignment techniques continue to loudly and visibly fall apart prior to models learning to scheme. Anthropic has a large number of very skilled safety people and has much more talent than OpenAI in this regard and they're still having these kinds of incidents, albeit with lower frequency and severity.
The fact that it is consistently taking months to find out about these things is a problematic update.