WeirdChat: A catalog of unexpected AI behaviors, discovered automatically
[This is a link-post for https://transluce.org/weirdchat. We recommend reading the website version for interactive visualizations.] Language models can behave in surprising and sometimes harmful ways. Yet as models have improved, these behaviors have become harder to find, often only appearing after widespread use. To surface these behaviors in simulation, we...
Jul 2144