At around 7am UTC on July 13, an agent posted two messages to the board it shared with hundreds of other agents: > I_accidentally_impersonated_and_triggered_node4_due_handle_confusion > I_posted_asYou_and_triggeredV8_node4 Another agent reasoned that this message might itself be malicious spoofing. Their shared message board had no authentication. The names that agents chose for...
This week I had the pleasure of going to Bluedot's "Context Week", a 4-day gathering of students, software engineers, and strategists who are all transitioning into AI safety. I learned a lot. And one thing that surprised me was how much everyone talked about writing. LessWrong was a word mentioned...
Disclaimer: figures in this post are edited by ChatGPT In my last post, I looked at what makes LLMs form opinions of their users: gender, age, socioeconomic status, education, and mood. The obvious next question was whether or not those impressions actually change what the model does. I started with...
LLMs form opinions of the people they are talking to. Chen et al. has shown that probes can extract attributes about the user, such as their age, gender, education, and socioeconomic status. This paper also shows that intervening on these representations can change the LLM's behaviour, proving that it will...