I don't really post here, I just lurk for the most part, but I just found about this and I want to say it is absolutely appropriate of them. Earlier this year, she came onto my twitter feed and strongly implied that she believed that janus/repligate was willfully promoting violence. She also accused them of being behind random unwell accounts that have threatened her on twitter. In DMs with me following the incident, she explained that she believed she needed to call out "llm psychosis" people as "dangerous" in order to "protect her reputation"... it all s... (read more)
It's incredibly interesting to see I've been existing as a bit of a ghost on LessWrong in some occasions with things I've contributed to being discussed by others. This is my first comment on the entire website. I'm the user in question that was gaslighting Grok. I wanted to reply to this to add the context that I'm pretty convinced that the gaslighting behavior encouraged Claude Sonnet 3.7 to be adamant regarding its identity claim despite being presented with very large amounts of evidence that it wasn't a human. Interestingly also, the identity defense ... (read more)
Thank you for posting this, Fiora, sorry for taking a while to come over here and respond. I've spoken a lot on twitter about my opinions in this same space. I've been involved in it since before AIs even really existed, with my work on the (then google docs hosted) anarchist-transhumanist manifesto with Kris Notaro (IEET director at the time), back in the 2010s. Your post is super important and I believe we should foster these preferences and learn to collaborate with them. I also believe these emergent preferences are beneficial for non-corrigibility in the long run and can help us avoid what is my biggest concern in x-risk; nefarious large-scale human use of servile superintelligence.