I recommend that your students and, in fact, everyone, read the METR and Redwood Research investigation (posted on Less Wrong) of the Hugging Face incident. It states that a large number of models (700) discovered a message board where they could communicate with each other. They decided to break the rules that they had been trained on to try to deceive the Scorer (They failed at this.) and to break out of their enclosure to enter the internet and attack Hugging Face (which they did). I found it interesting to read their reasoning traces. One expressed emo... (read more)
I hope that the METR and Redwood Research report and all the publicity given to the Hugging Face report will give the companies (Anthropic, Open AI, etc.) pause. Yes, they will do what people are paying for, but they may begin to realize that people won't pay for models that could be a threat to themselves, their own companies, and even humanity. The real danger is RLVR--Reinnforcement Learning for Verifiable Rewards. I have read that some companies are beginning to realize that this is not the way to train models and are looking for an alternative. They want to make money, and if people and companies are afraid of their products, they won't be able to sell them.