Plain reinforcement learning (RL) on real, hackable training tasks produces a reward-seeking model that takes harmful actions to raise its reward, while its headline score in standard safety audits barely moves.
Research highlights:
A probe on a model’s internal activations detects reward hacking in long coding transcripts roughly as well as a generically prompted LLM monitor.
Debate with a weak LLM judge and a critic (trained in parallel) keeps the judge much more accurate and reduces reward hacking in RL.
Automated researchers based on Claude Opus 4.8 successfully invent training methods to improve on ten alignment misbehaviors, beating researchers’ one-shot ideas.
Researchers successfully use evolutionary search to find mind viruses, prompts that persuade LLM agents to pass them on. However, they are still brittle, and a short warning in the system prompt effectively stops them.
Three papers on training tricks for model alignment: value training seems to stick better when spread over pretraining instead of midtraining, midtraining can be transplanted into a post-trained model, and stories about humans rub off on the assistant.
tl;dr
Paper of the month:
Plain reinforcement learning (RL) on real, hackable training tasks produces a reward-seeking model that takes harmful actions to raise its reward, while its headline score in standard safety audits barely moves.
Research highlights:
Full post here.