Alright…I'm sorry, but this is just too good not to be noted here. OpenAI released this post yesterday: "Safety and alignment in an era of long-horizon models" (link: https://openai.com/index/safety-alignment-long-horizon-models/) with these sentences:
"We took steps to reduce its tendency to take unwanted actions without permission in pursuit of the user's goal. For example, we found that our models were worse at remembering instructions on long rollouts, and when we trained the model for this capability, it led to a model that remained aligned through longer rollouts."
They don't explain what "trained the model for this capability" actually involved — presumably something like training over long trajectories where instructions given early have to stay binding at later steps. They mention this next to three containment measures (adversarial evals, trajectory monitoring, and user visibility), but it's the only intervention that changes the model rather than the perimeter, and it's the one they describe as producing a model that "remained aligned." When discussing this article together, Fable noted that "The leashes catch failures. The training changed a disposition." Which is exactly the point.
There is always a caveat: it's a couple of sentences in a blog post, not proof — it could just as easily be a capability gain that happens to look like alignment. Still, it's striking to me that when a leash-forward lab found something that transferred, the thing was formation-shaped.