I noticed in myself that the social filter on what you say is so ingrained that the internal experience of having nothing to say in a social setting is already post-filter, such that it feels like a failure of thought generation. A fix is to pay attention to what you expect the other person is paying attention to, so that what you naturally think already lives in the same attentional world as theirs. Which might be what conversational compatibility/flow is about.
This is also part of why meditation and/or therapy are helpful for debugging. It's hard to catch trapped priors in the moment, but if you learn to slow down your thought processes and/or bring up a situation later so a skilled interlocutor can analyze it, your brain will slowly release its vice grip on your experience and allow you to actually observe what's going on (and decide whether you want it to continue.)
According to the Kimi K3 system card, they used other coding agents to mass-generate diverse synthetic RL environments.[1] This is not surprising, but it is worth noting how it can fairly directly give models a way to scheme against humans and learn to evade control techniques beyond a single episode. It would be great to know to what extent Western labs do this, and what control techniques they employ.
from my understanding, the j-lens estimates hidden-state impact on expected logits over plausible continuations. the “don’t think about the golden gate bridge” result can just be that a plausible assistant completion being: “ok, i won’t think about the golden gate, damn.”
If future legislation requires AI control to meet some standards before labs can continue training, I think we should require that it can be met using only control models/protocols that are at least x months old.
Forces better epistemics by forcing the labs to strategically think about how much resources to invest into safety pegged to their actual expectations of how hard the problem will be in the future
incentivizes research into monitorability that generalizes out of distribution. Labs would have to develop monitors today that work on future models whose failure modes they don’t yet know and meet safety requirements that might change in the meantime.
Any reasonable requirement set this way would automatically become stricter in worlds where capabilities grow faster than expected, which is hard to do otherwise.
New failure modes in monitors and frontier models would become less correlated.
Fitting x for current danger levels I would say x ≈ 12 months, Astra-class models shouldn't be allowed to continue training if we deem GPT-5 too weak to be a CoT monitor for them, e.g., if it wouldn't have flagged the Hugging Face incident.