gregory.ruddell@snailsafe.ai
Message
Independent researcher and engineer working on inference-time interpretability and model dynamics. My current work focuses on how internal answer trajectories form, compete, stabilize, and sometimes reconverge during generation, using open-weight models and causal steering experiments.
My background includes defense electronics, including traveling-wave tubes, global remote server...
1
Hey there… I’ve been running open-weight steering experiments for a few years now. One of the tests I run is to push the internal answer—or rather the “candidate state”—toward a different answer. It’s more than just changing the token-level preference. What I find is that the answer frequently gets reversed or flipped back downstream. My thought is that the downstream computation is recovering the original answer.
With respect to your study, have you considered whether each unit of measured serial depth really corresponds to a distinct reasoning step?
It see... (read more)