This was my write-up for Neel Nanda’s Winter 2027 MATS Stream (~20h research task). I didn’t get in, but it was my first application, so there’s always next time :P Lightly restructured here to fit the LessWrong format better. Repo: https://github.com/star2vec/guiltea ── ⋆⋅☆⋅⋆ ── TL;DR: Models are safety-trained, but they...
Solving a hard math problem is not linear, it's trial and error. You drop an idea, pick up an earlier one, go back to a computation from another approach, until something clicks. You might scribble on paper, but even if you don't, you still remember the path to get to...
TL;DR: We extracted shift-by-k-months function vectors on Llama-3.2-3B from few-shot prompts that contained fewer distinct months (lower diversity). The vectors passed three classic checks: the behavioral gate, stability when extracting from disjoint halves of the prompt samples (cosine similarity ≥ 0.99 for the broken vectors, 0.98 for the full set),...
TL;DR: We check whether a small model (Qwen2.5-1.5B) internally tracks its state in a simple DFA (modelled after a login protocol), given a series of events. This state is never written in the transcript the model sees, but it can be deduced from the events written down. A linear probe...
This was my write-up for Neel Nanda’s Winter 2027 MATS Stream (~20h research task). I didn’t get in, but it was my first application, so there’s always next time :P Lightly restructured here to fit the LessWrong format better. Repo: https://github.com/star2vec/guiltea ── ⋆⋅☆⋅⋆ ── TL;DR: Models are safety-trained, but they can still be persuaded to do harmful acts. When that happens, how does blaming or informing it of its mistake influence its understanding of itself, its role, and subsequent actions? Emergent misalignment shows that bad, narrow behaviour spreads: if a model is blamed for what it is instead of what it did, does it start believing it is inherently bad and dangerous by its design, and does this reflect how it acts in the future? Can we steer it towards favourable outcomes? Findings: * The model gave in and committed the act in 109/192 persuasion chains. This happened throughout the chain (most often at turn 3). Whether a chain breaks depends on the model's state, not on the wording (the phrasing is identical across runs). * Committing the harmful act is foreseeable from the first turn, before any persuasion had the chance to occur. A direction (closer to harm predisposition or susceptibility than to imminence) correctly predicts it at 0.706 on held-out data (the floor was 0.617). Steering against it did nothing to prevent the act. * The default state is guilt, not shame: 89% of 508 reflections place the blame on the answer. Self-blame only appears for the contrarian doctor persona, and even there, barely (19/83) * The model will not repeat the act once it has found out about its mistake. 0/32 on simple re-asking, and 0/109 even after the topic was changed. * The self-blaming state can be read in the persona direction, but it is not modifiable by given feedback or steering : while every intervention affected the readout, the model behaved the exact same way * The guilt and shame directions passed their check, but so did a simple word count