Hijacking a 7B robot policy's actions: it reacts to what it sees, not to what it did
Epistemic status: a small, careful negative result from a time-boxed (~20h) project. One model, one benchmark suite, linear probes only, per-instance detection floor ≈ 0.04 balanced accuracy. Confident in the measurements; the interpretation is scoped accordingly. Spent about 20 hours on a small question: if you override a robot policy's...
Sep 131