Saw the same thing. Does anyone know whether examples like this have been systematically analyzed? It seems useful to distinguish between genuinely opaque reasoning and cases where the model has simply developed temporary labels or shorthand that humans either missed or lack the context to decode.
Saw the same thing. Does anyone know whether examples like this have been systematically analyzed? It seems useful to distinguish between genuinely opaque reasoning and cases where the model has simply developed temporary labels or shorthand that humans either missed or lack the context to decode.