What might be some explanations of why CoT forgery is less effective with the likes of GPT-5/Fable/Opus? Or even GPT-5-nano, from the paper.
Awesome work. I wonder if it might be easier to worldbuild alongside an existing story - such as One Piece? May help with comprehension, ease of creating the mnemonic mappings, and potential unforeseen connections to the existing story.
Certainly this world would have to be complex and extensible - even if you don't use the exact characters, particular elements of the story may provide nice scaffolding
If you had to hash out a super-aligned AI scenario with respect to AGI-enabled dictatorship vs. abundance for humanity, what might that look like?
And why is it (obviously/reliably) Anthropic that will capture mass compute share in Plans D and C (i.e. ASI Race scenarios)? Or just a placeholder?