Summary I measure GPT-6 Astra’s performance on multi-hop tasks when prompted with a secondary CoT-control instruction: to reason using only dots, or to reason steganographically. Astra demonstrates covert reasoning capabilities with task performance beating that when using no reasoning or filler tokens for reasoning. This work agrees with findings from...
Subliminal Learning [Cloud et al] demonstrated the transfer of owl preference by performing SFT on lists of numbers generated by a teacher model that was system prompted (or fine tuned) to imbue its outputs with a love of owls. In this work I * Propose a methodology for studying subliminal...
This post should be understood as a follow up from Public evidence of the OpenAI-HuggingFace AI attack. The huggingface hacking incident left some additional traces exposed to the public. Investigating this data gave us some additional insight into the attacks performed by the AI agents. We also came across some...
Overview This is my writeup for my BlueDot Impact - Technical AI Safety Project. In this project I aimed to demonstrate that there is capacity for LLMs to take on steganography capabilities. In terms of AI safety, steganography is of particular interest as it may be used by a misaligned...