Abhishu Oza's Shortform
Sep 281
Link to the repository of this work: https://github.com/abhishuoza/confessions-small-scale Summary OpenAI have proposed a new training technique for AI safety called Confessions (Joglekar et al., 2025), in which they train a separate output channel which is rewarded solely based on its honesty. In this post, I describe my work in reproducing...