Off-policy honesty training generalizes better than on-policy honesty training
This work was done by Purvi Chaurasia with Daniel Tan and Chloe Li as part of the SPAR Program for Spring 2026. All code related to the blog can be found in this repo. We investigate self-report fine-tuning (SRFT), a technique proposed by Li et al. to improve models' honesty....
Aug 109