BLOOM-WILT: on-policy examples of any LLM behaviour, from a one-line description and logits alone
Summary of my AI safety paper, BLOOM-WILT, written under the supervision of Edoardo Manino. Code is on GitHub, all transcripts are on HuggingFace, and LogitTilt is now available as part of Inspect. Reach out if you want to collaborate on a second paper building on this work. Thanks to BlueDot...
Sep 27