Fixing rewards for NLA to reduce confabulation
Hello, This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place. Note: 1. This post is 100% human-written. 2. Full paper in preparation for ICLR 2027 Anthropic's NLA(Natural Language Autoencoder) is mostly confabulated... until I fixed the reward. The NLA(The Natural Language...
Jul 248