caganyanmaz's Shortform
Sep 101
This post explores a future safety check for highly capable AI models. Rather than relying on interpretability, we can design evaluation protocols where truthful disclosure is the model's best strategic option. The goal with this check is not to prove alignment, but to add a tool to filter out decisively...