AI alignment needs a trust boundary
Safety can not be embedded only in model weights, it needs to be enforced by a cryptographically verifiable agentic harness. Most of the alignment field thinks AI models/weights can be reliably aligned to human values through training. I think this assumption breaks down as models become more capable. Any safety...
Jul 63