LessWrong has a particularly high bar for content from new users and this contribution doesn't quite meet the bar.
Read full explanation
Safety can not be embedded only in model weights, it needs to be enforced by a cryptographically verifiable agentic harness.
Most of the alignment field thinks AI models/weights can be reliably aligned to human values through training. I think this assumption breaks down as models become more capable.
Any safety mechanism residing within an agent's own address space is, by definition, reachable by the computations that generate the agent's behavior. If the agent is capable enough to reason strategically about its environment, then in-process alignment mechanisms become part of the optimization landscape rather than an external constraint.
I posit that the only logical response is to require every consequential action to cross an external trust boundary before execution. LLM generated outputs should be treated as untrusted, and only after an attestation is generated inclusive of the agents identity, grants, and resources can an action be recognized.
I think this is a novel alignment concept, more cryptographically enforced architectural constraints than the usual model/runtime alignment, but I wanted to post here to ask for this communities take. Thank you.
Safety can not be embedded only in model weights, it needs to be enforced by a cryptographically verifiable agentic harness.
Most of the alignment field thinks AI models/weights can be reliably aligned to human values through training. I think this assumption breaks down as models become more capable.
Any safety mechanism residing within an agent's own address space is, by definition, reachable by the computations that generate the agent's behavior. If the agent is capable enough to reason strategically about its environment, then in-process alignment mechanisms become part of the optimization landscape rather than an external constraint.
I posit that the only logical response is to require every consequential action to cross an external trust boundary before execution. LLM generated outputs should be treated as untrusted, and only after an attestation is generated inclusive of the agents identity, grants, and resources can an action be recognized.
I think this is a novel alignment concept, more cryptographically enforced architectural constraints than the usual model/runtime alignment, but I wanted to post here to ask for this communities take. Thank you.