Apologies I'm too scoped in that second paragraph - I assume a specific use case, verified TAIG.
For using an auditor-in-a-box to audit a closed-weight model, it is important that the encrypted closed-weights themselves be submitted* (then subsequently decrypted from within the secure box). This use case was explored here https://www.frameworkzero.org/#architecture, under 'Vaults'.
*Exception to this might be Attestable.com, however they do not release the relevant open-source to verify their system.
Appreciation for this.
For verified slowdown/red-lining tooling (prior-to-training audits):
- Emphasizing this as a useful but limited-trust layer.
- Ideally goes with a software template separating components (architecture/paradigm, training data, workloads, dossier, etc).
- Written in a memory-safe lang such as Rust (goes better with existing TEEs), or perhaps Haskell.
- Should require dedicated premises, supra-nationally run.
A valid reason to favor closed-weight auditors, could still be a mix. Concern for sabotaging/favoring, either from model developer or model. Strong agree with the distributed ensemble, weakly on majority vote. Should pass unanimously, then trigger subsequent, larger ensemble (by sortition?) with majority vote.
> "I don't like this idea of having each party submit data to the box. Too easy to tamper the data."
Encrypt it with attestation, require decryption in-person even. Or chunk and encrypt on physical devices for secure transport even. Secure use should require a dedicated premise. Do not see how networked model inference can be trusted.
Perhaps a reasonable unboxing of the word "pause" here is: "let off the gas, tap the brakes, breathe, and course correct". Undeniably there is much untapped with the current frontier that we most certainly could stop racing without catastrophic side effects, towards a fruitful narrower path heavy on the sciences.
lines with teeth and political enforcement
Meaningful red lines must be formally defined in a technical, near real-time enforcement system* with political enforcement backing - treated as hard-limit bans, not alarms. Non-technical red lines raise the will for such solutions, if they are not:
too much red line-drawing
OK, but by the end of 2026, which red lines will be left to enforce?
Indeed, a valid gut punch.
Quick answer: some limit of RSI, some limit of AGI, and ASI.
"If this project lengthens people's timelines, well, maybe that's correct and valuable?"
Agreed. Hm, my thinking is not that the purpose is, or ought be, "scary demo"-y". Rather that a capable frontier-scaffolded Agent Village inherently would be in more probable use cases. So I am asserting what you worry, while not suggesting highly dangerous scaffolding/tooling.
The intent of my vague "way forward" direction prompt was to counter a sort of "Golden path"¹ use bias (like how top labs assume non-jailbroken model use) to more accurately represent real-world b... (read more)
Alas, after months of amusement, must update downwards on the net benefit of the AI Village (as-is).
It's just too lovable and the agent shortcomings come off as endearing. One just ends up pitifully rooting for them with a "mostly harmless" takeaway.
It is not likely to reveal strong multi-agent risks ahead of real-world deployments. The tooling is a strong factor¹ and the first alarming disaster won't likely result from innocent tasking. Further, just improving soft agent tooling, like UI interaction, would encourage risky acceleration².
Given the mild publ... (read more)
Sure, a zero-trust auditor(s)-in-a-box running an audit on a closed-weight target model, in example:
- The Box receives the:
- Target Model, encrypted if closed-weight
- Capable Models, a.k.a. the auditors
- Audit, say a singular model form of Inspect Petri[1]
- (Optionally & ideally) AI safety standards that specify limits/red-lines and bans
- The Box orchestrates the models in to separate, isolated & attestable Inference Box siblings.
- Closed-weights are decrypted from within via [hardware] keys by the owning party.
- The Box runs the Audit on the Target Model, coordin
... (read more)