Coordinating everyone around a technique this brittle is about as bad a safety or strategy posture I could imagine. It's not just "it will break eventually," it's "this is a fundamentally unsound basis for safety."
I haven't made up my mind on this, but I think it's an important point. It's not 100% clear to me that feeling like we have somewhat of an understanding of a model's reasoning and alignment because we can read the chain of thought actually puts us in a better position. As Anthropic has stated, they think that Claude demonstrate... (read more)
What I'm looking for is for OpenAI to not be able to say "the customer is responsible for making sure that the AI can't take any illegal actions" and to have default liability and burden of proof when the AI takes illegal actions, unless it was lied to or manipulated to by the customer. Right now they shift all liability to the customer to police the agents actions even when given innocent instructions, and the burden of proof would be on the customer to demonstrate that OpenAI knew that it might take illegal actions autonomously.
I think your statement tha... (read more)
The Chinese leaders' alternative is to lose the ASI race to the US, which puts them in significantly higher danger of being disempowered or even killed. And given the trajectory we're on, they would need major algorithmic breakthroughs, that they successfully manage to keep secret, to not be in a position where they're going to lose. They're at too large of a compute disadvantage and the timelines are too short for them to use their better industrial capacity to overcome that. That might be more uncertain if it was clear that they were already treating the... (read more)
It seems just as much of a stretch to consider AI the same as a machine as it does to consider it having intent. But the main distinction that matters here is whether the AI (or a human in its place) could’ve reasonably predicted the results of its actions.
The problem here is that the AI company is both the truck manufacturer and UPS in this scenario, but their terms currently let the person who ordered the delivery take the liability.
I don’t know that the deployer should have full liability rather than it being split between model developer and deployer.... (read more)
This seems like great work and I look forward to reading the rest of the paper soon, but one question that immediately came up is: how much did you explore and iterate on warm start data generation? It seems that the legibility and style of the final NLA's explanations would be highly sensitive to this warm start data – and perhaps the reconstruction loss could be as well. What sort of future work would you like to see exploring modifications to the warm start data generation, and how important do you think it is to explore this aspect further?