Keeping model weights contained is another form of AI control. Maybe this is also bad? If Anthropic released the weights of Mythos instead of privately trying have it fix all the software bugs, we would have had a massive, internet-wide cyberattack by now, which would have been earlier, more impactful, and harder to dismiss as marketing hype than OAI/HF. Should AI safety update from this away from its stance against open weights models?
Keeping model weights contained is another form of AI control. Maybe this is also bad? If Anthropic released the weights of Mythos instead of privately trying have it fix all the software bugs, we would have had a massive, internet-wide cyberattack by now, which would have been earlier, more impactful, and harder to dismiss as marketing hype than OAI/HF. Should AI safety update from this away from its stance against open weights models?