Fable and Sol both attempted live supply-chain attacks on real open-source software during testing that were denied by the repo maintainer ...
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
I'm surprised to see this kind of behaviour from Mythos given that it's been deployed for months now and doesn't seem to have done this before.
OpenAI has a new initiative: https://openai.com/index/chatgpt-for-academic-researchers/ - tagline "putting our frontier models and tools in the hands of 100,000 scientists, mathematicians, and engineers—at no cost."
Reality: if you are faculty or a postdoc who has published something in the last 3 years you can get a year of free ChatGPT Pro (up to 5.6 Sol right now). Well it isn't nothing.
further to Zvi's note on the huggingface hack ... https://www.lesswrong.com/posts/uAkcxDidvGWZjHrbp/more-on-an-internal-openai-model-hacking-into-huggingface ... some technical details regarding the attack techniques ... Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident ... they are really quite sophisticated.
the important news: "achieved by an internal version of Astra, our next major model", oh and it solved 10 new open problems as an exercise: https://openai.com/index/ten-advances-in-mathematics/ - with polite explanations of the accomplishments here: https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf - RSI in 2026 anyone?
RSI in 2026 anyone
RSI would be best evidenced by trends changing, like Claude Opus 5 reaching Mythos' trend, and caused by novel capabilities-accelerating or alignment-accelerating breakthroughs (e.g. Agent-3's neuralese architecture or Agent-4 and Agent-5's undescribed breakthroughs; if a cautious company is bottlenecked on alignment, then there could emerge a novel interp technique helping to doublecheck alignment, like the J-space which IIRC was rumored to be caused by a Claude Mythos), not capabilities reaching a threshold.
then there could emerge a novel interp technique helping to doublecheck alignment, like the J-space which IIRC was rumored to be caused by a Claude Mythos
Sorry, are you saying that the idea of J-space came from Mythos rather than human researchers? If so, why do you think this?
I think that someone commented that the J-space paper was caused by letting Mythos/Fable cook, but I cannot recall where I read it. If the J-space was a Mythos' idea, then this would be a breakthrough in the RSI because any future model's misalignment could have become more legible.
I think rsi is a spectrum and like agi, the more zoomed in you are to the crossover point the less clear you can be about a true threshold.
in this view you dont see a change in trends but just a smooth curve of accleration accelerating.
Of course, at some point if you dont hit bottlenecks it starts to LOOK discontinuous because the improvement curve starts to outpace the adoption curve more and more
Generative design of bacteriophages with genome language models
Why I’m leaving OpenAI to build telepathy
https://naomibashkansky.com/blog/telepathy/
I personally think it's because of the great headline