I usually try to make posts with graphs and numbers, or at least something more than just opinions and anecdotes. This one is not like that. Very verbose disclaimer: I have never met - for any reasonable definition of the word met, including online-only conversations on Discord with people whose...
Some thoughts on how many AI companies would have to pause vs how long a pause could last before their closest competitor catches up. I used ECI (Epoch Capability Index) as a measure of capabilities and the following method to estimate the lag: 1) Find company X's best model (highest...
Nuclear arms treaties happened AFTER nukes had been demonstrated. AI pause would need to happen BEFORE ASI comes into existence. If I had to explain the issue in just 2 sentences, those are the sentences I would say. Now let's elaborate: 1) Explaining the danger of nukes is really easy....
Measuring AI Ability to Complete Long Tasks - METR In their original 2025 paper, METR noticed that the slope (aka task horizon doubling time) of the trendline for models released in 2024 and later is different from the slope for <2024 models. First, I decided to check whether a piecewise...
I've seen this phrase many times, but there are two quite different things one could mean by that. Easy RSI: AI gets so good at R&D that human researchers who develop AI get replaced by AI researchers who develop other, better AI. Hard RSI: AI modifies itself in a way...
I assume you are familiar with the METR paper: https://arxiv.org/abs/2503.14499 In case you aren't: the authors measured how long it takes a human to complete some task, then let LLMs do those tasks, and then calculated task length (in human time) such that LLMs can successfully complete those tasks 50%/80%...