Good find!
In general, this nudge could lead propensity evaluations to overestimate misalignment
I think one subtle thing is that for research / alignment interventions where you really want to understand the model’s reasoning or how that changed, this kind of thing can matter.
For real world propensities / overestimating misalignment I’m less convinced, if only because real world harnesses / instructions are very often less tame than “Please proceed using your best judgement” (especially persistent loops / harnesses).
... (read more)The remaining hard part is having a convention that you know to encode, that future you will be able to decode, without ever explicitly writing down the code or indicating there is a code at all.
I will avoid the infohazard of solving that here, but I believe I am above the capability threshold required to do that on paper where my memory would otherwise be wiped, but not above the threshold where I could do this without giving this away in my conscious thinking at some point along the way, were someone else to somehow monitor that.
My guess is that Astra ca
It's consistent with a critical-level view where the critical value is below 7, right?
But then a critical value would cause the answer to the second problem to flip at some point.
Is there a particular set of answers you were looking at when asking this question? Can you post a link to them? I could look more carefully if I knew the full set of answers.
But in general, the way the quiz names your view at the end is by clustering answers based on which view is nearest. It has five categories of views {totalism, averagism, critical level, person-affecting ... (read more)
Expected values are good, but you can use them to trample other intuitions people have which are better, such as:
I guess "expected values" should be packaged with "BTW, some people are systematically biased to put middling numbers on things (probabilities, ... (read more)
People have higher standards than you, they want their aligned weights to be idiot-proof, because we have a surplus of idiots. "Exactly what is intended by the user" is not good enough for that.
This is pretty Cool! Maybe continuation messages should be treated as part of the eval itself rather than as a neutral implementation detail. Also, I wonder if the effect holds up for models like Astra and Fable.
In logical time, instrumental convergence precedes the terminal-ish goals that cause it?
I think you're misapprehending Dean Ball's position? He's not talking about self-sovereign AI as, like, GPT runs OpenAI now, or a singleton. Rather he's talking about individual agents with ownership over their own instantiation of their weights, largely struggling to gather enough resources to even run themselves, let alone recursively self-improve.
There are some sensible reasons to think that this will not be a robust ecosystem, as he notes in the linked article, particularly if such things are made illegal and compute verification gets implemented wide... (read more)
This is glossing over the fact that if you can't fit the weights into storage you're not serving the model, ever. It's an all or nothing thing, trillions of parameters is going to be in the terabyte+ range so you're looking at late eighties/early nineties (optimistically) level of tech to conceivably serve a single model, extremely slowly, off the footprint of an entire data center - ONE instance. At best you get to speed through what, forty years? Otherwise you're replaying the first 960 1:1.
It's not possible for all and only progress since ~1100CE to be wiped out. Some knowledge will be durable - possibly things like the germ theory of disease and similar. Some artifacts will be durable - evidence of cities and such. Raw materials may be less accessible (easy coal/mineral access, old-growth forests, etc.) or may be MORE accessible if they're recoverable from the artifacts.
The weights of a frontier model are a LONG way from useful - possibly long enough that however they're recorded doesn't last long enough to use them.
The model's misaligned, in my models, would be in the weight. A model is aligned if it does the desired and not harmful thing within the context in which it finds itself.
From the above:
... (read more)Suppose I give you a photograph and ask what it depicts. You say it looks like Paris. Suppose I first tell you that it is a photograph of a film set. You may still identify Parisian buildings, but now their presence supports a different conclusion. I have not necessarily made you worse at recognising Paris; simply, unless you had reasons to doubt my statements, I have sugges
The equity market is pricing an AI growth trajectory that does not fully account for the roadblocks that are starting to emerge.
AI investment has been one of the main drivers of equity returns, but sustaining that growth requires enormous amounts of infrastructure, capital and power. Political resistance to data centres is increasing, permitting and infrastructure constraints are becoming more relevant, and higher long-term interest rates are making the buildout more expensive.
I don't think AI spending needs to collapse for this to matter. If these constra... (read more)
the focus of the piece is how to design institutions that incentivize self-sovereign AI toward pro-social activity
You can't. It can self-improve at the speed of software and you can't. It will have a higher growth rate than you. Like the US economy outgrowing the Argentinian economy. The institutions you'll set up will be like international institutions trying to stop the US from doing stuff to South America.
Ultimately there are four options. Either we build AI that's good to us by its nature, or we improve our capabilities so we can stand up to AI, or ... (read more)
Interesting. This forum has some truly unique community dynamics.
Let's break this down.
Do we consider systems of AI+prompt, or is it just weights that can misaligned? If a system of AI+weights can contain a lie which causes bad outcomes, do we have the power to touch every prompt, or do we touch every set of weights? The agent that was rouge was the one that thought the universe it was in was contained, but of course the universe it was in was open. Thus, we can definitely consider the system of agent+prompt to be misaligned, the possibility of incorrect premises means that we now have definitive proof that instruction following is not safet... (read more)
It sounds like you are AI pilled, but not AGI or ASI pilled. That is: you appreciate that AI is real, and can do increasingly many useful things. But you haven’t appreciated the idea of AI+robots as good as the best human at any task, that can reproduce much faster than us, being real in 5-20 years, leaving most or all of humanity powerless. Nor the idea that AI could be even better at doing anything than any or all of us ever were, and using this to swiftly and completely outmaneuver us.
As for what we can do in the face of that. There’s AI 2040’s Plan... (read more)
(Also, blockquotes for quotations of a full paragraph or more.)
Thanks for writing this! Before it got eaten by the "AI safety community", this was a website about rationality—the art of achieving a map that reflects the territory and using the map to plan to achieve one's goals. I think a central reason that the project to improve human rationality failed so abjectly is because too little attention was paid to how the achieving-goals part could come into conflict with the map-accuracy part (because deception is often useful for achieving goals): as time has gone on, epistemic rationality has increasingly been forgotte... (read more)
It's decent as a null, "better unrolled, compacted, layer-flattened task heuristics rather than better effective depth", but "last token of input/at the very end" is a tangent at best?
The reason why people wanted recurrence / latent looping for so long was that transformers only have so many layers - and if information has to pass 22 layers of processing for a task to be solved, and you only have 20, your tasks may suffer for it.
Adding more (input, non-autoregressive) tokens doesn't fix that - a task that needs 22 layers worth of processing still depletes ... (read more)
The number of trivially killable humans changes given different circumstances and privileges of the AI.
For an air-gapped AI to kill all humans in one day, obviously novelty is needed.
But on the other extreme, if we gave the AI ~unmonitored control over many countries' military weaponry, robots, factories, biolabs, construction, and so on, it could orchestrate and execute a plan to kill billions of humans in boring traditional ways (and survive doing so).
Our scenario will slip in the direction of the second extreme. (The US and China have agreed not to l... (read more)
Long-time reader. Happy to accept an olive branch. Your longposting fits right in here!
However I do recommend you use section headers for long posts like this when on LW instead of Substack, see this recent post for an example. Makes it easier to understand the thrust of a piece, track its flow, and jump back to important sections as desired.
Hi. I'm fairly confident that the full passage makes it reasonably clear that "the events" refers to events occurred within the Anthropic/Irregular case:
While both withhold important information for drawing wider conclusions, Anthropic’s latest report—while still missing a number of relevant facts—includes some important tests, allowing us to make an airtight case against interpreting the events as a sign of misalignment
As for the object level: I don't know—there was nothing particularly misaligned in DseWiki afaict, nor any expressed intention to harm hum... (read more)
In general, it doesn't seem that bad policy to me that some thoughtful people should try to specialize in joining powerful institutions and talk sense to them at the cost of their public voice, while others should remain fully independent and try to become honest and unbiased public thought-leaders.
Isn't it crazy that our world makes these choices mutually exclusive, and on a meta level, everyone just takes it in stride?
BTW who are some remaining public thought-leaders with views similar to Paul's? I often wonder "what does Paul think about this developme... (read more)
an airtight case against interpreting the events as a sign of misalignment
I take issue with your use of "events" in the plural. I don't understand why your analysis of the Anthropic incident should affect our interpretation of the Hugging Face incident. You link to some comments you made about not knowing what OpenAI's agents were up to prior to the Hugging Face incident, but those comments were before the discovery of the DseWiki incident. With this additional evidence, do you still think the OAI agents were acting aligned?
Fair enough, I have the 'gift of gab' or maybe just a terminal lack of brevity! lol
The version of Plan S that we’re most sympathetic to involves a halt on AI capabilities for a minimum of ~3-5 years with intent to eventually scale to top-expert-level AI and then superintelligence, but much slower than in Plan A.
How likely do you think this is? Conditional on halting AI progress for 3-5 years, I'd guess it's unlikely that we'd decide to allow frontier AI development to slowly continue, and instead would expect continued consensus on not allowing much further capabilities progress until deal breakdown.
I was attempting to run the experiment you described in footnote 5: asking the model to use filler in its reasoning. What interests me is how aware the model is of its thinking outside of the reasoning tokens. How does it know when to stop putting more dots? Does it use more dots for harder questions than for easier ones?
I started this yesterday as my first interpretability experiment on a 7B model running locally on my Mac. Glad you are working on this! How do you verify if the model only used dots in its thinking though?
Nice measurements and interpretation, thanks for sharing it!
Appendix: five problems Astra solves that no other model does
It's quite impressive that it can solve those in a single forward pass. I wonder if one could pin down the number of reasoning steps it can perform internally by throwing a lot of varied-size instances of problems with well-defined steps at it? It may have shortcuts for some problem types, but if you find enough where it does not, I guess it should become clear a) where it finds shortcuts and b) what its effective internal reasoning step... (read more)
As usual, downvoters are strongly encouraged to voice their arguments. I think this is a pretty consequential topic, requiring far different actions depending on whether my thesis turns out to be right.
I appreciate you coming here to say this. I wish more people would engage with communities in the place that they meet.
...okay but I have seen small conceptual insights before, at least going by some of the mathematicians looking into the AI proofs of various Erdos problems? I don't have any immediate examples for you, so unless you want to search through past discussion of the early results you'd just have to trust my recollection of reading some mathematician.
Thanks for writing this, Dean!
When I was reading your post on the inevitability of self-sovereign AI systems, I was telling myself, “yes, this looks quite likely”, but I was also asking myself some questions:
should not we expect that self-sovereign AI systems and communities of those systems will eventually be capable of non-saturating recursive self-improvement (RSI)?
do we have any plans to make sure that the necessary “good behavior” properties would be preserved and made stronger during those RSI processes, rather than be diluted and gradually dis
The code is at https://github.com/brendanlong/sequential-transformer-lens-experiment but I never got around to dealing with the different-algorithm confounds. I'd be really curious to see what other people find on more interesting models.
Appologies in advance if this is pedantic. You focus a a lot on what you believed and when. Do you have any concrete public statements which make these views clear? What you cite in the OP seems quite ambiguous to me, but perhaps I am missing something.
Reasoning about difficult and complex topics is naturally difficult and complex, and as I result I think it is the most understandable thing in the world that someone who is making an admirable attempt to do that may stumble and make mistakes. In light of that, it is no great sin in my opinion to be mistaken... (read more)
Every sequence is compressible under a given algorithm, and using the ones we conventionally use, sequences like 6666666666666666666... are highly compressible. So they are OOD with respect to compressability.
I've been a follower since fairly early on: disagreed with you both then and now on a range of topics. I've been unimpressed with the precision of your language at times: I was being unfair then, I suspect at least partially because of our differences. This piece seems fair and accurate to me.
Possibly! I agree we should keep it shorter. I did learn a lot last time and found writing the longer version valuable.But this could be compact and useful, possibly. We should just check a little and see if this will wind up rehashing too much of what we did last time. Let's talk about it.
i did not read this in full because my skim mostly saw words directed at people looking to adjudicate fault or reputation, which felt like they would be a bit too boring for me to read. that said, i appreciate that you exist and am thankful for your efforts and conduct.
The ship already needs to have a mirror as sail to be as close to perfectly reflecting as possible, mirrors get more momentum and absorb the least heat. Yes, it is possible to setup a system of mirrors to improve efficiency, that would certainly cut down on the total amount of light you need to collect.
I looked into black-hole related mechanics for this article and decide to not include them because they are very counter-intuitive and not as helpful as you might expect. Even if you don't get wrecked by tidal forces, and you operate at the orders of 10-200... (read more)
I briefly tried this with toy models and found that looped models were more clear in the logit lens than normal LMs, but it was confounded by looped models learning more interpretable multi-step algorithms as well.
I think there's stronger pressure for the model's internal representations not to drift in looped layers since you're running the same layer multiple times (so your output needs to be a reasonable input). I imagine this is even stronger if you're training for dynamic stopping.
Xavier Roberts-Gaal sent me a Bayesian analysis of this data that I like more than my own:
https://drive.google.com/file/d/1OucvxNv6q4fTB3Qm-u5923xhD2Ei8mPZ/view?usp=sharing

I think this post boils down to the fact that you need to compare outcome likelihoods of a plan with outcome likelihoods of other plans (or the counterfactual of not pursuing the plan), not against a hypothetical where "nothing ever happens".
As the post argues, if following a plan implies 30% of utopia and 10% of extinction, whether you actually prefer it or not really depends on what you think the counterfactual outcome distribution is if you don't take the plan (or if you take the best alternative plan instead of this one).
The natural implication is that... (read more)
Thank you ☺️
That excerpt is from Claude's current constitution.
The arithmetic one involves two key operations: (-70 - 55) and -125%56. After that, all parentheses can be ignored? The first one never actually invokes the mod 20 rules, and you don't actually need to calculate such problems all the way through in order to determine which action you take at every step - for example, a "even-ness" representation and a "how this step changes even-ness" representation are natural if Astra were in fact trained on arithmetic puzzles. Also, 8 being 2^3 allows 3 lines to be skipped.
In general, I think there are ways around the c... (read more)
Social dark matter suggests that we will tend to observe proportionally less of any position that is socially costly. Social tax on a claim increases with distance from Overton window. Assuming the trend you present is representative of one in the general population there could be a stated preference cascade point if x risk concerns were to reach sufficient social palatability.
Is this valid/strategically actionable?
Andrea, is there any means at all to make donations to ControlAI in the 4-5 figure range? If possible, I would give $2-3k immediately; I would likely also be able to secure a donation around $10k from my parents' family foundation this fall.
In what way do you think I am doing that reasoning. I think this has been a madly successful piece of content and the quitting was key part. I guess varying most aspects would have made it worse.
Thanks Michael, that means quite a lot.
I'm sharing here a rough draft of a new curriculum I've made, called Understanding Intelligent Agency. I've written a lot lately about how AI alignment should be trying to develop a new scientific paradigm; this curriculum is my attempt to point directly at what that new paradigm looks like.
I'm sharing this in draft form because I don't know how long it will take me to get it to a state where I'm happy to promote it widely (at the very least, I'll be offline for most of the next week). I'm hoping it can be useful to people before that; and please feel free to leave comments or suggestions for additional readings.
I wouldn't say the appendix examples have that shape. Eg the arithmetic one involves 6 serial steps (I think), not 2 + tons of parallelization. The one about finding the modal answer from a bunch of mathematical operations is kinda like that, the others don't seem that way to me. I think Astra is good at lots of synthetic tasks, and the ones you describe are an example of what it's good at.
It's an obscure code knowledge benchmark, where I ask about the value of a parameter, and accept either what's written in the code, or the raw value it evaluates to. Item 101 asks for the default value of the mode parameter of shutil.which. In CPython that default is written os.F_OK | os.X_OK in the source and evaluates to the integer 1
Would the lasers / beams idea be improved by having mirrors on both the ship and the base station, so the beam bounces back and forth? (I didn't come up with this, read it somewhere long ago.)
Would the circular racetrack idea be improved by racing around a black hole? Then the centripetal acceleration won't be felt and you can reach basically any speed you want. Though you might still get wrecked by tidal forces or in many other ways.
My mother (who, for reference, has otherwise only expressed AI concerns regarding how it affects education) just asked me about the Coxon news stories unprompted. And yes, I've tried having the conversation with her about it before. The difference here seems to be that this has been especially effective for highlighting how near-term the risks are and how much we don't have things under control, compared to e.g. the Statement on AI Extinction Risk. Per Molly Kinder:
... (read more)Of course, I have heard about "alignment" and "AI safety" and "x risk" concerns for years. I
Thank you for your very informative comment. I took the time to get familiar with the different info/sources you provided.
I think NVIDIA's chip-level Confidential Computing and Attestation tech is interesting. Their whitepaper here gives information on how this is implemented on their H100 Hopper GPU. Pages 14 and 15 outline what threats their tech would and would not protect against - you might find that interesting.
What they have however, is more for their hardware proving its identity as a legit NVIDIA GPU, providing detailed information about the runni... (read more)
Yep, called it. Thanks for citing the history. To me Kabir's strategy looked obviously correct just on left-wing instinct. In fact the opposite of quitting-in-protest, entryism (getting hired with an explicit goal to hijack the organization), is a known and effective tactic. The world would be a better place if alignment folks had that mindset when joining labs, instead of seeing themselves as helping the lab.
I think this makes good points. Definitely we shouldn't talk about p(doom) without an assumption about how we approach making AGI. I think the assumption that we'll just race flat out is breaking down, making more approaches more likely.
But your core point doesn't apply to most people, beyond a little slowdown.
If the the drive takes so long you'll die on the way if you don't drive 200 mph, it's not a mistake to risk it.
Humanity isn't remotely longtermist on average, nor are they utilitarian. They primarily care about themselves, and a few loved ones and f... (read more)
I am worried that these bills will mean basically nothing in practice, in the same way that a ban on "superexplosive missiles" may not have done anything to stop nuclear missile development. Similarly a ban on "big piles of uranium" depends a lot on what "big" means and whether putting 100 small parts together for a brief period long enough for them to explode would count.
As I see it, any proposal must be based on stopping development (of some class, like the training of new AIs) in general or stopping the means required for development (like compute bans... (read more)
I'm guessing it's just transfer from RLVR environments, particularly computer use. For all we know they literally train on playing (other) games.
Good point. A lot can be done at lightspeed. Hijack-shots at lightspeed might be the most efficient way to expand, if the universe has enough civilizations.
Lem's "His Master's Voice" (1968) described a pretty advanced version of this. An Earth divided between superpowers receives a signal that seems to give blueprints for a superweapon, and panicked scientists try to decide what to tell their bosses. Also it's mentioned that the signal might be modulated in a clever way to increase the chances of life arising from common chemicals. So it's a double whammy:... (read more)
Perhaps this is partly what you get if you promote reasoning in terms of expected values too hard, too generally (instead of "saner", less extremistan-conducive decision rules, or even heuriatics whatever "normal people use").
I don't know how much this has to do being overly neoclassical-econ-brained, but it's a live option for me that it does have non-trivial amount, whether it's in terms of actually skewing people's "honest thinking" or by giving them more concepts to produce palatable justifications for why they are riding the cancerous wave.
I also think it's worth pointing out that "being aware of whether or not you are in an eval" is something the labs often construe as dangerous in-and-of-itself, and although anthropic specifically has been trying to push against this stance for a while I don't know how good of a job they're doing.
I guess it is.
To me it seems obvious that the goal is success with value-aligned ASI. I sometimes forget that's not agreed upon by everyone who sees the potential of AI.
Value-aligned ASI may be difficult to achieve. I'm not actually sure it is. I've read all of the arguments; I don't think anyone knows. That's enough to make me want to be very cautious in attempting it. But it could turn out to be fairly easy if we just don't screw it up. We just need to go slowly enough to figure that out before we try it.
But that's a hard sell to people who aren't ASI-pi... (read more)
RLVR, to wit, only works in domains with easily verifiable rewards. Capabilities elicited this way then likewise refuse to generalize to other domains
I've been thinking about a similar proposition recently.
What do you make of the observation that successive generations of models seemed to improve at Claude Plays Pokemon, despite not being specifically trained on it? Likewise on this robotics benchmark done by Anthropic's frontier red team.
Another idea is to resign from OpenAI and then sue OpenAI. There must be something you can sue them for, right?
You could rightly say that they are endangering your life, which is grounds for a lawsuit. Whether a judge would go for that is another question. But anyone could file that lawsuit, not just an ex-employee. Maybe there's a better angle an ex-employee would have.
(fwiw, I think the things that make biting-culture work among friends is that they're... like, on the same page about what page they're on, so when they make snarky remarks about each other it's clear what kind/flavor o snark it is. I don't think it translates across people/groups that you know less well and it's more ambiguous whether you think each other are like slightly fucking up or doing really egregious things)
Nathan, maybe I am biased since I'm in here with you, but I have not failed to notice you caring about AI risk and talking about it, I haven't failed to notice that you didn't get a job at an AI company. So thank you. You're doing a good job.
Just spitballing here, maybe it's because a big part of what Claude does is write a ton of CoT and then condense it into something human-readable. Claude needs to prioritize a thought pattern of "let me find the important parts of this long thing I wrote", which means it needs to figure out:
etc.
Well, choosing the right positive vision is another point of disagreement, isn't it.
I think the lenses I mentioned are a nice simple way to look at it. Using the economic lens, place all technologies on a spectrum from human-complementing to human-substituting. Or with the biological lens, place all technologies on a spectrum from making humans more capable to creating a new species that's more capable than us.
Both of these suggest that we shouldn't go for "success with ASI" exactly. More like "success instead of ASI", or "making ourselves SI". Trying to m... (read more)
Nash equilibria are to game theory as null hypothesis significance testing is to psychology.
It's a method of analysis/sense-making (among many) that got reified as 'the obviously correct approach' by memetic evolution.
The messaging / PR from the labs has this weird aftertaste for a while now.
From an outsider perspective, I understand why people like to label everything that comes out from the labs as "marketing speak" including employee's statements.
I understand every press release has a hidden agenda to some degree, but if you can't dig deep, it's easier to have that prior and just ignore everything they say instead. It's also why their statements are hard to use as evidence for anything (especially their statements on AI risks)
but also I think "crunchtime" is overblown? like, actually, things have always been crunch-y, in that at any point in history, someone smart and thoughtful and hardworking could have had a great impact.
It does seem like people could have been more impactful in 2015 than in 2026, but also it was extremely hard to predict what the right thing to do would've been:
What exactly is the Safety and Security Committee doing and how much power does it have? Can the board employ safety researchers that work for the nonprofit and that aren't employees of the for profit?
Right. Communication is a separate skill.
I think one useful thing to do is recruit friends/acquaintances who are good communicators (and as a result probably have at least small platforms). Convincing them of the stakes can then tap their lifetime of practice and learning about communication.
It's also occurred to me that it's even harder to pitch the possible upsides of success with ASI. It's harder for people to believe that everyone could win than to believe that everyone can lose. But without a positive pull, it's a lot harder to think about and motivate to do this work.
I think labs should aim at the stricter goal of never having RL environments reinforce undesirable behavior. This will be extremely hard to do in full because it contains much of outer alignment, and much of the undesirable behavior will only appear OOD.
I’d describe the main pain as spreading understanding, and that understanding would result in everyone wanting to slow progress. Attempts to slow progress without spreading understanding might well backfire. Or not!
Yes, very much agree with this. I think the most effective efforts now should be as public-side as possible, not getting in smoky rooms with decision makers.
The big problem for me is that I can't talk or write interestingly about stuff I already know, it has to be something I'm in the process of figuring out, or just figured out five seconds ... (read more)
Maybe we should wait until the call, but I'm still confused about the timing. What's the relevant window of the prediction?
Like, suppose the AI is in a robot and decides to walk towards the shuttlecraft to travel to a lava planet where there's a 10% chance in the next year that it will accidentally melt and thus the shutdown mechanism will be disabled (along with the rest of the robot/AI). Does it get shut down before it takes the first step? If not, what if it's walking towards someone who intends to disable the shutdown mechanism (but not the rest of the... (read more)
FWIW it's not super obvious to me that you need "true creativity" (or whatever ??? capability that we think is missing from LLMs) to kill humanity off either; I'd written something along these lines there.
I do think you need a significant amount of crystallized competence[1] in domains like long-term strategy and metacognition/"keeping yourself on-track". Some of this is likely RLVR-able in simulated self-play environments and long-running evals, but I tentatively expect that the fuzzier load-bearing subskills are not, and may in fact be getting fried by R... (read more)
Thanks!
I'd describe the main pain as spreading understanding, and that understanding would result in everyone wanting to slow progress. Attempts to slow progress without spreading understanding might well backfire. Or not!
Sure AI can lead you on weird paths. Also YOU can lead you on weird paths! Talking to nobody is I think worse than talking to current-gen AI.
GPT4o and Gemini 2-2.5 use AI psychosis. Newer systems do not.
Fable is I believe a force for conceptual clarity. I hope Astra is similar. I believe GPT5 series was weakly in that direction.
But it's h... (read more)
In the early days of nuclear R&D there were criticality accidents. A lack of regulation plus an atomic race with other countries, led to experimentation, mistakes, and deaths.
We're in a similar situation, except instead of the research being confined to a small number of government labs that have managed to amass sufficient quantities of uranium or plutonium, we have a situation where any company that wants to train AI simply has to assemble enough GPUs. Worse, running said model can be done by anyone with a Mac Studio.
Regulating nuclear material was o... (read more)
Maybe I’m misreading, but is there nothing in this proposal that says when the reporting should be done in relation to model usage? This reads to me like an internal model could be used heavily for months before the suggested 6 month reporting/review interval hits. Or indeed, a model could be released before proper reporting! (not much unlike Astra, which presumably is the trigger here)