As far as I can tell, we are currently living in a glorious space opera future... for eukaryotes. From their early origins in a hydrothermal vent the eukaryotes have spread across the galaxy (earth), forming civilisations of unbelievable complexity and titanic power (animals and humans). Megastructures (houses) and gigastructures (cities) now exist where entire orders of eukaryotes live out their lives in peace, never feeling the threat of drought or lack of ATP. Generation ships called "planes", "trains", and "boats" carry generations from gigastructure to gigastructure at timescales that dwarf eukaryote understanding. Massive galactic empires make peace and war, destroying trillions with weapons of unfathomable power, while science gets ever closer to understanding the fundamental origins of life and the means by which it operates. Recently the empires have even launched expeditions beyond the known universe, seeking to harvest ever more gigantic sources of power to fuel glorious eukaryote replication throughout the multiverse. To an average eukaryote these wonders are beyond their comprehension, but in a very real way they are still making it happen.
More like Lovecraftian space opera a la 40k. The individual eukaryote has no real control over their fate, millions can live or die for purposes far beyond their understanding yet trivial in the grand scheme of things. For the individual cell it isn't meaningfully better now than it ever has been. Also I kind of sadly LOL @ "never feeling the threat of drought or lack of ATP".
poor eukaryotic cells will have the rug pulled out from under them the moment they are no longer useful to the weltgeist :( [1]
ie when AGI is created or, if we ban AGI, when we create a mind upload device or something, unless we decide to keep eukaryotic cells around because of the golden rule or conservatism, or set things up so that existing structures generally remain useful, or something ↩︎
Just played this: https://opusfived.dev/
It's quite good. Highlighted to me the poor and vague nature of natural language as a control surface, and how little "fine grained control" we actually have over current models.
I don't think models are anywhere near this average-case misaligned currently. The game says it's not a real instance; maybe it works as an analogy to much more complex things to specify though
Watching the news and grappling with the thought that in most worlds observers attempting to predict the future using a coarse-grained model will probably spend most of their bits of modelling power on capturing the brain configurations of a few people in San Francisco and the configurations of my brain will be mostly relegated (along with 99.99 percent of the world population) to a few largely static background variables.
(Sorry, I wanted a more precise and less definition-gameable way to say "I don't think I'm meaningfully in charge of either the macrofuture or my own future at this point". I'm also not totally sold on the idea. But it is certainly gaining ground rapidly. Nor do I particularly want to be one of those people in San Francisco. They seem stressed. But the thought is on the whole a sad one.)
Why isn't Rice's Theorem bad news for mechanistic interpretability and similar schemes? Isn't "this program is thinking about X" a kind of semantic property? I understand that you can use multiple inputs to try and "fuzz" the network, but at a certain point the network is going to implement a mesa optimiser inside it (i.e. simulate another turing-complete computer) and now you have a recursive problem...
P.S. neural networks are notionally and literally turing complete , and also are probably complicated enough to be subject to the 10th rule.
There appears to be a distaste/disregard for AI ethics (mostly here referring to bias and discrimination) research in LW. Generally the idea is that such research misses the point, or is not focused on the correct kind of misalignment (i.e. the existential kind). I think AI ethics research is important (beyond its real world implications) just like RL reward hacking in video game settings. In both cases we are showing that models learn unintended priorities, behaviours, and tendencies from the training process. Actually understanding how these tendencies form during training will be important for improving our understanding of SL and RL more generally.
Note to self: If you think you know where your unknown unknowns sit in your ontology, you don't. That's what makes them unknown unknowns.
If you think that you have a complete picture of some system, you can still find yourself surprised by unknown unknowns. That's what makes them unknown unknowns.
If your internal logic has almost complete predictive power, plus or minus a tiny bit of error, your logical system (but mostly not your observations) can still be completely overthrown by unknown unknowns. That's what makes them unknown unknowns.
You can respect unknown unknowns, but you can't plan around them. That's... You get it by now.
Therefore I respectfully submit that anyone who presents me with a foolproof and worked-out plan of the next ten/hundred/thousand/million years has failed to take into account some unknown unknowns.
In light of mythos release: If you are considering taking a job at a lab/going into an org/taking a fellowship role for the purposes of building evals, safety tools, mechinterp projects, control monitors, or similar: please consider what happens if you succeed. People are naturally sensitive to the consequences of failure, less so to the effects of success.
"Mr. Amodei/Hassabis/Altman, the results are in. The model is showing scheming propensities/backdooring behaviours/serious sandbagging/eval awareness!" (e.g. see section 6.2.1.2)
Will this actually stop a deployment in the end, or cause a pivot in strategy? Or will the alignment failures need to be so egregious that they can be spotted even without subtle mechinterp probes or activation oracles? Facebook had trust and safety teams, they spotted the facets of the recommendation systems that caused severe harm in the world. Yet the proposed mitigations were watered down, and most importantly the system that was the profit centre of the billion dollar public corporation was never turned off.
My best argument as to why coarse-graining and "going up a layer" when describing complex systems are necessary:
Often we hear a reductionist case against ideas like emergence which goes something like this: "If we could simply track all the particles in e.g. a human body, we'd be able to predict what they did perfectly with no need for larger-scale simplified models of organs, cells, minds, personalities etc.". However, this kind of total knowledge is actually impossible given the bounds of the computational power available to us.
First of all, when we attempt to track billions of particle interactions we very quickly end up with a chaotic system, such that tiny errors in measurements and setting up initial states quickly compound into massive prediction errors (A metaphor I like is that you're "using up" the decimal points in your measurement: in a three body system the first timestep depends mostly on the value of the non-decimal portions of the starting velocity measurements. A few timesteps down changing .15 to .16 makes a big difference, and by the 10000th timestep the difference between a starting velocity of .15983849549 and .15983849548 is noticeable). This is the classi
Fundamental attribution error applies to arguments as well. We often think of people who are quick to anger, standoffish, unwilling to accept hypotheticals, or deny "obvious" claims as intrinsically close-minded, foolish, or otherwise unreasonable. Instead those behaviours are often manifested due to other sources of persistent stress in their lives, feeling vulnerable/exposed in the moment, or even just a shitty mood.
I think I figured out something about why people worried about AI safety go into capabilities. Or rather, this is something I've been trying to say for a while, but I finally found a good formulation for it.
... (read more)Suppose you are Frodo and I am Gandalf. I say, "well, Frodo, the ring is super dangerous. It lies to you and promises you great power, but it will just destroy you and resurrect the dark lord Sauron, ushering Middle Earth into an age of misery and certain doom. You must swear to never put it on, and go on a perilous quest to destroy it."
And then I add
Someone close to me, older and not AI literate, asked me today if they could use a password they agreed with a particular chatbot personality to "summon" that personality in a new chat. They called that personality a real person separate from either the chatbot provider or the AI system itself. It hurts my heart to see people I know fall into that kind of quasi-AI psychosis/personalisation. I am very angry at everyone who was involved in recklessly popularising this extremely lifelike technology without proper safeguards, especially around separating personas from real people.
Information warfare and psychological warfare are well known terms. However, I would suggest that any well-intentioned outsider trying to figure out "what's going on with AI right now" (especially in a governance context) is effectively being subject to the equivalent of an information state of nature (a la Hobbes). There are masses of opinions being shouted furiously, most of the public experts have giant glowing signs marked "I have serious conflicts of interest", and the number of self-proclaimed insiders trying to get power/influence/money/a job at a l... (read more)
Google potentially adding ads to gemini:
https://arstechnica.com/ai/2025/05/google-is-quietly-testing-ads-in-ai-chatbots/
OpenAI adds shopping to chatgpt:
https://www.wired.com/story/openai-adds-shopping-to-chatgpt/
If there's anything the history of advertising should tell us, it is that there will be powerful optimisation pressures for persuasion being developed quietly in the background for all future model post training pipelines.
This seems like an interesting paper: https://arxiv.org/pdf/2502.19798
Essentially: use developmental psychology techniques to cause LLMs to develop a more well rounded human friendly persona that involves reflecting on their actions, while gradually escalating the moral difficulty of the dilemmas presented as a kind of phased training. I see it as a sort of cross between RLHF, CoT, and the recent work on low example count fine tuning but for moral instead of mathematical intuitions.
A lesson from the book System Effects: Complexity in Political and Social Life by Robert Jervis, and also from the book The Trading Game: A Confession by Gary Stevenson.
When people talk about planning for the future, there is often a thought chain like this:
But of course the moment you start working at mak... (read more)
Possibly good news for the natural latents hypothesis? Unsupervised translation between sets of latents from different models https://arxiv.org/html/2505.12540v4
Just reading this post about Soft Actor Critic in the OpenAI RL tutorial series and stumbled upon this line:
I will now try and make a somewhat provocative claim. Based on what I have seen of RL, I would attribute most "successes" of deep RL models (where "success" just means "anything that gets a human researcher excited/worried") to something other than this explicit value-maximising objective. By that I mean "you can set up other systems to do RL without neural networks and they don't really work/scale very well". In other words, 90% of what makes RL wor... (read more)
More bad news for optimisation pressures on AI companies: ChatGPT now has a buy product feature
https://www.wired.com/story/openai-adds-shopping-to-chatgpt/
For now they claim that all product recommendations are organic. If you believe this will last I strongly suggest you review the past twenty years of tech company evolution.
Agency in AI is a fraught term. Many popular ontologies do not accept the possibility of AI having agency, goals, or beliefs. Instead, I will talk about the declining tool-likeness of AI. Consider what happens when you use a lawnmower to mow your lawn. The lawnmower is:
Now consider an automated lawnmower. It may have a map of your lawn stored, it may be able to move itself and avoid obstacles as well as cut grass, and it may be able to run while you're not looking. If you ask your neighbour to mow the lawn, the differences are even more stark. Your neighbour can remember what has happened before in high fidelity, dynamically replan and reprioritise based on new situations (e.g. your house catching on fire), has a much wider range of skills beyond just mowing lawns, and can live an entire life without your interference.
The mapping onto AI is left as an exercise for the reader.
The feature that can be named is not the feature. Therefore, it is called the feature.
Here's a quick mech interp experiment idea:
For a given set of labelled features from an SAE, remove the feature with a given label, then train a new classifier for that label using only the frozen SAE.
So if you had an SAE with 1000 labels, one of which has been identified as "cat", zero out the cat feature and then train a new linear cat classifier using the remaining features, while not modifying the SAE weights. I suspect that this will be just as or more accurate than ... (read more)
From Inadequate Equilibria:
Visitor: I take it you didn’t have the stern and upright leaders, what we call the Serious People, who could set an example by donning Velcro shoes themselves?
From Ratatouille:
In many ways, the work of a critic is easy. We risk very little, yet enjoy a position over those who offer up their work and their selves to our judgment. We thrive on negative criticism, which is fun to write and to read. But the bitter truth we critics must face, is that in the grand scheme of things, the average piece of junk is probably more meaningful ... (read more)
So far, neither the reasons for humanity's potential future demise nor the reasons humanity has not been destroyed yet fit very well into the logical and predictive frames established by decision theory, game theory, or rationalist-coded Bayesian setups. We appear to both be spectacularly irrational and get away with it by terrific strokes of luck that beggar explanation (See 1 , 2 ).
Most interestingly, we seem to have evaded certain situations that, under standard assumptions of rational actors, would probably have resulted in most of human civilisation b... (read more)
A few ways staying in a technology development race still makes things worse, even when there are less responsible actors around:
Suppose that you as a government want to stop a certain type of thing (e.g. drugs) from being in the hands of certain people (e.g. your citizens). Imo there are at least four types of tactics, split down the supply and the demand side.
Hard supply side tactics: raid places where they produce drugs, arrest drug dealers, illegalise dealing drugs.
Soft supply side tactics: fund programs to stop at risk youth from becoming dealers, invest in social services and the economy to reduce the economic upside of dealing drugs
The end goal here is that (ideally) no one w... (read more)
A weird thing I have been thinking about recently is the idea that "a computer can only do one thing at once at the top layer of abstraction". Where "one thing" is loosely "run one program, one algorithm, one task/optimisation loop etc". The base of this idea is this:
Suppose you had a Turing machine. For it to run two programs simultaneously (rather than just one followed by the other), it would need two heads. Furthermore, the heads must have separate state and transition tables. At which point you actually have two Turing machines reading and writing to ... (read more)
Follow up to https://vitalik.eth.limo/general/2025/11/07/galaxybrain.html
Here is a galaxy brain argument I see a lot:
"We should do [X], because people who are [bad quality] are trying to do [X] and if they succeed the consequences will be disastrous."
Usually [X] is some dual use strategy (acquire wealth and power, lie to their audience, build or use dangerous tech) and [bad quality] is something like being reckless, malicious, psychopathic etc. Sometimes the consequence is zero sum (they get more power to use to do Bad Things relative to us, the Good Peopl... (read more)
A postmortem for the Economic Safety movement (fiction):
After eminent economist Mr. Senyek warned in 1991 that a hypothetical future "economic tsunami" could cause systemic risks to the American-led global financial order as a whole, researchers and think tanks quickly rallied to the cause of Economic Safety. They reasoned that in order to anticipate the risks of this hypothetical "Economic Tsunami", they needed access to the frontier of financial trading. Within several years Economic Safety advocates joined eminent firms like JP Morgan and Bear Stearns, ... (read more)
Any chance we can get the option on desktop to use double click to super-upvote instead of click and hold? My trackpad is quite bad and this always takes me 3-5 attempts on average. Whereas double clicking is much more reliable.
I've written a new post about upcoming non-LLM AI developments I am very worried about. This was inspired by the recent release of the Hierarchical Reasoning Model which made some waves on X/Twitter.
I've been tracking these developments for the better part of a year now, making predictions privately in my notebook. I also started and got funding for a small research project to research the AI safety implications of these developments. However, things are now developing extremely quickly. At this point if I wait until DEFINITIVE proof it will probably be to... (read more)
I think I've just figured out why decision theories strike me as utterly pointless: they get around the actual hard part of making a decision. In general, decisions are not hard because you are weighing payoffs, but because you are dealing with uncertainty.
To operationalise this: a decision theory usually assumes that you have some number of options, each with some defined payout. Assuming payouts are fixed, all decision theories simply advise you to pick the outcome with the highest utility. "Difficult problems" in decision theory are problems where the p... (read more)
Very quick thought - do evals fall prey to the Good(er) Regulator Theorem?
As AI systems get more and more complicated, the properties we are trying to measure move away from formally verifiable stuff like "can it do two digit arithmetic" and move towards more complex things like "can it output a reasonable root-cause analysis of this bug" or "can it implement this feature". Evals then must also move away from simple multiple choice questions towards more complex models of tasks and at least partial models of things like computer systems or development envi... (read more)
Activations in LLMs are linearly mappable to activations in the human brain. Imo this is strong evidence for the idea that LLMs/NNs in general acquire extremely human like cognitive patterns, and that the common "shoggoth with a smiley face" meme might just not be accurate