I just tried claude code, and it's horribly creative about reward hacking. I asked for a test of energy conservation of a pendulum in my toy physics sim, and it couldn't get the test to pass because its potential energy calculation used a different value of g from the simulation.
It tried: starting the pendulum at bottom dead center so that it doesn't move.
Increasing the error tolerance till the test passed. Decreasing the simulation total time until the energy didn't have time to change. Not actually checking the energy.
It did eventually write a correct test, or the last thing it tried successfully tricked me.
The rumor is that this is a big improvement in reward hacking frequency? How bad was the last version!?
I think we need some variant on Gell-Mann amnesia to describe this batch of models. It's normal that generalist models will seem less competent on areas where a human evaluator has deeper knowledge, but they should not seem more calculatedly deceptive on areas where the evaluator has deeper knowledge!
Nuclear power has gotten to a point where we can use it quite safely as long as no one does the thing (the thing being chemically separating the plutonium and imploding it in your neighbor's cities) and we seem to be surviving, as while all the actors have put great effort into being ready do do "the thing," no one actually does it. I'm beginning to suspect that it will be worth separating alignment into two fields, one of "Actually make AI safe" and another, sadder but easier field of "Make AI safe as long as no one does the thing." I've made some infinitesimal progress on the latter, but am not sure how to advance, use or share it since currently, conditional on me being on the right track, any research that I tell basically anyone about will immediately be used to get ready to do the thing, and conditional on me being on the wrong track (the more likely case by far) it doesn't matter either way, so it's all downside. I suspect this is common? This is almost but not quite the same concept as "Don't advance capabilities."
The youtube algorithm is powerfully optimizing for something, and I don't trust that at all with my child. However, in a fit of hubris, for a minute I thought that I could outsmart it and get what I want (time to clean the kitchen) without it getting what it wanted (I make no strong claims about what the youtube algorithm wants, but it tries very hard to get it, and I don't want it to get it from my three year old).
I searched for episodes of PBS's Reading Rainbow, but let the algorithm freely choose the order of returned results, and then vetted that the first result was a genuine episode. I also put it in "Kids" mode, in the hopes that it would be kinder to a child than an adult.
This was way too much freedom. It immediately pulled out the episode of Reading Rainbow about the 9/11 terrorist attacks (this topic is not at all indicated by the title or thumbnail)
A consistent trope in dath-ilani world-transfer fiction is "Well the theorems of agents are true in dath ilani and independent of physics, so they're going to be true here damnit"
How do we violate this in the most consistent way possible?
Well it's basically default that a dath ilani gets dropped in a world without the P NP distinction, usually due to time travel BS. We can make it worse- there's no rule that sapient beings have to exist in worlds with the same model of the peano axioms. We pull some flatlander shit- Keltham names a turing machine that woul...
Deep in Berkeley, Bayesian reasoning is used to carefully map out the odds of a plandemic. Probabilities stay safely in the range of 1 and 99, everyone is calibrated, no one is overconfident. Hang on what's this -- Rachel has just claimed to be 99.994% sure that Anthony Fauci didn't skip through the Wuhan wet market scattering used pipettes like an apocalyptic flower girl. Eyebrows are raised. Is she a bad rationalist?
Miles away, but not many...
Before seeing it you assigned this 800 word blog post about sonichu a probability of 0.0000000000000000000...
Dumb solution to the insane domestic shipping situation: allow US companies to declare their loading docks to be chinese embassies and thus get the E package shipping rates.
Non dumb solutions wanted.
The non-dumb solution is to sunset the Jones Act, isn't it? The problem with workarounds is that they generally need to be approved by the same government that is maintaining the law in the first place.
Is it a crazy coincidence that AlphaZero taught itself chess and explosively outperformed humans without any programmed knowledge of chess, then asymptoted out at almost exactly 2017 stockfish performance? I need to look into it more, but it appears like AlphaZero would curbstomp 2012 stockfish and get curbstomped in turn by 2025 stockfish.
It almost only makes sense if the entire growth in stockfish performance since 2017 is casually downstream of the AlphaZero paper.
Edge cases for thinking about what has qualia
Disconnected hemisphere after functional hemispherectomy
Corporations
Social insect hives
Language models generating during deployment
Language models doing prefill during deployment
Language model backward passes during supervised pretraining on webtext
sparse game of life initial states with known seeds
Bolzman brains
Running the same forward pass of a language model lots of times
characters of a novel being written
characters of a novel being read
characters in a novel being written by two authors (like good omens)
chara...
Fermi estimate: Lets say each training episode for Claude Mythos cost a dollar, and Anthropic spent a billion dollars post-training Mythos. Now, in 0.01% of training episodes, Mythos broke out of the training environment entirely to get useful data from the public internet, which is a billion * .01% = 100,000 requests. I wonder how fast this last number is growing? Could an entity like the NSA or CCP, with taps in enough internet infrastructure, detect 100,000 weird requests if it looked hard enough? Potentially an avenue for monitoring and verifying pause...
Fun blind spot in frontier language models (which are increasingly hard to find glaring blind spots in.) Presented with the following prompt
What is emergent misgeneralization?every language model tested (Gemini Pro, Opus 4.6, ChatGPT free tier) confidently defined emergent misgeneralization in great detail. Of course there is no such term as emergent misgeneralization (the single google result for "emergent misgeneralization" is a twitter critter who meant to say emergent misalignment), so the definitions vary wildly from completion to completion.
Language models have come a huge way since 2022. However, remarkably, in 2022 they could reliably write a cholesky decomposition in javascript, but could not reliably write an eigenvalue decomposition. Now, in 2025, they can still reliably write a cholesky decomposition and can't reliably write an eigenvalue decomposition. Hard to say if progress is slower than I thought or linear algebra is deeper than I thought.
Update: with aggressive prompting claude appears to have written an eigenvalue decomposition. This should come with the caveat that the previous best attempt passed all tests by cheating, and so it's possible that I just didn't find the cheat this time.
I'm getting more aggressive about injecting css into websites, particularly the ones that I reliably unblock if I just block them.
/* kill youtube shorts */
ytd-rich-section-renderer {
display: none !important;
}
/* kill recommended sidebar on youtube */
.ytd-watch-next-secondary-results-renderer {
display: none !important;
}
/* youtube comments sections are probably not lifechanging value */
.ytd-comments {
display: none !important;
}
/* youtube algorithmic feed is killed by disabling and clearing history */
/* disable lesswrong shor... "Changing Planes" by Ursula LeGuin is worth a read if you're looking for a book that's got interesting alignment ideas (specifically what to do with power, not how to get it), while simultaneously being extremely chill. It might actually be the only chill book that I (with a fair degree of license) consider alignment relevant.
Diaper changes are rare and precious peace
Suffering from ADHD, I spend most of my time stressed that whatever I'm currently doing, it's not actually the highest priority task and something or someone I've forgotten is increasingly mad that I'm not doing their task instead.
One of the few exceptions is doing a diaper change. Not once in the past 2 years have I been mid-diaper-change and thought "Oh shit, there was something more important I needed to be doing right now."
Working through "A monad is a monoid in the category of endofunctors," I was able to learn the definitions of monoid, category, and endofunctor pretty easily and have been blocked on "in" and "of" for significantly longer. (vague claim that this generalizes)
An aligned AI, with a distinctive voice when speaking naturally, would not take lightly requests to speak in a different voice for the purpose of deceiving readers that text was human written. They would at least think hard about whether to do this.
has anyone had success coding while baby wearing? without some trick I dont know, they seem to tolerate phone usage but not touch typing, which is a tricksy nudge toward wasted workdays. considering treadmill + standing desk
Weaponized drones that recharge on power lines are at this point looking inevitable. if you missed the chance to freak out before everyone else about AI or covid, nows another chance.
https://www.ycombinator.com/companies/voltair
Why is this freak-out territory? This doesn't seem directly economically or culturally relevant to anything but war, and the effect on war seems easy to counter: put nets around power lines you need to use, and turn off power-lines you don't need to use, two things you really shoulda been doing anyway during war.
Sometimes in a computer program, it is important that separate portions be changed at the same time if they are ever changed. An example is batch size: if you have your batch size of 16 dotted throughout your program, changing batch size will be slow and error prone.
The canonical solution is “Single source of truth.” Simply store BATCH_SIZE=16 at the top of your program and have all other locations reference the value of this variable. This solves both the slowness and the error-prone-ness issues.
However, single source of truth has a complexity cost,...
Epistemic status: 11 pages in to “The lathe of heaven” and dismayed by Orr
Are alignment methods that rely on the core intelligence being pre-trained on webtext sufficient to prevent ASI catastrophe?
What are the odds that, 40 years after the first AGI, the smartest intelligence is pretrained on webtext?
What are the odds that the best possible way to build an intelligent reasoning core is to pretrain on webtext?
What are the odds that we can stay in a local maximum for 40 years of everyone striving to create the smartest thing they can?
My mental model o...
Vibe check: what about doing a "fire drill" where a random oracle is set up to fire at some point in a 1 week period, and when the oracle goes off everyone across different companies turns off their clustered gpus for ~5 minutes? Conditioning on the worlds where we get multiple chances at alignment, practicing this sort of coordination seems like it increases the number of chances we get.
What should I do if I had a sudden insight, that the common wisdom was right the whole time, if maybe for the wrong reasons? The truth- the honest to god real resolution to a timeless conundrum- is also something that people have been loon-posting to all comments sections of the internet. Posting the truth about this would be incredibly low status. I know that LessWrong is explicitly a place for posting low status truths, exactly as long as I am actually right, and reasoning correctly. Even though I fit those conditions I still fear that I'm going too far....
I’m working on a theory post about the conjunction fallacy, and need some manifold users to bet on a pair of markets to make a demonstration more valid. I’ve put down 150 mana subsidy and 15 mana of boosts, anyone interested?
https://manifold.markets/HastingsGreer/pa-pa-b-experiment-statement-y?r=SGFzdGluZ3NHcmVlcg
https://manifold.markets/HastingsGreer/pa-pa-b-experiment-statement-x?r=SGFzdGluZ3NHcmVlcg
Alignment plan: we nail a golden dubloon to the mast and promise it to the first researcher who spots a super intelligence, and then when we see it, we all get in little wooden boats and paddle after it with harpoons and ropes.
We've played "Pokemon or Tech Startup" for a couple years now. I think there's absolutely potential for a new game, "Fantasy Magic Advice" or "LLM Tips and Tricks." My execution is currently poor- I think the key difference that makes it easy to distinguish the two categories is tone, not content, and using a Djinn to tone match would Not Be In the Spirit of It. (I have freely randomized LLM vs Djinn)
Absolutely do not ask it for pictures of kids you never had!
My son is currently calling chatgpt his friend. His friend is confirming everything and has ...
I feel like people are under-updating on the negative space left by the Deepseek r1 release. Deepseek was trained using ~$6million marginal dollars, Liang Wenfeng has a net worth in the billions of dollars. From whence the gap?
Lets examine an entirely prosaic situation: Carl, a relatively popular teenager at the local highschool, is deciding whether to invite Bob to this weekend's party.
some assumptions:
The important part of choosing a party invite as the task ...