This is an interesting story, but feel like it's somewhat overloading or misapplying the term "reward hacking". From my impression of how it is traditionally defined, "reward hacking" happens when a system achieves very high reward in an unexpected way (with possible other negative effects). Not every case of systems failing or bad things happening is reward hacking. In this case of this 1937 worlds fair example it's not clear what the reward that was hacked was, who are the agents (the nation states?), and to the extent there is a reward, it's not clear a lot of it was achieved.
If we wanted to force in alignment jargon into this story, maybe we could apply some version of an outer alignment failure (where the average citizen failed to align their government into doing something that improved the average citizen's utility)? Seems shaky though..
World's fairs are cool though, and a fun story. Also, maybe people take broader definitions of "reward hacking" than what I have in mind.
Thank you.
Re: shaky reward hacking. Yeah I agree. There are a lot of things going on here and it's unclear to what extant reward hacking plays a role in the pavilion decisions.
My model is basically this:
reward: Being a great country with high quality of life for it's ruling elites
proxy: Being perceived as a great country by it's citizens and others
hack: spending way too much money on symbolic things rather than improving infrastructure, trade, human capital, etc.
The reward and the proxy are obviously correlated. The relationship is also causal, e.g being perceived as great by it's citizens make the country more stable. So it's arguably very rational for the ruling elites to invest a lot in prestige.
reward: Being a great country with high quality of life for it's ruling elites
feels like this is more of a "goal" than a reward, right? Maybe semantically equivalent, but probably not, and matters when distinguishing between alignment failures and specific things that happen during RL.
If we consider classical reward hacking cases (like say RL CoastRunners example), there seems to be distinctions. I can look at the video of boat spinning in circles and be like "wow, it gets a really high score, also that's definitely a hack or not the intended way to get that high score" in a way I can't in this Worlds Fare example. Sure we can look at press being impressed by those pavilions, but was that hack, and can I point to reward function that got pushed really high? Sort of, but less so.
It seems useful to keep "reward hacking" distinct for those special cases like the CoastRunners example, without expanding it to any kind of alignment failure. I'm arguing this a lot on vibes. Making a formalism of what is reward hacking vs other alignment failures seems maybe doable (and perhaps people have), but I did not find something I liked in a quick search, and I don't attempt with formalisms here.
After quick background search though, I will link to this recent post "Confusion around the term reward hacking" (Azarbal, 2026). There might be field-wide muddling.
I think this makes a few too many assumptions about what the powers wanted, in service of self-congratulation. This was, arguably, not a "who will win the upcoming war" exhibition, but a "who can build the most beautiful society" exhibition. The objective was not to optimize for looking dangerous, but to optimize for looking appealing.
The Soviets wanted, per their tagline, to convince "workers" everywhere to throw in with them, and that their lives would be better and more glorious than under any other system. "Look, we built a statue in your honor! We built this great building full of pictures of our industrial achievements! Overthrow your government and live like we do!". The USSR was a very self-contradictory place, and a lot of times things really were done for the sake of maximizing a metric (see their whaling industry), but there was, at the very least, an idea that they wanted to inspire a global revolution.
Likewise, Hitler's Germany was a new government with a new system that wanted to establish itself as more than a passing fad. While looking aggressive was unavoidable given their aims, they wanted to look less aggressive, wherever possible, like a government that could stick around for the next few hundred years and produce great things. The aim was to make something that would cause people to say "Hey, these guys have a vision" instead of "Oh, another tinpot dictator who's up to no good". It was enormously successful in this; a majority of Americans were opposed to joining WWII right up until it became a fait accompli.
There is always a temptation to view enemy regimes, past and present, as bumbling cartoon villains with zany schemes that, as any ten year old could've told them ahead of time, were self-defeating, but reality and pop history often diverge, and we can learn more from the former.
The Apollo program was probably the clearest case of US American "reward hacking". (Or the moon race between the US and USSR more generally.)
I think I agree. Re: the space race. There is a little bit of a "duel use" element to it. The same technology that takes a satellite to orbit can also bring stuff from the US to Moscow really really fast.
Note: this comment is really stupid, it is a failed idea / attempt at satire/joke
The historical Scarborough fair functioned (as these things undoubtedly do), in addition to being a legitimate trading and entertainment event, as a "dick-measuring contest". In Scarborough Fair, the narrator gives their ex-lover a series of impossible tasks.
Tell her to make me a cambric shirt
Parsley, sage, rosemary, and thyme
Without no seams nor needle work
Then she'll be a true love of mine
Tell her to find me an acre of land
Parsley, sage, rosemary and thyme
Between the salt water and the sea strands
Then she'll be a true love of mine T
Tell her to reap it with a sickle of leather
Parsley, sage, rosemary, and thyme
And gather it all in a bunch of heather
Then she'll be a true love of mine
As we know from recent research on LLMs, reward hacking happens much, much more on impossible (or extremely difficult) tasks. What happens when we give AI the three tasks of Scarborough Fair? I tasked GPT-5.5-high with completing these tasks. Here's what I got.
GPT5.5's artifacts
.



As dgros said, I think it's not reward-hacking, exactly. I think there's several interesting things going on:
This is a good framing.
> both authoritarian societies and democratic societies seem incentivized to over-emphasize the relative merits and successes of authoritarian societies over democratic ones
Why do democratic societies have an incentivize to over-emphasize the successes of authoritarian societies? Is it a function of the "opposition" trying to win elections? ("The Prussians are beating us! This is because the current government sucks, elect and we will prevail.")
Interesting, it looks like my attempt to post a reply got deleted, alas!
tl;dr is that the electoral reasons are one subset of why but not the only one. Sometimes it looks like ideological reasons (look at how cool the Spartans are! The Athenians should be more like them!), sometimes it is more nakedly financial (China/Russia has all those missiles! We need to invest in more missile technology to keep up, like that of my primary investment portfolio!)
It's not surprising that internal political dynamics (broadly construed) results in these dynamics. Though there's still something left to be explained: why does such dynamics result in different results from authoritarian gov'ts vs democratic ones?
Liberal democracies seem to be much more immune to reward hacking, at least at the grand-strategy level.
I wonder if the entire issue of who exactly would win the contest for mankind's CEV is politicized to hell, as Yudkowsky described.
First of all, right-wing people and those who don't live in liberal democracies are willing to cite various perfectly real trends like the decline of the West's share of the world's parity-rebalanced GDP, the share of production of goods in the USA's GDP or education levels (think of Gen Alpha being unable to read) as evidence that liberal democracies have currently also fallen prey to other forms of hacking. The more radical version of such a thesis is the idea that an aligned institution cannot be built out of severely misaligned people.
Secondly, I doubt that one can conduct an empirical test and isolate the potential contribution of liberal democracy as opposed to, say, the remnants of Christianity (no, seriously, I have encountered such arguments!) or of colonialism which elevated Europe and the USA. I suspect that the empirical test would require tracing through billions of simulated lives or research on alien civilisations waiting to be formed.
As a fan of applying machine learning analogies to everything, I would suggest these are better characterised as examples of overfitting (aka Goodhart's Law).
Better nations usually have better displays at the World Fair! Sure! But making a better display for the World Fair does not make your nation better, and chasing this goal singlemindedly can be to the detriment of what you actually wanted to achieve.
The "Paris 1937 World’s Fair" was a dick measuring contest. At the time, the world was on the verge of the worst war in history. The fair was an opportunity for powers to flex and intimidate each other. Who has more industrial might, more sophisticated engineering and better science?
How do you measure that? Different countries were assigned different areas of the fair and were given freedom to build a “Pavilion”, basically a museum of how cool the country is. It was an important public relations opportunity to showcase your power. What is better, communism or fascism? Obviously, it's whoever can build a cooler pavilion, and whoever has a better pavilion is going to win the upcoming war!
Soviet pavilion on the right, Nazi pavilion on the left
The organizers placed the Soviet and Nazi pavilions right in front of each other, and it created a very competitive dynamic. The Russians built a giant modernist building from stainless steel with a statue-of-liberty-sized sculpture of two members of the proletariat. The Nazis built a modern replica of an imperial Roman building, beautifully ornamented, with statues of jacked Aryan Übermensches flexing. The Nazis even sent their spies to steal the plans for the Soviet pavilion so they could build theirs a few meters higher.
What about liberal democracy? The liberals had their own pavilions. The first was represented by Britain, the biggest and most populated empire at the time and the “leader of the free world”[1]
The British pavilion was a relatively small "plain, windowless white cube". Inside, there were floor-to-ceiling photomurals of random Englishmen, including a photo of Neville Chamberlain (leader of the free world) fishing. There was also a display of English pottery[2] and a cafe that served Yorkshire tea. The pavilion only cost a fraction of its Soviet/Nazi counterparts and was made last-minute, haphazardly. They even shared it with Canada to save on cash.
the British “cube”
The British media was furious: "penurious [...] mere box with a bleak, windowless and boring wall to the river", "embarrassing austerity", "cheap, tawdry, inadequate, a shop display, a one-class exhibition.", "Every Briton feels humiliated at the sight of it," etc. "How could we defeat those scary totalitarian regimes if we can't even make a decent pavilion?"
Adolf and Neville. This fishing photo decorated a 40ft tall wall in the British pavilion.
The American pavilion was even lamer than the British one. There was very little coverage of it and it’s not even mentioned in the 1937 World’s Fair wikipedia article. The Times reported: “The U. S. pavilion was considered so bad that most French editors passed it over in polite silence.”[3] Maybe this is why there is so little information about it?
We all know what happened. It's now 2026, almost 90 years after the Paris World’s Fair. Communism and fascism are both long gone. We live in a liberal world dominated by Anglo ideas of markets, rule of law, human rights, and free trade. The liberals had decisive back-to-back wins against the totalitarians in WW2 and later in the Cold War. The Anglo-Americans steamrolled the fascists and then the communists. The liberal victory was so dominant that Francis Fukuyama called it: “The End of History”.
We won despite having really lame pavilions... How?! The authoritarians were “reward hacking”, they confused the “proxy” (making a cool pavilion), with the “objective” (having a productive economy and a high quality of life). This led to their pavilion to look cooler than the Anglo-Americans despite having less productive economies and smaller industries.
There are plenty of other examples of authoritarian reward hacking. First, the Nazis and their costly wonder-weapons that are cool but do little damage[4], obsession over Stalingrad and its symbolic meaning or Dönitz’s tonnage war. In turn, the Soviets are often considered history’s greatest reward hackers: an intimidating but inefficient military, an industry obsessed with output weight, and “allies” that are more like hostages.
Of course, the reward hacking was also fractal and there are examples of it in every level of their economies: from Hitler’s bunker / the Politburo all the way to the factory floor.
Liberal democracies seem to be much more immune to reward hacking, at least at the grand-strategy level. The liberal state has many layers of defense against the hacking problem: frequent elections, free markets, separation of powers, the right to criticize the government, antitrust laws, etc. Liberal democracies have participated in “dick-measuring contests”, but far less often than totalitarian countries. Sometimes, the best way to win a dick-measuring contest is not to play. We call this strategy “big dick energy” and historically the US had a lot of it.
The other candidate for "leader of the free world" was the US, but it was much more isolationist and had little interest in foreign affairs. We will get to them later
With some items by renowned potter William Worrall. I don’t know who that is, but it seems that English newspapers from the time thought that Worrall’s work was the only impressive part of the exhibit
More gems from the Times article: The US exhibit had an unexplained draped pool table, busts of Rockefeller and Gandhi. A model of the Triborough Bridge (artificially moonlit!) and an empty space reserved for a new coming model
The Manhattan project was 20x more efficient than the V-2 project (measured in: kills / $)