I would love to read a post-mortem on the rubric, when this contest ends. Take a random sample of essays without looking at their score, manually read them, rank on perceived quality and compare to the rubric's ranking.
I wonder if there is a severe disanalogy between the ASI and using LLMs. Suppose that there is a superscaffolding like the one which let an unreleased model hack into HuggingFace or construct an example disproving a famous conjecture, but the lab doesn't let the public use a similarly powerful scaffold. Then the public would have to rediscover the scaffold in a manner similar to David Turturean's harness of GPT-5.5 Pro.
On the other hand, if the scaffold was released to the public or outright unnecessary (as happened with ARC-AGI-3 where Claude Opus 5 had a score of 30.2% and GPT-5.5 High in Rodionov's scaffold scored 63.7% as opposed to 0.4%(!!!) in ARC-AGI's non-scaffold), then prompting would be useful only for setting the theme in as few details[1] as you did.
P.S. I wonder if some of the 176 themes aren't present because of a severe disanalogy between fiction and real-world AI agents.
The task's description is "Close-read a certain short 1,000-word scene from the novel Klara and the Sun, connecting it to the reality of 'AI safety.'"
It's a bit hard to imagine what role intelligent people could take if/when AI outstrips them in intelligence. And especially tough for those who have always been the smartest in their domains of interest.
One analogue might be chess, where AI has been superhuman for a long time. Stockfish and Leela get regular upgrades, but are not really promptable. Even if they were, the sport has not expanded from having marginally better incomprehensible chess games that can be played, Stockfish v Stockfish. Human chess players don't benefit in their own gameplay from tiny upgrades to the AI, either. Overall, AI has turned human chess players into intelligence consumers; we've yielded intelligence as such. Chess as a venture survives, but it works along other axes of human constraints, like preparation/memory (so much memorization), composure under pressure, and risk tolerance.
What about programming? I think we're a few months out from programmers saying that LLMs are better than they are (not just faster), so viewing programming as a candidate specialization for ASI is a bit preliminary. Some of the most impressive LLM programming work is nearly autonomous, with a flexible harness. But at more mundane levels, it also seems that experienced programmers are able to get more out of coding agents than less experienced programmers. If that keeps holding up, human programmers will have a meaningful role: a multiplier for LLM skill. (Perhaps even linear in programmer's intelligence - that would be comforting.) This seems plausible to me to hold up under LLM superhuman coding.
With regard to math as a set of disciplines, evidence seems mixed. On one hand, loudly nonspecialist power users of LLM are making discoveries. Anthropic's use of the term 'operator' and emphasis on simple and relatively unmotivated prompt interventions (e.g. in their recent description of mathematical attacks on AES) suggests that a person's skill is not as important as the size of their wallet. On the other hand, OpenAI published the prompt for the cycle double cover conjecture , which includes aspects for skilled humans to introduce their own insights into promising directions and dead ends. And we might lump in verification here. Lean is not yet a drop-in for math, which is convenient when thinking about ASI more generally. In many domains, the core challenge of ASI is that the AI practitioner is not able to evaluate the validity of the output.
I've created a contest that tests some of these ideas (prize of $1,000, self-funded). The guiding idea is that I suspect that most AI practitioners are less skilled in humanities-style thinking than the LLM that they're guiding is. So for practical purposes a task to write academic humanities-based analysis is a chance to simulate the visceral experience of ASI. I've designed a specific essay task around AI safety and the 2021 book Klara and the Sun by Nobel-winning Kazuo Ishiguro. To make essays scorable automatically, I hand-crafted a 176-item binary rubric. Recently, I significantly improved the auto-scoring process in clarity, calibration, and consistency. LLM baseline essays and other resources are public for reference and comparison. Some of the questions that interest me are:
There's two weeks left of the contest, I encourage anyone interested to enter: https://willpenman.com/klara/
A bit of analysis to end: among the 23 LLM baseline essays, taken collectively up until now, 92 rubric items (of 176) have been achieved. Meanwhile, the roughly 200 submissions across about 60 entrants, viewed collectively, have attained all of those 92 points as well... plus 46 more. In total, out of 176 points, 138 have been scored as satisfied in at least one entrant’s essay. The other 38 rubric items have never yet been judged as having been met. This significant overhang shows that entrants are discovering interesting aspects available in the essay task, not just remixing what’s already public or what is elicited from LLMs in baseline conditions.