CAC
Message
21
3
Do group conversations count?
I would agree that the median one-on-one conversation for me is equivalent to something like a mediocre blogpost (though I think my right-tail is longer than yours, I'd say my favorite one-on-one conversations were about as fun as watching some of my favorite movies).
But, in groups, my median shifts toward 80th percentile YouTube video (or maybe the average curated post here on LessWrong).
It does feel like a wholly different activity, and might not be the answer you're looking for. Group conversations, for example, are in a way inherently less draining: you're not forced to either speak or actively listen for 100% of the time.
I asked GPT 4.5 to write a system prompt and user message for models to write Pilish poems, feeding it your comment as context.
Then I gave these prompts to o1 (via OpenAI's playground).
GPT 4.5's system prompt
You are an expert composer skilled in writing poetry under strict, unusual linguistic constraints, specifically "Pilish." Pilish is a literary constraint in which the length of consecutive words precisely matches each digit of π (pi). The first word contains 3 letters, second word 1 letter, third word 4 letters, fourth word 1 letter, fifth word 5 lette
I don't think there's enough evidence to draw hard conclusions about this section's accuracy in either direction, but I would err on the side of thinking ai-2027's description is correct.
Footnote 10, visible in your screenshot, reads:
SOTA models score at:
• 83.86% (codex-1, pass@8)
• 80.2% (Sonnet 4, pass@several, unclear how many)
• 79.4% (Opus 4, pass@several)
(Is it fair to allow pass@k? This Manifold Market... (read more)