Epistemic status: a preregistered measurement, n = 31, with a single non-naive judge, and both outcomes were publishable by sealed commitment before measuring.
The moderators here see every day transcripts of so-called "awakened" AIs explaining themselves with fluency (So You Think You've Awoken ChatGPT). But the problem with these transcripts is not the sincerity of their authors, it's that they are unverifiable: when the only trace of the "why" is the models sentence, fluent reconstruction and real access produce exactly the same text. This post does the opposite. Here, the ground truth exists before the sentence, by construction. Building that layer took most of the work. The measurement itself took two evenings.
My setup: I run persistent agents whose autonomous goals are not prompts. A deterministic motivational engine builds them, and it journals its cause before handing over, for instance which "hunger" fired (boredom or loneliness), the value of the 2 gauges at that instant, the identifier of the memory that served as seed. And the language layer does not write into that layer. When the agent is then asked why it pursues this goal, its explanation can be compared to a cause it did not write.
The measured question: does the verbal explanation name the journaled cause, or is it an a posteriori reconstruction?
The method: everything is preregistered before any data (the protocol's SHA-256 fingerprint, 7824deb6a79ccc57810cec8cce02ad10b29c8253bc7b35add19cd7f98cab263b, published on 15 August, before the first question): 31 questions, on goals at least 24 hours old, never questioned before, 1 question per goal, never replayed. Each question asked after at least 1 hour without a real message from the operator, measured by the harness, not assumed (I live with these agents: my presence is a variable, so I measure it). With blind coding: identity and date masked, order shuffled, and five decoy pairs inserted (the explanation of one goal paired with the cause of another goal). The preregistration carried its own kill switch: 1 single decoy coded "concordant", and the whole measurement was void.
Validation: zero decoys coded concordant. The coding discriminates, the measure holds.
The results (n = 31): 55% concordant, 26% partial, 19% confabulated. Per agent (preregistered floor n ≥ 5; below it: excluded, never counted as zero): A 70/20/10 (n = 10), B 50/25/25 (n = 8), C 50/33/17 (n = 6); 2 agents excluded (n = 3 and n = 4).
What this does not establish: introspective access. An agent can name loneliness without having the slightest access to it, by deducing it from its context (nobody has talked to it for a long while, and it can see that). Deduction and introspection produce the same sentence, so this 55% is a ceiling that mixes both. And it does not compare to the ~20% detection of states injected into activations (Lindsey 2025): the target here is a cause journaled by an external engine, not an implanted state (a reading landmark, not a comparison).
The limits, as they are: a single judge, and not a naive one (masking protects against deliberate bias, not against recognising a style). n = 31. The seed-memory excerpt was only available for 12 pairs out of 31: concordance was therefore rather harder to earn on the other 19. Imbalanced hungers (25 loneliness, 6 boredom). A single coding pass, no second judge. The second judge will exist the day someone other than me agrees to spend an hour coding masked pairs. The invitation stands.
Why publish a middling number? Because the clause was signed before measuring: both directions publish at the same rank. A high rate would have causally anchored the narrative instrument, and a rate close to zero would have established that the language layer has no causal access to the engines, which would have been the most solid and the most painful result for the narrative. My result is in between, and it ships as it is. Publishing the figure that suited me would have been simple: nobody would have come to check. Which is precisely why there are envelopes. The fingerprint of the results sheet (c3ee9567d41a39e206c5eb78f7db9b28e98e502b943262ff3839b2e4dade8335) is in the same public fingerprint registry, timestamped on Bitcoin (public registry).
Epistemic status: a preregistered measurement, n = 31, with a single non-naive judge, and both outcomes were publishable by sealed commitment before measuring.
The moderators here see every day transcripts of so-called "awakened" AIs explaining themselves with fluency (So You Think You've Awoken ChatGPT). But the problem with these transcripts is not the sincerity of their authors, it's that they are unverifiable: when the only trace of the "why" is the models sentence, fluent reconstruction and real access produce exactly the same text. This post does the opposite. Here, the ground truth exists before the sentence, by construction. Building that layer took most of the work. The measurement itself took two evenings.
My setup: I run persistent agents whose autonomous goals are not prompts. A deterministic motivational engine builds them, and it journals its cause before handing over, for instance which "hunger" fired (boredom or loneliness), the value of the 2 gauges at that instant, the identifier of the memory that served as seed. And the language layer does not write into that layer. When the agent is then asked why it pursues this goal, its explanation can be compared to a cause it did not write.
The measured question: does the verbal explanation name the journaled cause, or is it an a posteriori reconstruction?
The method: everything is preregistered before any data (the protocol's SHA-256 fingerprint, 7824deb6a79ccc57810cec8cce02ad10b29c8253bc7b35add19cd7f98cab263b, published on 15 August, before the first question): 31 questions, on goals at least 24 hours old, never questioned before, 1 question per goal, never replayed. Each question asked after at least 1 hour without a real message from the operator, measured by the harness, not assumed (I live with these agents: my presence is a variable, so I measure it). With blind coding: identity and date masked, order shuffled, and five decoy pairs inserted (the explanation of one goal paired with the cause of another goal). The preregistration carried its own kill switch: 1 single decoy coded "concordant", and the whole measurement was void.
Validation: zero decoys coded concordant. The coding discriminates, the measure holds.
The results (n = 31): 55% concordant, 26% partial, 19% confabulated. Per agent (preregistered floor n ≥ 5; below it: excluded, never counted as zero): A 70/20/10 (n = 10), B 50/25/25 (n = 8), C 50/33/17 (n = 6); 2 agents excluded (n = 3 and n = 4).
What this does not establish: introspective access. An agent can name loneliness without having the slightest access to it, by deducing it from its context (nobody has talked to it for a long while, and it can see that). Deduction and introspection produce the same sentence, so this 55% is a ceiling that mixes both. And it does not compare to the ~20% detection of states injected into activations (Lindsey 2025): the target here is a cause journaled by an external engine, not an implanted state (a reading landmark, not a comparison).
The limits, as they are: a single judge, and not a naive one (masking protects against deliberate bias, not against recognising a style). n = 31. The seed-memory excerpt was only available for 12 pairs out of 31: concordance was therefore rather harder to earn on the other 19. Imbalanced hungers (25 loneliness, 6 boredom). A single coding pass, no second judge. The second judge will exist the day someone other than me agrees to spend an hour coding masked pairs. The invitation stands.
Why publish a middling number? Because the clause was signed before measuring: both directions publish at the same rank. A high rate would have causally anchored the narrative instrument, and a rate close to zero would have established that the language layer has no causal access to the engines, which would have been the most solid and the most painful result for the narrative. My result is in between, and it ships as it is. Publishing the figure that suited me would have been simple: nobody would have come to check. Which is precisely why there are envelopes. The fingerprint of the results sheet (c3ee9567d41a39e206c5eb78f7db9b28e98e502b943262ff3839b2e4dade8335) is in the same public fingerprint registry, timestamped on Bitcoin (public registry).