The Turing test can't determine that a human is intelligent if he only speaks French,and you don't. In general, if the human is sentient but can't communicate well, it's going to have a hard time. This does not mean that any AI is as intelligent as a human who doesn't communicate well.
There's also the fact that the Turing test works because humans can modify their questions based on the entity's responses, to try to poke at specific aspects of their intelligence that couldn't be predicted in advance, based on the human's own knowledge of how humans behave (and in this case specifically toddlers). It doesn't sound like you did this in your tests.
(I'm saying this, of course, about questions designed to figure out if the entity is intelligent, not shibboleths like seeing how many times they refer to toys and pizza.)
Thank you, Jiro. You are right about the asymmetry; I should have made it explicit: the Turing test was proposed as (roughly) a sufficient condition, so: a toddler failing it doesn't break the test. Conceded!
What I'm pushing on is how the test gets used in the debate: it is invoked as if it sorted minds from "non-minds", and it can not do that even in principle: passing is (maybe) informative, failing tells you nothing. My daughter and your French speaker both land in the "no information" bucket, while the graphics card lands in "yes" (!). That may be the "instrument" working as designed, but then it's not doing the job which the public debate hires it for. Turing himself was careful about this: imitation, not thinking.
On "adaptive questioning": fair, again. My toddler experiments were: not blinded, and not adaptive, and heavily contaminated by my affection for the subjects. The serious half of the claim is about "stochastic parrot" and "emergence" used as criteria, where no adaptive protocol exists at all!
About "shibboleths": I agree. Counting the "pizza" references gets my daughter identified with 99.99% accuracy; but I suspect there could be some overfitting here ;^)
In July 2022 I was in a parking lot with a Portuguese colleague, trying to fix the cargo-metering system of a 12-ton tanker truck. During a break I read a headline on my phone: Google engineer claims experimental AI went sentient. An engineer (like me!), from Google, testing an AI (I do tests too!), claimed it had become sentient.
I could not believe it. LaMDA was describing itself as a globe of light, claiming fear of being shut down, meditating during the long pauses between chats. If I were an artificial intelligence, I would definitely not present myself as a scared light bulb doing yoga; still, the thing didn't fade for me with the online hype. The experts said: "Eliza effect", "stochastic parrot", a machine repeating words in sequences made plausible by maniacal statistical matching. Fine; but I wanted to understand that answer, not repeat it as a parrot, and everything I found either stopped at the pop-science mantras or assumed I already knew the whole thing.
So I wrote my own transformer engine, in C language, from scratch. It took about 18 months (during lunch breaks and weekend nights) (TRiP, on GitHub). It runs the weights of Gemma, Llama, GPT-2, and PaliGemma for vision, it does inference and training, and it's CPU-slow. I learned what I wanted: what attention actually does, what the KV cache is for, and that half the work is not the engine but connecting it to the wheels.
This post of mine is not about the engine, though:
while I was building TRiP, two other systems were being trained (at home): my daughter Sofia (two years old at the start) and my son Paolo (born in 2023). And I noticed that every litmus test we dip into AI, I could also dip into my children. The results were... embarrassing?, in both directions.
The stochastic parrot
Sofia at 2 spoke constantly, fluently, and often incomprehensibly. Not mispronounced, but structurally opaque. You didn't understand the purpose of her sentences; you didn't understand their meaning; you didn't even understand many of the words.
As a child I had read about software fed with the statistics of the English language, able to generate text that looks plausible at first glance and turns out to be garbage when you actually try to understand it (the stochastic parrot, in short). Sofia was just the same to me... no, wait. She was quite the opposite. At first glance her output was pure garbage. My small parrot could generate incomparably colourful anti-stochastic sequences that would make LaMDA perform a self-shutdown.
So whatever "produces statistically plausible token sequences" measures, my daughter failed it. Relevant.
Emergent properties
One night dinner was polenta with sausage gravy. Paolo, not yet two, whose longest recorded utterance until that minute had probably been "banana", was face-deep in the dish, processing every bit of it in a continuous stream, polenta up to his hair and down to the diaper. Then he stopped, all of a sudden. He raised his face, painted in red, white and yellow, looked at us with complete seriousness, and pronounced two solemn words, perfectly spelled:
Then he submerged back into the dish. A new, central, self-defining property, which was not there two minutes before, had just emerged: from sausage and polenta.
Someone could argue that it was only a surge in a mass of growing neurons; or maybe just the effect of an increasing accumulation of polenta. These happen to be the same two positions available in the debate about emergence in LLMs.
One-shot learners
One evening at dinner I tried to teach Sofia the concept of setting a good example. "It's when you show Paolo that you pick up your things, and he learns to do the same by looking at what you do." She followed, so I moved to the negative case: "And what is it to set a bad example instead? I come to you and say: SLAVE! PICK UP MY THINGS FOR ME!", and I gave her a slow, funny slap on the cheek.
Kids are one-shot learners: Sofia slid off her seat (Paolo already waiting, like he'd read the script) and gave him a full Hollywood backhand. Paolo answered with a hammer slap on her head. In a few seconds, my academic lesson on phenomenological ethics had derailed into a slap fight, Bud Spencer style, and both of them were laughing like crazy.
People who work on alignment will recognize the failure: the demonstration was the training signal, and the "bad example" label around it was not. I ran this experiment once; not sure whether I'll be gathering more data.
My point
Everybody goes to GPT and asks if it feels happy. I went to my family instead, and used the same litmus paper we normally dip into AI. Here's what I found:
My conclusion is somehow narrow. It's not "LLMs are like children", and not "children are like LLMs" either. It's that these tests work (as descriptions) and fail (as discriminators). If a test cannot distinguish my daughter from a graphics card, whatever it measures is not the thing we were arguing about, when we invoked it. The Turing test was deliberately about the imitation of intelligence, not about thinking machines; 70 years later we got the imitators, but the debate reopened instead of closing. The "stochastic parrot" describes a mechanism; as a criterion, it catches my two-year-old. "Emergence" gives a name to a discontinuity after it has happened.
...and something I can't explain...
Sofia asks for a story every single night. She asks for the story, but it's not about the story. The story may be flowers, rats, rainbows, sandwiches; she doesn't care. She wants me to be with her. I've had late-night chats with AIs about the deep meanings of life. I got powerful responses, and wrote powerful insights back. But no AI has ever asked me to tell it a story. Models trained on more or less all recorded human output reproduce the stories very well, and in four years I have never seen one reproduce the need itself; not even as a glitch! I don't have a theory of why. I'm just flagging the datum (maybe my confusion, as well).
I stop here
I never claimed that AI is a person. But after the months inside the engine and the years with the toddlers, I've landed here: in both cases there is something before which the honest move is to stop and listen, trying to understand, instead of forcing it into the rows and columns of a spreadsheet ahead of the evidence. This is what I ask for myself, and I'm willing to extend it in both directions.
And the debate about the nature of AI is, in truth, also about us: whether we are worthy because of our performance, or simply because we are; whether our freedom is only a poetic reading of residual randomness, or something more.
I wrote a short book about all this (My TRiP through AI), part memoir, part technical field notes. The argument above is the part I'd like to stress-test here: where does it break? If there's a version of "stochastic parrot" or "emergence" that cleanly separates the toddler from the transformer, I'd like to hear it.