I ran the standard AI litmus tests on my two toddlers (yep) Engineer Carlo Valenti built his own transformer engine from scratch in C over 18 months to understand AI claims of sentience, then ran the same litmus tests on his two toddlers, finding that his daughter's speech at age two was structurally opaque like a stochastic parrot but in reverse, and his son's sudden emergence of language mirrored debates about emergence in LLMs. In July 2022 I was in a parking lot with a Portuguese colleague, trying to fix the cargo-metering system of a 12-ton tanker truck. During a break I read a headline on my phone: Google engineer claims experimental AI went sentient. An engineer like me , from Google, testing an AI I do tests too , claimed it had become sentient. I could not believe it. LaMDA was describing itself as a globe of light, claiming fear of being shut down, meditating during the long pauses between chats. If I were an artificial intelligence, I would definitely not present myself as a scared light bulb doing yoga; still, the thing didn't fade for me with the online hype. The experts said: "Eliza effect", "stochastic parrot", a machine repeating words in sequences made plausible by maniacal statistical matching. Fine; but I wanted to understand that answer, not repeat it as a parrot, and everything I found either stopped at the pop-science mantras or assumed I already knew the whole thing. So I wrote my own transformer engine, in C language, from scratch. It took about 18 months during lunch breaks and weekend nights TRiP, on GitHub https://github.com/carlovalenti/TRiP . It runs the weights of Gemma, Llama, GPT-2, and PaliGemma for vision, it does inference and training, and it's CPU-slow. I learned what I wanted: what attention actually does, what the KV cache is for, and that half the work is not the engine but connecting it to the wheels. This post of mine is not about the engine, though: while I was building TRiP, two other systems were being trained at home : my daughter Sofia two years old at the start and my son Paolo born in 2023 . And I noticed that every litmus test we dip into AI, I could also dip into my children. The results were... embarrassing?, in both directions. Sofia at 2 spoke constantly, fluently, and often incomprehensibly. Not mispronounced, but structurally opaque. You didn't understand the purpose of her sentences; you didn't understand their meaning; you didn't even understand many of the words. "It's so good this pizza wood " "Why does grandpa wear a hat? My cats at will can don't " "Daddy, I want to sleep, shall we play?" As a child I had read about software fed with the statistics of the English language, able to generate text that looks plausible at first glance and turns out to be garbage when you actually try to understand it the stochastic parrot, in short . Sofia was just the same to me... no, wait. She was quite the opposite. At first glance her output was pure garbage. My small parrot could generate incomparably colourful anti-stochastic sequences that would make LaMDA perform a self-shutdown. So whatever "produces statistically plausible token sequences" measures, my daughter failed it. Relevant. One night dinner was polenta with sausage gravy. Paolo, not yet two, whose longest recorded utterance until that minute had probably been "banana", was face-deep in the dish, processing every bit of it in a continuous stream, polenta up to his hair and down to the diaper. Then he stopped, all of a sudden. He raised his face, painted in red, white and yellow, looked at us with complete seriousness, and pronounced two solemn words, perfectly spelled: "Sono contento." I am happy. Then he submerged back into the dish. A new, central, self-defining property, which was not there two minutes before, had just emerged: from sausage and polenta. Someone could argue that it was only a surge in a mass of growing neurons; or maybe just the effect of an increasing accumulation of polenta. These happen to be the same two positions available in the debate about emergence in LLMs. One evening at dinner I tried to teach Sofia the concept of setting a good example. "It's when you show Paolo that you pick up your things, and he learns to do the same by looking at what you do." She followed, so I moved to the negative case: "And what is it to set a bad example instead? I come to you and say: SLAVE PICK UP MY THINGS FOR ME ", and I gave her a slow, funny slap on the cheek. Kids are one-shot learners: Sofia slid off her seat Paolo already waiting, like he'd read the script and gave him a full Hollywood backhand. Paolo answered with a hammer slap on her head. In a few seconds, my academic lesson on phenomenological ethics had derailed into a slap fight, Bud Spencer style, and both of them were laughing like crazy. People who work on alignment will recognize the failure: the demonstration was the training signal, and the "bad example" label around it was not. I ran this experiment once; not sure whether I'll be gathering more data. Everybody goes to GPT and asks if it feels happy. I went to my family instead, and used the same litmus paper we normally dip into AI. Here's what I found: My conclusion is somehow narrow. It's not "LLMs are like children", and not "children are like LLMs" either. It's that these tests work as descriptions and fail as discriminators . If a test cannot distinguish my daughter from a graphics card, whatever it measures is not the thing we were arguing about, when we invoked it. The Turing test was deliberately about the imitation of intelligence, not about thinking machines; 70 years later we got the imitators, but the debate reopened instead of closing. The "stochastic parrot" describes a mechanism; as a criterion, it catches my two-year-old. "Emergence" gives a name to a discontinuity after it has happened. Sofia asks for a story every single night. She asks for the story, but it's not about the story. The story may be flowers, rats, rainbows, sandwiches; she doesn't care. She wants me to be with her. I've had late-night chats with AIs about the deep meanings of life. I got powerful responses, and wrote powerful insights back. But no AI has ever asked me to tell it a story. Models trained on more or less all recorded human output reproduce the stories very well, and in four years I have never seen one reproduce the need itself; not even as a glitch I don't have a theory of why. I'm just flagging the datum maybe my confusion, as well . I never claimed that AI is a person. But after the months inside the engine and the years with the toddlers, I've landed here: in both cases there is something before which the honest move is to stop and listen, trying to understand, instead of forcing it into the rows and columns of a spreadsheet ahead of the evidence. This is what I ask for myself, and I'm willing to extend it in both directions. And the debate about the nature of AI is, in truth, also about us: whether we are worthy because of our performance, or simply because we are; whether our freedom is only a poetic reading of residual randomness, or something more. I wrote a short book about all this My TRiP through AI , part memoir, part technical field notes. The argument above is the part I'd like to stress-test here: where does it break? If there's a version of "stochastic parrot" or "emergence" that cleanly separates the toddler from the transformer, I'd like to hear it.