Carlo Valenti is a firmware engineer and a father of two. To understand transformers, he built TRiP, a transformer inference engine written from scratch in C (github.com/carlovalenti/TRiP). He also holds a degree in Theology, which turns out to be handy when a machine starts raising questions about freedom and self-awareness. His first book, “My TRiP through AI”, is part memoir, part essay, part technical field notes. He lives in Italy.
Carlo Valenti
Thank you, Jiro. You are right about the asymmetry; I should have made it explicit: the Turing test was proposed as (roughly) a sufficient condition, so: a toddler failing it doesn’t break the test. Conceded!
What I’m pushing on is how the test gets used in the debate: it is invoked as if it sorted minds from “non-minds”, and it can not do that even in principle: passing is (maybe) informative, failing tells you nothing. My daughter and your French speaker both land in the “no information” bucket, while the graphics card lands in “yes” (!). That may be the “instrument” working as designed, but then it’s not doing the job which the public debate hires it for. Turing himself was careful about this: imitation, not thinking.
On “adaptive questioning”: fair, again. My toddler experiments were: not blinded, and not adaptive, and heavily contaminated by my affection for the subjects. The serious half of the claim is about “stochastic parrot” and “emergence” used as criteria, where no adaptive protocol exists at all!
About “shibboleths”: I agree. Counting the “pizza” references gets my daughter identified with 99.99% accuracy; but I suspect there could be some overfitting here ;^)
I ran the standard AI litmus tests on my two toddlers (yep)
Hello everyone! I’m Carlo Valenti, a firmware engineer from Italy; to understand how LLMs work I spent 18 months building a transformer inference+training engine from scratch in C (“TRiP”, on GitHub). Along the way, I kept comparing what I saw in models with what I saw in my two toddlers, and I ended up writing a short book about it. I’m planning a longer post here, about what building an engine (and raising the two!) from scratch taught me; happy to answer questions in the meantime.
For gold: the condition and the object are separate things: the streak is evidence for composition, which we can check by other means, and you can see whether you managed to rub the sample. The Turing test has no access to a mind except through communicative performance. There is nothing, behind the conversation, that the test reaches; the conversation is the whole measurement.
So, the applicability clause does a different work in the two cases. “Works when you can rub it” is checkable. “Works when the subject can communicate” is not: to tell “can’t communicate” from “no mind there” you need independent knowledge of the subject, and that knowledge is what the test was supposed to give us. “I know my daughter’s failure means nothing” only because “I already know she has a mind” ; the test borrows its answer from what I knew before running it.