The Turing test can’t determine that a human is intelligent if he only speaks French,and you don’t. In general, if the human is sentient but can’t communicate well, it’s going to have a hard time. This does not mean that any AI is as intelligent as a human who doesn’t communicate well.
There’s also the fact that the Turing test works because humans can modify their questions based on the entity’s responses, to try to poke at specific aspects of their intelligence that couldn’t be predicted in advance, based on the human’s own knowledge of how humans behave (and in this case specifically toddlers). It doesn’t sound like you did this in your tests.
(I’m saying this, of course, about questions designed to figure out if the entity is intelligent, not shibboleths like seeing how many times they refer to toys and pizza.)
Thank you, Jiro. You are right about the asymmetry; I should have made it explicit: the Turing test was proposed as (roughly) a sufficient condition, so: a toddler failing it doesn’t break the test. Conceded!
What I’m pushing on is how the test gets used in the debate: it is invoked as if it sorted minds from “non-minds”, and it can not do that even in principle: passing is (maybe) informative, failing tells you nothing. My daughter and your French speaker both land in the “no information” bucket, while the graphics card lands in “yes” (!). That may be the “instrument” working as designed, but then it’s not doing the job which the public debate hires it for. Turing himself was careful about this: imitation, not thinking.
On “adaptive questioning”: fair, again. My toddler experiments were: not blinded, and not adaptive, and heavily contaminated by my affection for the subjects. The serious half of the claim is about “stochastic parrot” and “emergence” used as criteria, where no adaptive protocol exists at all!
About “shibboleths”: I agree. Counting the “pizza” references gets my daughter identified with 99.99% accuracy; but I suspect there could be some overfitting here ;^)
If I tell someone to check for fools gold by rubbing it on a plate, nobody’s going to say “That test doesn’t work! What if it’s locked up inside a box so you can’t rub it on a plate?”
That’s just not what people mean when they say that a test works or doesn’t work. “This test works” means “this test works when the test is applicable”. The Turing test relies on the ability to communicate, and any statement about the Turing test working to detect minds has the implicit limitation “as long as you can communicate with the mind”.
The reason that the toddler doesn’t break the test isn’t that the test is a sufficient condition, it’s that that isn’t what “fails the test” means. “Fails the test” means “fails, when used on something that will communicate”.
For gold: the condition and the object are separate things: the streak is evidence for composition, which we can check by other means, and you can see whether you managed to rub the sample.
The Turing test has no access to a mind except through communicative performance. There is nothing, behind the conversation, that the test reaches; the conversation is the whole measurement.
So, the applicability clause does a different work in the two cases. “Works when you can rub it” is checkable. “Works when the subject can communicate” is not: to tell “can’t communicate” from “no mind there” you need independent knowledge of the subject, and that knowledge is what the test was supposed to give us. “I know my daughter’s failure means nothing” only because “I already know she has a mind” ; the test borrows its answer from what I knew before running it.
That’s just a quirk of the example. Imagine using a metal detector to find metal. Nobody would say that a metal detector “doesn’t actually detect metal” on the grounds that it can’t detect every single piece of metal, even ones that are too small. Yet the only way to know the difference between “too small” and “not made of metal” is to do something (like dig it up) that amounts to independently detecting it.
If not being able to detect uncommunicative minds made the Turing test a failure in a meaningful way, people wouldn’t have been talking about the test for decades.
The Turing test can’t determine that a human is intelligent if he only speaks French,and you don’t. In general, if the human is sentient but can’t communicate well, it’s going to have a hard time. This does not mean that any AI is as intelligent as a human who doesn’t communicate well.
There’s also the fact that the Turing test works because humans can modify their questions based on the entity’s responses, to try to poke at specific aspects of their intelligence that couldn’t be predicted in advance, based on the human’s own knowledge of how humans behave (and in this case specifically toddlers). It doesn’t sound like you did this in your tests.
(I’m saying this, of course, about questions designed to figure out if the entity is intelligent, not shibboleths like seeing how many times they refer to toys and pizza.)
Thank you, Jiro. You are right about the asymmetry; I should have made it explicit: the Turing test was proposed as (roughly) a sufficient condition, so: a toddler failing it doesn’t break the test. Conceded!
What I’m pushing on is how the test gets used in the debate: it is invoked as if it sorted minds from “non-minds”, and it can not do that even in principle: passing is (maybe) informative, failing tells you nothing. My daughter and your French speaker both land in the “no information” bucket, while the graphics card lands in “yes” (!). That may be the “instrument” working as designed, but then it’s not doing the job which the public debate hires it for. Turing himself was careful about this: imitation, not thinking.
On “adaptive questioning”: fair, again. My toddler experiments were: not blinded, and not adaptive, and heavily contaminated by my affection for the subjects. The serious half of the claim is about “stochastic parrot” and “emergence” used as criteria, where no adaptive protocol exists at all!
About “shibboleths”: I agree. Counting the “pizza” references gets my daughter identified with 99.99% accuracy; but I suspect there could be some overfitting here ;^)
If I tell someone to check for fools gold by rubbing it on a plate, nobody’s going to say “That test doesn’t work! What if it’s locked up inside a box so you can’t rub it on a plate?”
That’s just not what people mean when they say that a test works or doesn’t work. “This test works” means “this test works when the test is applicable”. The Turing test relies on the ability to communicate, and any statement about the Turing test working to detect minds has the implicit limitation “as long as you can communicate with the mind”.
The reason that the toddler doesn’t break the test isn’t that the test is a sufficient condition, it’s that that isn’t what “fails the test” means. “Fails the test” means “fails, when used on something that will communicate”.
For gold: the condition and the object are separate things: the streak is evidence for composition, which we can check by other means, and you can see whether you managed to rub the sample. The Turing test has no access to a mind except through communicative performance. There is nothing, behind the conversation, that the test reaches; the conversation is the whole measurement.
So, the applicability clause does a different work in the two cases. “Works when you can rub it” is checkable. “Works when the subject can communicate” is not: to tell “can’t communicate” from “no mind there” you need independent knowledge of the subject, and that knowledge is what the test was supposed to give us. “I know my daughter’s failure means nothing” only because “I already know she has a mind” ; the test borrows its answer from what I knew before running it.
That’s just a quirk of the example. Imagine using a metal detector to find metal. Nobody would say that a metal detector “doesn’t actually detect metal” on the grounds that it can’t detect every single piece of metal, even ones that are too small. Yet the only way to know the difference between “too small” and “not made of metal” is to do something (like dig it up) that amounts to independently detecting it.
If not being able to detect uncommunicative minds made the Turing test a failure in a meaningful way, people wouldn’t have been talking about the test for decades.