It already happens now, without science fiction and without holograms. A chat opens, someone replies with a joke, gets a detail wrong, uses a slightly crooked phrase, maybe puts that lightness of a real person that makes you lower your guard. For years we thought we recognized artificial intelligence by its perfection, by its overly smooth responses, by its mechanical courtesy, by that sort of automatic reception smile. The trouble is that the most advanced models are learning something else: the small social imperfection.
A new study published in Proceedings of the National Academy of Sciences and conducted by researchers from the University of California San Diego, reports the Turing test in a form very close to the idea formulated by Alan Turing in 1950: a person converses simultaneously with two interlocutors, one human and one artificial, then must understand which of the two is the real person. Nearly 500 participants were involved in the tests, including university students and a larger online sample, with text conversations lasting five or fifteen minutes.
The result has a certain effect. GPT-4.5 was judged to be humane in 73% of casesthus more often than the actual person it was being compared to. LLaMa-3.1-405B reached 56%, a value statistically indistinguishable from human interlocutors. The systems used for comparison remained much further behind: ELIZA, the historic chatbot from the 1960s, at 23%; GPT-4o at 21%.
The strange part is in the tone
The most uncomfortable thing about the study concerns the reason for the result. The most convincing models worked best when they received a “person” prompt, that is, precise instructions to assume a character, a way of speaking, a conversational posture. Without that mask, GPT-4.5 dropped from 73% to 36%, while LLaMa-3.1 went from 56% to 38%.
Here the discussion shifts. What deceived the participants was not pure intelligence, understood as the ability to solve problems or reel off information. It was there social similarity: tone, irony, hesitations, naturalness, fallibility. Cameron Jones, author of the study, explains that with the right prompts, great language models can display human-like tone, immediacy, humor and imperfections. Ben Bergen, co-author of the research, adds that the Turing test today increasingly measures “perceived humanity”, rather than the brute strength of reasoning.
And it is precisely here that the matter becomes more everyday. AI doesn’t have to look like a genius to pass as human. She just needs to look normal enough. An answer that is too perfect can raise suspicion; a somewhat lateral response, with a half-successful joke, with an ordinary chat expression, can have the opposite effect. In practice, the machine does not win when the computer plays better. Wins when any person plays.
A five minute chat is enough
The detail of the times weighs. The conversations lasted five minutes, or fifteen in the replay. We are not talking about endless interrogations, laboratory tests far from real life. Let’s talk about the duration of a normal online exchange: a message on a forum, a conversation on a social network, a request for information, a profile that comments under a post, someone who writes to you with a trustworthy air.
Jones puts it quite bluntly: it’s relatively easy to tell these models how to become indistinguishable from humans, and when talking to strangers online we should be much less sure that we’re talking to a person. Bergen brings the reasoning to more practical ground: those who want to use bots to convince someone to share personal data, support a party or buy a product find this ability to be a very powerful tool.
For Italy, the reference to personal data immediately translates into scenes already seen: suspicious links, fake operators, messages asking for codes, credentials, OTPs, banking access, documents, digital identities. The difference is that until now many scams were betrayed by rigidity, gross errors and poorly translated formulas. A model capable of modulating tone, patience, confidence and small imperfections makes that threshold much more slippery.
This doesn’t mean that every online profile is a bot, nor that every chatbot is a threat. The study says something more precise and more useful: our confidence in recognizing the human from conversation is becoming fragile. For decades we have used style as implicit proof of authenticity. If someone joked well, made mistakes well, hesitated well, they seemed like a person. Now that evidence holds up much less.
The Turing test changes face
The Turing Test began as a question about machine intelligence. Today he returns with a different question, dirtier and closer to our habits: how much is enough to seem human in a chat? The study’s response is not very reassuring. Sometimes all you need is a good personality.
The distinction remains fundamental: appearing human does not mean experiencing emotions, having conscience, desires, intentions, real biographical memory. It means producing a conversational form that we interpret as presence. And the human being, faced with a credible presence, tends to complete the rest on his own.
Perhaps the most useful lesson lies here. There is no need to imagine sentient machines replacing us wholesale. We need to look more carefully at the tiny normality of online conversations. The “hello” written well. The ironic answer. The fake embarrassment. The phrase that seems to have come from a tired person in front of the screen. The next great imitation could look the most banal in the world. And that’s the hard part to see.