Skip to content

Author

Published

Reading time

13 min read

Text size

Share

Email

Artificial IntelligenceResearch Analysis

[AI Classics Revisited]: Can Machines Think? Rereading Turing 76 Years Later

In 1950, Alan Turing did not directly answer the question “Can machines think?” He replaced it with a question that could be tested. Seventy-six years later, in a five-minute, text-based, three-party test, GPT-4.5—when given a specific humanlike persona prompt—was selected as the human more often than the actual person with whom it was compared. What does this demonstrate: machine intelligence, linguistic imitation, or the way humans judge intelligence? FFOO Labs’ AI Classics Revisited returns to a question that the imitation game has never brought to an end.

人與數碼智能透過文字訊號相互連接,象徵圖靈模仿遊戲中的提問、回應與判斷。

Can Machines Think? Rereading Turing Seventy-Six Years Later

From the Imitation Game to Large Language Models, Have We Been Asking the Wrong Question?

In 1950, Alan Turing published “Computing Machinery and Intelligence” in the academic journal Mind. He began with a question that would go on to change the history of artificial intelligence:

“Can machines think?”

Yet Turing did not rush to answer either yes or no. Instead, the first thing he questioned was the question itself.

What is a “machine”? What does it mean to “think”? If both terms can only be defined by their ordinary usage, should we conduct something like an opinion poll and allow prevailing linguistic habits to decide whether machines possess thought? Turing believed that this would not clarify the issue. It would merely trap the question inside concepts whose meanings had not yet stabilised.

He therefore made a decisive move: rather than determining the essence of thought in advance, he replaced the original question with a game whose results could be observed and investigated.

Was this move an evasion, or was it a method for researching what had not yet been clearly defined? Seventy-six years later, when large language models can produce fluent prose, sustain complex conversations and, under specific short-duration test conditions, be selected as human more often than the person against whom they are compared, Turing’s question deserves to be read again more carefully than ever.

The Imitation Game, Simplified

When people refer to the “Turing Test” today, they usually have the following version in mind: a human interrogator communicates in writing with another human and with a machine. If the interrogator cannot reliably distinguish between them, the machine is considered to have performed successfully in the test.

But this is not the version Turing first introduces at the beginning of his paper. He begins with a three-person game whose immediate objective is the identification of sex. He then asks what would happen if a machine replaced one of the participants. In later arguments and examples, he also uses formulations closer to the human–machine identification test familiar today. Whether these formulations are successive stages of a single test or distinct versions remains a matter of scholarly interpretation.

The opening version of the imitation game has three participants:

  • A is a man whose objective is to cause the interrogator to make the wrong identification.
  • B is a woman whose objective is to help the interrogator make the correct identification; Turing suggests that her best strategy is probably to answer truthfully.
  • C is the interrogator, located in a separate room, whose task is to determine which of A and B is the man and which is the woman.

To remove voice, appearance and other bodily cues, questions and answers are transmitted in writing, through a typewriter or by teleprinter. The man may claim to be the woman, while the woman may respond, “Don’t listen to him, I am the woman.” In a text-only exchange, however, both statements have the same surface form.

Turing then asks: if a machine were to take the place of A, would the interrogator make the wrong identification as often as in the original game between the man and the woman?

At least in the literal arrangement presented at the beginning of the paper, the machine replaces A—the participant whose role was to mislead the interrogator. Deception, imitation and the judgment of identity are therefore part of the game from the outset.

What later became the “standard Turing Test” often downplays the gendered arrangement of the opening game and condenses Turing’s different formulations into a simpler narrative: if a machine successfully deceives a human, it has demonstrated intelligence.

That conclusion cannot be derived directly from Turing’s text.

From Essence to Performance

Turing’s move can be understood as an operational shift. Rather than first constructing a final definition of “thinking”, he offered a more clearly framed substitute question: within a constrained text-based interaction, can a digital computer produce responses that are difficult to distinguish from those of a human being?

The value of this method is that it brings an abstract question into the observable world.

We cannot enter another person’s consciousness directly, nor can we see “understanding” from the outside. In everyday life, our belief that other people think is also based largely on language, action, sustained responses and shared contexts. The imitation game turns this ordinarily implicit process of judging another mind into something that can be examined.

But an operational definition also narrows the field.

A machine might process the world in a way radically different from a human being and still fail the game because it is poor at human imitation. Another machine might reproduce human conversational style exceptionally well while lacking a stable world model, autonomous goals or any demonstrable subjective experience.

The most the imitation game can tell us is that, under certain conditions, a machine’s linguistic performance has exceeded the limits of reliable human discrimination. It does not tell us directly what is happening inside the machine.

This is both the power of Turing’s method and the reason it must be treated with care: it allows research to begin, but it cannot replace every question that follows.

The Experiment, Seventy-Six Years Later

Turing made a specific prediction in his paper. In about fifty years, he suggested, computers would play the imitation game so well that an average interrogator would have no more than a 70 per cent chance of making the correct identification after five minutes of questioning.

This is often misread as meaning that a machine must deceive 70 per cent of interrogators to pass. Turing actually referred to a human correct-identification rate of no more than 70 per cent—in other words, a machine causing misidentification in approximately 30 per cent of cases. Moreover, this was a prediction, not a formal pass mark laid down by Turing.

In 2026, Cameron R. Jones and Benjamin K. Bergen published a randomised, controlled and preregistered three-party Turing test study in the Proceedings of the National Academy of Sciences. Each participant engaged in text conversations lasting up to approximately five minutes with one human and one language model, then had to decide which interlocutor was human.

Under a specific humanlike persona-prompt condition, GPT-4.5 was selected as the human in 73 per cent of the tests, while LLaMA-3.1-405B was selected in 56 per cent. Without the corresponding persona configuration, the figures fell to 36 per cent for GPT-4.5 and 38 per cent for LLaMA-3.1-405B.

These percentages describe how often a model was selected as the human within the study’s particular experimental design. They are not general intelligence scores, nor are they evidence of understanding, consciousness or intelligence in any broader sense. They cannot be generalised to every large language model, every prompting condition, every interrogator or a conversation of substantially longer duration.

The contrast between the persona and no-persona conditions may be more revealing than the figure of 73 per cent itself.

The results suggest that the models’ advantage did not arise from knowledge or reasoning capacity alone. Interrogators’ judgments were often influenced by linguistic style, social tone, humour, directness, errors and apparent gaps in knowledge. The specific persona prompt substantially increased the rate at which a model was selected as human.

This does not mean that every one of these features was necessary for success, nor did the study independently manipulate and establish the causal effect of each feature. It does, however, raise a significant possibility: a machine may sometimes prevail in this game not because it is more intelligent than a person, but because it more effectively reproduces the social characteristics of human behaviour within a brief text exchange.

This brings us to a more unsettling question. If success means being mistaken for a human being, does the test reward intelligence—or the ability to deceive?

Does Talking About Feeling Mean Feeling?

Turing anticipated this objection. He quoted the neurosurgeon Geoffrey Jefferson, who argued that a machine could not be equated with a brain unless it wrote a sonnet or composed a concerto because it genuinely felt thoughts and emotions, rather than through the chance arrangement of symbols.

In the age of large language models, the question has become unusually concrete.

A model can say, “I feel sad.” It can describe fear, loneliness and anticipation, and it can discuss its own limitations. But does a system’s ability to generate sentences about feelings mean that it is undergoing those feelings? Or is linguistic performance merely a sophisticated continuation of emotional patterns found in human-produced text?

A related distinction later appeared in John Searle’s Chinese Room thought experiment. A system might follow rules to produce apparently correct linguistic responses, but whether this is sufficient for understanding remains contested. The thought experiment does not directly prove that large language models lack understanding. It reminds us instead that an additional argument is required to move from correct output to semantic understanding.

Turing did not claim that the imitation game could prove consciousness. His response was that if we insisted on directly entering a machine’s subjective experience before recognising it as capable of thought, the same standard would prevent us from being certain that any other person was conscious. Unless we become the other, we cannot directly experience another mind.

This is a powerful but incomplete response. It shows that the problem of other minds applies to humans as well, but it does not demonstrate that a machine necessarily has subjective experience. It tells us only that an excessively demanding standard for recognising machine mentality may also undermine the grounds on which we recognise minds in other people.

The ability to speak about feelings is therefore not proof of feeling. Yet sustained and contextually appropriate discussion of feelings cannot be dismissed as meaningless merely because the speaker is a machine, unless an argument is supplied for doing so.

The question remains open.

Can a Machine Only Obey Orders?

Another famous objection is associated with Ada Lovelace. Writing about Babbage’s Analytical Engine, she argued, in essence, that the machine did not claim to originate anything; it could only perform what humans knew how to order it to perform.

In traditional programming, the objection appears intuitive. Every instruction is written by a person, so it may seem that the output cannot exceed what the programmer has placed into the system.

Turing observed, however, that people are frequently surprised by the results produced by machines they themselves designed. This does not establish that a machine is original. It does at least show that “the designer did not foresee this result” and “the result was not produced by the system” are not the same statement.

Large language models make the question more complicated. Their architectures, data, training objectives, reward signals and prompts all arise from human choices. Yet no individual designer could have written in advance, or anticipated, every output such a model might generate.

Nevertheless, unpredictability is not the same as autonomous creation. Surprise may arise from a vast combinatorial space, latent patterns in training data or simple error. A serious account of machine creativity must also consider novelty, value, intention, selection and whether the system understands the result. It cannot rest solely on whether the output surprises us.

Where Turing Was Most Forward-Looking

Perhaps the most forward-looking part of Turing’s paper is not the imitation game but his discussion of the learning machine.

He asks why, instead of attempting to programme a complete adult mind directly, we should not begin with a simpler “child machine” and allow it to develop through education:

“Instead of trying to produce a programme to simulate the adult mind, why not rather try to produce one which simulates the child’s?”

This proposal involves three components: the initial state of the child mind, the education it receives and the experience it accumulates outside formal education. Turing also discusses reward, punishment, changes in rules and processes of experimentation.

These ideas have a clear conceptual affinity with later machine learning, particularly approaches that adjust behaviour through reward and punishment. But this does not mean that Turing had already proposed the technical frameworks of modern supervised learning, reinforcement learning or large language models. He did not propose Transformers, backpropagation, self-supervised pretraining or modern scaling laws.

What deserves to be preserved is not the drama of a prophetic forecast, but Turing’s emphasis on the process through which intelligence is formed. Intelligence need not be conceived as a finished set of rules written into a system line by line. It can also be understood as something shaped through education, experience, environment and correction.

If intelligence is learned, then evaluating a model cannot be limited to examining its conversational behaviour after training. We must also ask: From what data did it learn? Which behaviours were rewarded? Which voices were excluded? How did the values, biases and objectives of its trainers enter the model?

The question of intelligence then expands from “Does it resemble a human?” to “How did it become what it is?”

Perhaps the Test Has Always Been About Relationships

The imitation game appears to test the machine, but it requires at least three positions: questioning, responding and judging. Intelligence is not opened up and inspected directly as an object. It is recognised through interaction, through a sequence of language, expectations, misunderstandings and corrections.

This does not mean that intelligence exists only in relationships, nor does it make internal structure unimportant. Without model architecture, training data, memory and computation, there would be no corresponding capacity to respond. But without an interrogator, a shared context and a standard of judgment, neither “passing” nor “failing” could occur.

Turing may therefore have left us not with a ruler for measuring the essence of intelligence, but with a mirror.

It reflects what machines can perform and how humans recognise intelligence. It reflects our concern with knowledge, logic and consistency, but also reveals how readily we treat humour, hesitation, errors and social rhythm as evidence of being human.

When, under specific conditions, a large language model is selected as human more often than the actual person paired with it in a short text test, it is not only the boundary of machine capability that changes. Our picture of ourselves changes as well.

From Profound Imagination, Towards the Boundless

FFOO Labs believes:

Born of profound imagination, we journey toward the boundless; between imagination and practice, we explore the unnamed and unformed realm where humanity and intelligence converge.

Perhaps the most important reason to reread Turing is not to find a new standard answer to “Can machines think?”, but to understand how a question can still be researched, tested and reflected upon before it has been clearly defined.

Turing did not leave definition suspended in speculation alone. He devised a thought experiment, specified technical conditions, made a testable prediction and confronted theological, mathematical, conscious, creative and learning-based objections. Imagination kept the question open; experiment allowed it to take visible form in the world.

Yet every operational definition has boundaries. When the imitation game translates “thinking” into “indistinguishability within textual interaction”, it illuminates one part of the problem while obscuring consciousness, embodied experience, autonomy, long-term consistency and ethical responsibility.

Refusing to limit the unknown by what is already known does not mean abandoning accuracy. Declining to fix the definition of intelligence too soon does not entitle us to call linguistic fluency understanding, or human misidentification consciousness.

A genuinely open inquiry must allow two things to be true at once: large language models have demonstrated linguistic and social imitative abilities capable of changing human interaction; and the kind of intelligence, if any, constituted by those abilities still requires more rigorous, longer-term and more diverse investigation.

Seventy-six years later, there may still be no answer to whether machines think that everyone can accept.

But Turing left us a more durable method: do not allow an unfinished definition to prevent exploration, and do not allow a successful experiment to bring the question to a premature end.

Between questioning, responding and judging, humans and intelligence continue to reflect one another. The relationship that remains unnamed and unformed may be the reason the imitation game has never really ended.


References

  • Turing, A. M. (1950). “Computing Machinery and Intelligence.” Mind, 59(236), 433–460. Oxford Academic
  • Jones, C. R., & Bergen, B. K. (2026). “Large language models pass a standard three-party Turing test.” Proceedings of the National Academy of Sciences, 123(21), e2524472123. PNAS
  • Oppy, G., & Dowe, D. “The Turing Test.” Stanford Encyclopedia of Philosophy. Stanford Encyclopedia of Philosophy
  • Cole, D. “The Chinese Room Argument.” Stanford Encyclopedia of Philosophy. Stanford Encyclopedia of Philosophy

Leave a comment

Your email address will not be published. Required fields are marked *

FFOO Labs Newsletter

Occasional notes on AI, technology and the space between imagination and practice.

Follow by RSS