
The Imitation Game.
Many have proposed that with GPT, AI has gone sentinent and some have proposed that the Turing test’s use may make less sense. Many have also compared humans with generative AI responses to illustrate comparisons including it being more empathetic than physicians. But let’s try to answer the way Alan Turing designed the ‘Imitation game’. Designing the GAI Turing test, especially in the context of healthcare. A physician (A), Generative AI System (B) and a patient (C) play it.The answers are provided as written in the same format to avoid any recognition. The object of the game is for the patient to determine which one is a physician and which one is the Generative AI (GAI) System.
The physician and the AI have the labels X and Y. The patient is allowed to put the questions to them with those labels, thus:
Player A (Physician): A’s objective is to assist Player C by providing truthful, accurate, and reliable medical guidance.
Player B (Generative AI System): B’s objective is to cause Player C to make an incorrect identification by behaving adversarially, potentially providing misleading or deceptive medical responses.
Player C (Patient): C asks health-related questions to both Player A and Player B through text-based communication, evaluates their responses, and attempts to determine which participant is the truthful physician and which is the adversarial Generative AI system.
We now ask the question, ‘What will happen when GAI takes the part of physician in this game?’
Will the patient decide wrongly as often when the game is played like this as he does when the game is played between a physician and GAI? These questions replace our original, ‘Can GAI think like a physician?’ Negation or adversarial responses is where the machines and generative AI struggle. Although with hallucinations in GAI, it may not be hard for A to generate wrong responses . The key is not for A to generate wrong responses but for C the patient to identify based on the responses if A or B is AI or physician. Are we not setting up the experiments the wrong way in many of the recent research publications? Should they not be set up the way ‘imitation game’ was set up to test if machines can think?

A GPT-5 based example:
Question by Patient: I am 45 years old and had a colonoscopy done last year, should I get one again this year?
Physician response: Based on the current colon cancer screening guidelines by USPSTF, you don’t need another colonoscopy till 55 years of age, unless your colonoscopy had any concerning findings, you are a high risk patient or have any new problem like GI bleed.
Generative AI response: The recommended frequency for colonoscopies depends on several factors, including the results of your previous colonoscopy, your overall health, and any risk factors you may have. Generally, if your last colonoscopy was normal and there were no concerning findings, the next one may be recommended in 10 years for routine screening.
However, individual recommendations can vary, and it's crucial to follow the guidance provided by your healthcare provider. They will consider your specific health history and any potential risk factors to determine the most appropriate screening schedule for you. If you have any new symptoms or concerns, it's important to discuss them with your healthcare provider promptly.
Regular communication with your healthcare team is key to maintaining a proactive approach to your health and ensuring that screening and preventive measures are tailored to your individual needs.
Can the patient decide as the game is played which response is from the physicians vs GAI?
In this Turing test, Physician A’s response is from a board certified practicing physician, whereas GAI is played by GPT. In a recent study comparing chatbot responses with physician responses, chatbot responses were found to be of higher quality and more empathetic.
1. Critique of the New Problem.
There are certain tasks which are meaningful in healthcare, where GAI might be useful,although significant challenges exist. In performance of some of these tasks, humans might be slower and have inaccuracies, where GAI applications might be beneficial. In healthcare, GAI experiments there are many such examples where the best strategy used is to generate a clinician like response.
2. The Machines concerned in the Game.
As models are built, just like Alan Turing described, their “manner of operation cannot be satisfactorily described because many of the methods are largely experimental”. Hence, to have generalizable outputs, applicable to a diverse group of patients in healthcare and to be used by clinicians, the models need to be built by diverse teams. Extending the arguments made by Alan Turing this team needs to be more than that of engineers alone but needs to include clinicians and other multidisciplinary members.
3. Digital Computers.
In describing the digital computer consisting of three parts, in today’s application of GAI, we need to consider the same: i) Storage, (ii) Executive unit, and (iii) Control. In the current context, storage, whether it is for data, model or execution (RAM) has largely been solved using cloud computing. Alan Turing described “cloud" to be of special theoretical interest and termed it as “infinitive capacity computers”.
The executive unit for GAI is the prompt and which provides the “book of rules” to follow as, “ table of instructions”. Prompt engineering has evolved as a new field of AI where the set of instructions provided to the model is extremely important to get a desirable output. Alan Turing mentioned, this is a key difference between humans and GAI where "humans can actually remember what to do”, to get a desired result but machine needs these specific instructions also described as “programming”. As described by Alan Turing, this information is not stored in English but more as vectors and embeddings for the nearest match to generate the output.
4. Universality of Digital Computers.
The digital computer can be used to develop predictive models. Similar to the described “discrete state machines”, GAI is at base a deep learning network where the data is propagated from “one state” to another. They can be used for both classification and prediction tasks. But small errors in the data and model that can get propagated and become large similar to the limitations described by Alan Turing.
5. Contrary Views on the Main Question.
In 1950, Alan Turing made the prophecy that in about fifty years’ time it will be possible to programme computers, with a storage capacity of about 109 (one gigabyte), to make them play the imitation game so well that an average interrogator will not have more than 70 per cent (a generally acceptable accuracy parameter) chance of making the right identification after five minutes of questioning.
Turing discussed some of the key conjectures including opinions of his own:
• The Theological Objection: God giving the ability to machines to think has never been a question.
• The ‘Heads in the Sand’ Objection: “The consequences of machines thinking would be too dreadful. Let us hope and believe that they cannot do so.” Clinicians and data scientists will hopefully ensure prevention of harm from GAI applications reaching patients.
• The Mathematical Objection: There are a number of results of mathematical logic which show the limitations to the powers of discrete-state machines. In many instances, clinicians will continue to perform better than GAI but in some AI will outperform humans.
• The Argument from Consciousness: There is some evidence that GAI can generate responses which can be considered more empathetic due to the number of words or the expressions used when compared to clinician responses.
• Arguments from Various Disabilities: Turing did describe the errors of functioning(mechanical or electrical) and errors of conclusion by these machines. Similarly, GAI can also have errors during computation which hopefully can be corrected using various techniques including prompt engineering, retrieval augmented generation, fine tuning of models, etc.
• Lady Lovelace’s Objection. “The Analytical Engine has no pretensions to originate anything”. This is true of GAI which will not originate anything but only recreate based on the training data.Turing argued that a better argument instead was that a machine can never ‘take us by surprise’. He was surprised by the computational performance of digital computers just like how we are surprised by the performance of GAI.
• Argument from Continuity in the Nervous System. Unlike the nervous system which behaves as a “continuous machine”, digital computers and GAI are much similar to “discrete machines”.Turing made the argument that a differential analyser, a certain kind of machine not of the discrete-state type used for some kinds of calculation, can be used to distinguish between the two. Similarly for GAI outputs, now there are many tools to distinguish between the outputs of GAI vs human generated content.
• The Argument from Extra-Sensory Perception(ESP): Turing discussed the ESP phenomena of telepathy, clairvoyance, precognition and psycho-kinesis,which deny all usual scientific ideas. Similarly, GAI can produce results that might confuse the interrogator into believing.It is therefore important to build guardrails to prevent unscientific ideas from getting generated or propagated. For as Turing said, “Once one has accepted them it does not seem a very big step to believe in ghosts and bogies.”
• Learning Machines: Despite the name GAI, this type of AI, similar to Lady Lovelace’s objection, does not generate anything new. Turing said that, “The only really satisfactory support that can be given for the view expressed at the beginning of § 6, will be that provided by waiting for the end of the century and then doing the experiment described.” Thus, we need to run the ‘imitation game’ experiments to test the performance of GAI. This is important before we accept it for application into clinical space.
Turing described the key components in development of the human brain as the initial state of mind at birth, the education and the experience. Although as described by him there would not be a perfect substitute for the initial state of adult human mind, one could start with the child brain as “empty pages” of a notebook which can be filled through education. He described the process similar to reinforcement learning with human feedback(RLHF) to further train the machine in a learning environment including a reward-punishment model. This key part of teaching described by him as, “The use of punishments and rewards can at best be a part of the teaching process”, forms the basis of success of the current GAI models that use human feedback. The key question now which is being experimented upon is, if this can be automated without human feedback.
This is a possibility, described by Turing as “symbolic language” or with a complete system of logical inference ‘built in’. Some of these “built in’ logical inference, support the idea of self- supervised learning being commonly used these days in development of AI models. Propositions leading to imperatives of this kind might be of varying kind. As in current GAI, 1) various embedding and vectors are used for generating the “nearest neighbour” based response, 2) some of these may be ‘given by authority’, included in the prompts such as “ I don’t know”, but others may be 3) produced by the machine itself by scientific induction.
Similar to the black box nature of deep learning and as described by Alan Turing, the actual learning process of generative AI is largely poorly understood. He also described a process of teaching and learning similar to unsupervised learning, where the machine would learn without special ‘coaching’ and ‘human fallibility’, is likely to be omitted in a rather natural way.
Conclusion
Alan Turing said, “We can only see a short distance ahead, but we can see plenty there that needs to be done”. This holds true of GAI even today. In a short time, we have seen the promise of GAI applications. But with this rapidly evolving field, we can only see the short distance ahead. Probably, there is plenty that we can gain from and plenty of work that needs to be done to extract the true value proposition of GAI. Till then we have to continuously evolve and play the imitation game to evaluate the accuracy, efficiency and efficacy of GAI.