CASE STUDY: Turing Tests in Chess: An Experiment Revealing the Role of Human Subjectivity

The Turing test, originally proposed by Alan Turing in 1950, determines whether a machine can mimic human intelligence so convincingly that a human evaluator cannot reliably tell it apart from a person. In a study published in Computers in Human Behavior Reports, researchers Yke Bauke Eisma, Robin Koerts, and Joost de Winter tested whether chess players could determine if their opponent was a human or a computer.
The Methodology and Procedure
The experiment involved 24 chess players who each played eight 5-minute Blitz games starting from balanced, pre-selected positions. Each participant faced four distinct opponents twice:
- A Human Player: An experimenter rated around 1100 Elo (above-average player).
- Maia: A neural-network engine specifically trained on millions of human games to replicate human-like moves and mistakes.
- Stockfish (Low Level): A downgraded version of Stockfish 16 that occasionally introduces artificial errors.
- Stockfish (Max Level): The world’s strongest chess engine playing at a superhuman level.
To standardize the experience, opponents executed moves at a fixed interval of 10 seconds. Participants spoke their thoughts aloud during gameplay, analyzing positions and reflecting on their opponent’s play style. After every game, players completed a questionnaire evaluating whether they believed their opponent was human or machine.
Visual Tracking with EyeLink and WebLink and Results
To observe visual attention and cognitive effort during web-based gameplay, the researchers recorded eye movements and pupil size using an EyeLink Portable Duo tracker (SR Research), sampling at 1000 Hz.
The experiment was run using SR Research WebLink, a screen-recording software that seamlessly captured synchronized gaze data, screen activity, browser navigation, and mouse movements during live chess matches. For example, pupillometry data captured by the EyeLink system revealed a sharp surge in pupil dilation immediately after an opponent executed a move, reflecting the instant jump in cognitive processing load as players evaluated new board positions.
Overall, the study highlighted the major role of human perception in evaluating AI:
- Superhuman Performance Fails: Max-level Stockfish was correctly identified as an engine 75% of the time because its play was too flawless.
- Human-Like Errors Deceive: Players frequently mistook Maia for a human opponent because it made natural, human-like mistakes.
The Takeaway—The Turing Test is About Mimicking Human Limitations
High-precision EyeLink hardware and WebLink software helped capture subtle shifts in cognitive burden during real-time human-AI interaction. The study proves that in a domain where computers perform at superhuman levels, passing the Turing test is no longer a measure of raw machine intelligence; rather, it is a test of whether an AI can adapt itself to convincingly replicate human flaws and match our expectations of human behavior.
For information regarding how eye tracking can help your research, check out our solutions and product pages or contact us. We are happy to help!
