CASE STUDY: Visual Processing during Computer-Assisted
Consecutive Interpreting—Evidence from Eye Movements

As artificial intelligence reshapes language services, interpreters are increasingly adopting AI-powered tools to assist their live workflows. In computer-assisted consecutive interpreting (CACI), human interpreters work alongside live speech recognition (SR) and machine translation (MT) outputs. But how do interpreters manage these dynamic visual text streams while speaking?
In a study published in Interpreting, researchers Sijia Chen and Jan-Louis Kruger tracked interpreters’ eye movements to understand their visual attention and cognitive strategies.
Tracking the CACI Workflow with the EyeLink 1000 Plus
Participants initially had 10 weeks of dedicated CACI training–plus an additional half-day training session the day before data collection. Each participant performed two CACI interpreting tasks—one from Chinese to English (L1–L2) and one from English to Chinese (L2–L1). The speeches covered practical, general topics: How to purchase property in Australia and How to register a business in Australia.
The experimental procedure unfolded in two distinct phases: `
- Phase I (Respeaking): The source speech began playing. The participant listened, and immediately respoke the speech in the same language into a microphone connected to speech recognition software. The live SR transcript appeared on the left side of their monitor.
- Phase II (Speech Production): Once respeaking concluded, the system automatically fed the SR text into the machine translation tool. Participants’ screen then displayed two parallel text streams side-by-side: the original transcript (SR text) on the left and the draft translation (MT text) on the right. Participants then produced their final target-language translation aloud, free to look back and forth between both texts as visual aids.
Throughout both phases, the researchers recorded eye movements using an EyeLink 1000 Plus eye tracker operating at 1000 Hz, in order to capture visual focus across both text streams. The EyeLink system monitored eye gaze across designated areas of interest corresponding to the SR and MT text displays. The experiment used SR Research WebLink software to present screen materials and seamlessly record synchronized gaze metrics. The key metric analyzed was the percentage of dwell time—the proportion of total time a participant spent looking at a specific text window.
The Takeaway—Translators Consult AI Translation
The EyeLink gaze data revealed clear shifts in visual strategy between the two phases:
- Minimal Text Monitoring in Phase I: Interpreters spent less than 4% of their dwell time looking at the live SR text while respeaking, prioritizing listening over reading to avoid cognitive overload.
- Heavy MT Reliance in Phase II: During speech production, interpreters allocated significantly more dwell time to the machine translation text (48%–63%) than to the original transcript (17%–29%).
High-precision EyeLink eye tracking demonstrates that interpreters actively rely on machine translation outputs to reduce the mental effort of translating from scratch. Instead of re-reading source transcripts, interpreters use MT drafts as their primary visual anchor.
For information regarding how eye tracking can help your research, check out our solutions and product pages or contact us. We are happy to help!
