CASE STUDY: Wave 2 of the Multilingual Eye-Movement Corpus (MECO): New Text Reading Data Across Languages

Reading research historically relies heavily on English speakers, creating a massive Anglocentric bias in cognitive science. To build a truly universal picture of the reading brain, the global Multilingual Eye-Movement Corpus (MECO) project is tackling this imbalance across three major phases. With Wave 1 already complete and Wave 3 on the horizon, Noam Siegelman, Sascha Schroeder, Victor Kuperman, and their colleagues have officially released MECO Wave 2. Published in Scientific Data, with the information on the MECO Official Website, this pivotal second phase captures real-time first-language reading behaviors from 654 participants across 13 native languages.
Standardized Experimental Methodology with the EyeLink Eye Trackers
To ensure direct cross-linguistic comparability, researchers used a uniform multi-step procedure across 16 labs in 15 countries:
- Language Profiling: Participants first filled out an abridged Language Experience and Proficiency Questionnaire (LEAP-Q) to document their demographic background and self-rated native language proficiency.
- Passage Reading Task: Next, participants silently read 12 encyclopedic, Wikipedia-style passages in their native language. To keep topic differences from throwing off the results, five texts were directly translated to share the exact same meaning in every language. The other seven texts shared the same topics, but were written naturally in each language.
- Comprehension & Cognitive Assessment: Immediately following each passage, participants answered four yes/no comprehension questions that served as attention checks. Finally, participants completed a non-verbal intelligence test (CFT-20 Matrices) alongside individual native-language reading skill assessments.
Achieving data quality across diverse global research environments required rigorous technical harmonization. Every participating laboratory standardized its data collection on EyeLink eye trackers (including EyeLink 1000, EyeLink 1000+, EyeLink Portable Duo, and EyeLink II systems). Participants rested on chin and forehead supports to minimize head movement while eye movements were tracked monocularly from the dominant eye at a sampling rate of 1000 Hz. Researchers programmed the experimental task using SR Research Experiment Builder software to guarantee millisecond-accurate timing uniform stimulus presentation.
Expansion of Linguistic Diversity and Geographic Reach
Wave 2 represents a massive leap forward in bridging the gap toward under-represented language communities:
- Unprecedented Coverage: Nine of the thirteen languages analyzed in this wave represent less than 1% of the languages historically studied in eye-movement research.
- New Writing Systems: The expansion features seven completely new languages and eight new written languages, bringing non-Latin scripts like Devanagari (Hindi) and logographic characters (Mandarin Chinese) into the global dataset
Differences in Oculomotor Patterns Across Languages
1. Character Density and Visual Information Unit
The physical area required to convey meaning varies drastically between scripts, directly influencing how far ahead a reader’s eyes jump (saccade) and how often words are skipped.
- Logographic Scripts (Mandarin Chinese): Chinese packs dense semantic information into compact, individual characters. Because a word often consists of only one or two visually dense characters, Chinese readers demonstrate significantly higher skipping rates and higher regression rates (gazing back at previous text) compared to alphabetic readers.
- Alphabetic Scripts (Latin, Cyrillic): Spelled-out words spread visual information across multiple letters, resulting in lower skipping rates and longer average saccade lengths measured in characters.
2. Morphological Complexity and Typology
How languages build words out of smaller units of meaning (morphemes) places distinct processing demands on the reader’s visual system.
- Agglutinative Languages (Basque, Turkish): Agglutinative languages build complex words by stringing together multiple prefixes and suffixes. Because individual words carry heavy grammatical loads and tend to be longer, readers of Basque evoke longer overall total fixation durations and a higher number of fixations per word.
- Synthetic / Fusional Languages (Hindi, German, Russian): Morphologically rich languages require extra cognitive processing to parse inflections and case markings. This structural complexity leads to increased refixation rates (looking at a word more than once during the first pass) and longer gaze durations.
3. Script Layout and Visual Complexity
Beyond word structure, the physical appearance and layout of a writing system impact initial visual uptake during reading.
- Abugida Writing Systems (Hindi / Devanagari): The Devanagari script combines consonants and vowels into single visual units (syllabic blocks) and uses a continuous horizontal hanging line (Shirorekha). This visually complex layout results in longer first fixation durations and more fixations per word on average compared to transparent alphabetic scripts.
- Orthographic Transparency: In highly transparent orthographies (like Serbian or Turkish), where letter-to-sound mappings are straightforward, readers process words faster than in opaque orthographies (like English or Danish), reflecting in shorter first-fixation durations.
Key Oculomotor Metrics at a Glance [Sam: Leave out?]
| MEASURE | LOGOGRAPHIC (CHINESE) | AGGLUTINATIVE, ABUGIDA (BASQUE, HINDI) | ALPHABETIC (e.g., ENGLISH, SPANISH) |
|---|---|---|---|
| Skipping Rate | High | Low to Moderate | Moderate |
| Fixation Duration | Short per character | Long overall | Moderate |
| Fixations Per Word | Low | High | Moderate |
| Regression Rate | High | Moderate | Low to Moderate |
Powered by reliable EyeLink hardware and standardized testing protocols, MECO Wave 2 provides an essential benchmark for reading fluency, word skipping, and fixations worldwide. As this second wave bridges the gap between the initial findings and the upcoming third installment, the project continues to push cognitive science toward truly universal models of human language processing.
For information regarding how eye tracking can help your research, check out our solutions and product pages or contact us. We are happy to help!
