A newly published study of Ethiopia’s primary school English textbooks has delivered a sobering verdict: most of the reading passages that children in grades 3 through 6 are expected to master are far too difficult for them, in many cases matching texts written for learners several years older. The research, conducted by Abebe Wubalem of Bahir Dar University and published in Discover Education, combined computational text analysis, cloze tests with nearly 300 students, and interviews with teachers to build one of the most comprehensive readability assessments yet attempted in an East African EFL context. Its conclusions challenge the assumption that difficult texts simply provide ‘positive challenges’ that push children to work harder.
The textbooks under scrutiny were written in 2020 and officially launched by Ethiopia’s Ministry of Education in 2022 for use across all primary grades. Since their introduction, parents and teachers have been divided. Some complained that the reading sections left children frustrated; others insisted the texts were appropriate or appropriately demanding. Wubalem set out to resolve this dispute empirically, asking whether the texts sit at the right readability level for their target age groups, which books are most affected, and what specific features of the texts drive the difficulty.
The methodological design was deliberately multi-pronged. Every reading passage in the four textbooks was assessed with Brown’s recalculated Flesch readability formula, an adaptation widely used in foreign language settings such as South Korea and Japan because it sets age-appropriate difficulty thresholds for EFL learners. A systematically drawn 35 percent sample of passages was then converted into cloze tests, in which every seventh content word was deleted, and administered to 280 learners, 70 from each grade. To separate genuine text difficulty from student weakness, the same students also completed standard cloze tests calibrated to A1/A2 levels of the Common European Framework of Reference. Finally, 64 teachers were identified through multi-stage sampling across the Amhara region, with 20 interviewed until thematic saturation was reached.
The Flesch-based results were striking. Grade 3 readers, ideally scored at a readability level of 98 for children aged roughly 8.5 to 10, produced scores ranging from 54.7 to 88, with several passages appropriate only for fifth, seventh, or even ninth graders. The grade 4 textbook fared comparatively well, averaging 84.6 against a benchmark of 95, and three of its passages met the minimum requirement. Grades 5 and 6 fared worst: grade 5 passages averaged 66.42 against an expected range of 75 to 90, and grade 6 texts averaged 67.4 against a benchmark of 80, with the most difficult passage suitable only for twelfth graders. Across the series, longer sentences and a surplus of polysyllabic words were the principal culprits.
Surface formulas, however, only reveal part of the picture, and critics have long argued that readability indexes built on sentence length and syllable counts ignore deeper cognitive dimensions of reading. A word like ‘television’ is polysyllabic yet concrete and familiar, while a monosyllabic word like ‘myth’ may defeat a young learner. To address this, Wubalem ran every passage through Coh-Metrix, a computational linguistics tool that quantifies lexical easability, syntactic density, syntactic cohesion, and semantic richness. The lexical analysis found that concreteness, imaginability, meaningfulness, and age-of-acquisition scores all exceeded expected cutoffs by roughly 3 to 5 points, meaning the passages were saturated with abstract words acquired late in a native speaker’s development. Grade 3 texts, for example, asked eight-year-olds to process terms such as ‘sustainability,’ ‘biodiversity,’ ‘photosynthesis,’ and ‘deforestation.’
Syntactic analysis told a similar story. Coh-Metrix norms suggest that young learners should encounter heavily modified noun, verb, and prepositional phrases in no more than about two sentences out of ten, with agentless passives and left-embedded clauses nearly absent. While grade 5 texts were only moderately dense, several passages in grades 3, 5, and 6 packed stretched modifiers, embedded clauses, and passive constructions into single sentences. One grade 6 passage on mobile phones contained a sentence with multiple embedded qualifiers and a final gerund clause requiring sophisticated parsing. From the perspective of cognitive load theory, such constructions overwhelm working memory precisely when grammatical processing has not yet been automatized, diverting mental resources away from meaning-making.
Interestingly, cohesion emerged as the texts’ one genuine strength. Measures of referential cohesion, stem overlap, argument overlap, and temporal cohesion were largely above the norms for each grade level, meaning the texts do supply explicit linguistic threads linking sentences. Yet the study identified what it calls a threshold effect: the extreme lexical and syntactic complexity exhausted learners’ cognitive resources before those cohesive scaffolds could be exploited. In other words, children could not benefit from well-connected discourse because they were already drowning in unfamiliar vocabulary and convoluted grammar. A related deficit appeared in semantic richness, with WordNet connectivity, latent semantic overlap, and intentionality falling below expected levels in most passages, forcing learners to make excessive inferences to construct coherent mental models of the content.
The cloze tests provided the human validation for these computational findings. Readability conventions classify texts as independent if students score above 60 percent on cloze measures, instructional between 40 and 60 percent, and frustrating below 40 percent. Learners in grades 3, 5, and 6 all scored below 35 percent on the textbook-derived tests, squarely in the frustration zone, while grade 4 students averaged 45 percent, reaching instructional level. Crucially, the same students scored more than 50 percent on standard cloze tests calibrated to their curriculum level, and independent-samples comparisons showed statistically significant gaps of 21 points in grade 3, 25 points in grade 5, and 26.5 points in grade 6. This pattern demonstrates that the failure lies in the materials, not in the children’s underlying comprehension abilities, and that the texts fall outside the learners’ zone of proximal development.
Teacher interviews reinforced the quantitative evidence. One grade 3 teacher described the texts as full of complex, abstract words that children could not process; another reported that repeated attempts to train inference and contextual word-guessing strategies failed because the passages were too complex to practice on. A third pointed to subordinate clauses, coordinating conjunctions, participle phrases, and abstract vocabulary blocking comprehension in grades 3 and 5. Educational administrators added an important caveat: learners’ difficulties were compounded by insufficient prior language input, pandemic-era schooling disruptions, and security problems in recent years, suggesting cautious interpretation of the absolute scores even as the text-based deficits remain clear.
The study’s implications are direct. Wubalem calls for aggressive simplification of overly complex passages, strategic replacement of the worst offenders, robust teacher training, and ongoing resources to help educators mediate texts that cannot be fixed quickly. Grade 4 demonstrates that appropriate leveling is achievable within the same national curriculum framework, making the failures of grades 3, 5, and 6 all the more conspicuous. The findings also speak to a broader global problem, since readability research remains scarce across much of Africa, Asia, and Latin America, where textbooks are often the only printed source of target-language input. As generative text-analysis tools like Coh-Metrix become more accessible, the study suggests education systems worldwide can no longer plead ignorance about whether their reading materials match the children assigned to read them.
Subject of Research: Readability assessment of reading passages in Ethiopian primary EFL textbooks for grades 3–6.
Article Title: An assessment on the accessibility of reading passages in Ethiopian primary EFL textbooks
Article References: An assessment on the accessibility of reading passages in Ethiopian primary EFL textbooks. (n.d.). https://doi.org/10.1007/s44217-026-02113-5
Image Credits: AI Generated
DOI: 10.1007/s44217-026-02113-5
Keywords: readability, EFL textbooks, Coh-Metrix, Flesch readability formula, cloze test, lexical density, syntactic complexity, semantic richness, reading comprehension, Ethiopia, primary education, text difficulty
Cite Scienmag News
Courtney Benton. (September 12, 2026). Ethiopian Primary English Textbooks Are Too Hard for the Children Reading Them. Scienmag. https://scienmag.com/ethiopian-primary-english-textbooks-are-too-hard-for-the-children-reading-them/
Courtney Benton. "Ethiopian Primary English Textbooks Are Too Hard for the Children Reading Them." Scienmag, 12 September 2026, https://scienmag.com/ethiopian-primary-english-textbooks-are-too-hard-for-the-children-reading-them/. Accessed 12 September 2026.
Courtney Benton. "Ethiopian Primary English Textbooks Are Too Hard for the Children Reading Them." Scienmag. September 12, 2026. https://scienmag.com/ethiopian-primary-english-textbooks-are-too-hard-for-the-children-reading-them/

