Every word a Greek child reads in elementary school is now catalogued, tagged, and dissected in unprecedented detail. A team of linguists and psychologists led by Anthi Revithiadou of Aristotle University of Thessaloniki has unveiled HelexKids 2.0, a dramatically expanded and linguistically annotated version of the HelexKids word frequency database, published in the journal Behavior Research Methods. The resource covers 67,802 distinct word types drawn from more than 1.3 million word tokens found in the official textbooks and workbooks used in primary schools across Greece and Cyprus, and it is freely available to researchers and educators worldwide.
The original HelexKids, released in 2017, was the first psycholinguistic database built specifically for Greek primary school children aged six to twelve. It compiled frequency information from 76 textbooks spanning six grade levels and subjects ranging from language arts to mathematics, science, and history. That earlier version revealed striking facts about the vocabulary children encounter: roughly half of the words at each grade level appeared only once in the entire corpus, words were shorter and more frequent than in adult Greek databases, and the top 100 words alone accounted for nearly 45 percent of all tokens. What it lacked, however, was grammatical and phonological depth, a gap that limited its usefulness for studies of reading development, speech and spelling interventions, and theoretical work on how sound and grammar interact.
HelexKids 2.0 closes that gap with a rich annotation scheme. Every entry now carries a part-of-speech tag, syllable count, a full phonetic transcription in the International Phonetic Alphabet, orthographic and phonological syllabification, word-level syllable templates, and detailed stress pattern information. Nouns receive additional morphological annotation covering case, number, gender, and inflection class. The database also retains all original frequency measures, including raw counts, Zipf values, orthographic Levenshtein distance, dispersion, and contextual diversity, while adding new metrics such as phonological Levenshtein distance and bigram and biphone frequencies.
Building the annotations was far from a purely mechanical exercise. The team first corrected systematic character-encoding errors introduced when scanned textbooks were converted to spreadsheets, in which Latin letters had been substituted for visually similar Greek characters. Initial part-of-speech tagging drew on two existing annotated resources, GreekLex 2 and the A-Clean database, but roughly 45,000 words, about two-thirds of the total, remained untagged. The researchers then tested spaCy’s machine learning models for Greek, which significantly underperformed: the word for ‘bag’ was tagged as a proper noun, and ‘despair’ was labeled a verb. Ambiguity posed a deeper problem, since the textbook words came without sentence context. A form like ‘διορθώσεις’ can be either the plural noun ‘corrections’ or the verb ‘you correct’, and words such as ‘ένα’ can serve as a numeral, an article, or a pronoun depending on context.
Faced with these challenges, the team chose to manually annotate the entire lexicon using linguistically motivated criteria drawn from authoritative Greek grammars and dictionaries. Where genuine ambiguity existed, they assigned compound tags such as ADV/NOUN or PCP/ADJ/NOUN to capture a word’s multifunctionality. Quality control was rigorous: three researchers independently annotated a random sample of 1,000 words, yielding a Krippendorff’s alpha of 0.789 for part-of-speech tagging. After refining the guidelines, the final scheme contained 61 distinct categories and category combinations. Morphological annotation of nouns achieved even stronger agreement, with alpha reaching 0.827.
The phonological layer demanded equally careful engineering. Custom R scripts automated syllabification, transcription, and stress assignment, but Greek’s notorious variability required human oversight at every step. Vowel sequences such as ‘ια’ may be realized as a single syllable with a palatal consonant, as in the word for ‘eyes’, or split across two syllables, as in some pronunciations of ‘I delete’. Monosyllabic content words like ‘earth’ carry stress without written accents, capitalized words lose their accent marks entirely, and clitic constructions produce double stress. Approximately 3.2 percent of word types received dual annotations to reflect legitimate pronunciation variation, a design choice the authors say acknowledges the gradient nature of phonological output.
The redesigned database is also a technical leap. Implemented as a relational MariaDB database with nine interconnected tables, HelexKids 2.0 assembles lexicons dynamically on demand rather than storing fixed lists. Users can query by grade level or cumulative grade range, part of speech, syllable count, stress pattern, frequency thresholds, neighborhood density, and phonotactic measures simultaneously, through a bilingual English-Greek web interface hosted on the GRADIENCE project webpage. Results export cleanly to CSV or Excel with proper handling of Greek and IPA characters.
The descriptive statistics already yield insights into how children’s linguistic input evolves across schooling. Vocabulary expands unevenly: the largest jump occurs between Grades 2 and 3, with types growing nearly 88 percent and tokens nearly 140 percent, coinciding with the introduction of more school subjects, while growth slows to about 7 percent for types between Grades 5 and 6. Words also lengthen as grades advance. Four-syllable and longer words rise from 35.1 percent of types in Grade 1 to 53.4 percent in Grade 6, and average word length grows from 3.25 to 3.75 syllables. The most common word template across all grades is a three-syllable, penultimately stressed word made entirely of open consonant-vowel syllables, though token-based analysis shows that short monosyllabic function words dominate actual reading volume.
Perhaps the most theoretically consequential findings concern stress. Greek is a morphology-dependent stress system in which accent can fall on any of the final three syllables and cannot be predicted from sound structure alone. The database shows that each major word class distributes stress differently: nouns favor penultimate stress, verbs show the highest proportion of antepenultimate stress among their types yet their most frequent token forms are penultimately stressed, adjectives overwhelmingly take final stress, and adverbs prefer penultimate stress. For nouns, penultimate stress gradually declines across grades while antepenultimate stress rises, a redistribution that leaves final stress untouched. The authors argue these patterns support accounts distinguishing morpholexically conditioned stress from phonological defaults, and reinforce proposals that nouns are prosodically more complex than verbs.
The implications extend well beyond Greek linguistics. HelexKids 2.0 joins a small international family of child-specific databases, including MANULEX for French, childLex for German, ESCOLEX for Portuguese, and CYP-LEX for English, but its phonological depth exceeds them all, offering a template for annotation projects in other languages. For educators, the resource promises practical tools: stress position is a documented difficulty for Greek children in both spelling and reading, and teachers can now generate targeted word lists organized by stress pattern, starting with high-frequency, predictable items before moving to rarer, unpredictable ones. The team acknowledges limitations, including the restriction of morphological annotation to nouns and the database’s reliance on textbooks from a specific 2007 to 2013 period, with new textbooks scheduled from 2027 onward. Thanks to its expandable architecture, the annotation pipeline can absorb new materials with minimal modification, positioning HelexKids 2.0 as a living resource for developmental psycholinguistics, morphophonological theory, and evidence-based literacy instruction.
Subject of Research: A linguistically annotated Greek child lexical database with part-of-speech, phonological, and morphological annotations
Article Title: HelexKids 2.0: A linguistically annotated lexical database building on HelexKids
Article References: Revithiadou, A., Terzopoulos, A., Niolaki, G., Markopoulos, G., Avdelidis, K., Mittas, I., & Kosmidis, K. (2026). HelexKids 2.0: A linguistically annotated lexical database building on HelexKids. Behavior Research Methods, 58(11), Article 311. https://doi.org/10.3758/s13428-026-03179-7
Image Credits: AI Generated
DOI: 10.3758/s13428-026-03179-7
Keywords: lexical database, Greek language, word frequency, psycholinguistics, stress patterns, morphophonology, elementary education, reading development, phonetic transcription, part-of-speech tagging, textbook vocabulary, open access resource
Cite Scienmag News
Glenn Wilkins. (October 7, 2026). Greek Children’s Textbooks Power a New Linguistically Annotated Lexical Database. Scienmag. https://scienmag.com/greek-childrens-textbooks-power-a-new-linguistically-annotated-lexical-database/
Glenn Wilkins. "Greek Children’s Textbooks Power a New Linguistically Annotated Lexical Database." Scienmag, 7 October 2026, https://scienmag.com/greek-childrens-textbooks-power-a-new-linguistically-annotated-lexical-database/. Accessed 7 October 2026.
Glenn Wilkins. "Greek Children’s Textbooks Power a New Linguistically Annotated Lexical Database." Scienmag. October 7, 2026. https://scienmag.com/greek-childrens-textbooks-power-a-new-linguistically-annotated-lexical-database/

