What happens in the mind when a reader encounters a string of letters that looks like a word but is not one? For decades, psycholinguists have treated such stimuli as little more than convenient foils in the lexical decision task, the classic experiment in which participants must rapidly judge whether a letter string is a real word or not. Words have always been the stars of the show, with researchers meticulously cataloguing how factors such as word frequency, length, and orthographic neighborhood shape the speed and accuracy of recognition. The strings that fail the word test, by contrast, have languished in the background, understudied and undertheorized. A new open-access resource published in Behavior Research Methods now aims to change that, offering the largest and most carefully controlled behavioral dataset ever assembled for English pseudowords and nonwords.
The English Pseudoword Lexicon Project, or EPLeP, was created by Fabio Marson, Rolando Bonandrini, Iva Šaban, Simona Amenta, Marco Marelli, and colleagues at the University of Milano-Bicocca as a registered report. The team collected lexical decision data for nearly 13,500 non-lexical stimuli and roughly 4,400 real words from 1,416 native English speakers recruited through the Prolific platform, all based in the United States. Across four online experiments, participants judged whether letter strings were words, generating more than 1.4 million individual trials. The sheer scale of the effort places EPLeP alongside the great megastudies of visual word recognition, but with a crucial difference: here, the non-lexical stimuli are the main event rather than an afterthought.
The design of the stimulus sets reflects a sophisticated taxonomy of non-words. The researchers distinguished three categories. Morphologically simple pseudowords are unattested letter strings that obey the graphotactic rules of English but cannot be decomposed into a stem plus an affix, items such as plief or dokstine. Morphologically complex pseudowords are built by combining real stems with real suffixes, yielding items such as plungable or snottant, and respecting the orthographic adjustments that English affixation typically requires. Finally, nonwords are letter strings that violate English graphotactics altogether, such as iprwvca or pghtmyg, making them conspicuously unlike anything in the language. Words in the database were likewise split into morphologically simple and morphologically complex sets, drawn from the MorphoLex database and filtered through frequency corpora including SUBTLEX-UK and SUBTLEX-US.
This careful construction matters because previous megastudies, however large, shared systematic limitations when it came to non-lexical material. None included graphotactically illegal nonwords, morphological complexity of pseudowords was never systematically controlled, and the lack of stimulus control opened the door to list effects, in which an item appears easier or harder to process simply because of what else appears alongside it. EPLeP addresses all three problems by running separate experiments for each pseudoword type, plus a fourth experiment in which all three types were intermixed with an equal number of items, allowing the researchers to compare between-experiment and within-experiment processing of the same stimuli.
The headline finding is that not all non-words are processed alike, and the differences are striking. Morphologically complex pseudowords triggered the deepest level of processing, taking the longest to reject, while graphotactically illegal nonwords were rejected fastest and with the shallowest engagement of lexical machinery. Simple pseudowords fell in between. This graded pattern suggests that the visual word recognition system responds to the degree of linguistic structure in a letter string: the more an item resembles a well-formed English word, complete with a plausible stem and suffix, the more it activates lexical and morphological processes, and the harder it becomes to dismiss as a non-word. The result replicates a long line of evidence showing that even nonexistent strings undergo morphological decomposition, much as real words like goodness are parsed into good and ness.
Perhaps the most consequential discovery concerns how the type of pseudoword in an experiment reshapes the processing of real words. When words appeared in lists containing morphologically complex pseudowords, the classic effects of orthographic neighborhood, length, and frequency on word recognition were amplified. The researchers interpret this as evidence of heightened lexical competition: because complex pseudowords are decomposed into stems and suffixes and can even evoke plausible meanings, they flood the system with lexical and semantic activation, forcing participants to scrutinize words more deeply. Conversely, when words were paired with graphotactically illegal nonwords, these psycholinguistic effects shrank, and in an ancillary analysis the usual advantage for morphologically simple words over complex ones was even reversed. In other words, everything the field knows about word recognition from lexical decision studies is partly a function of the pseudowords used in each study, a conclusion the authors describe as a call for a contextual reframing of the literature.
The dataset also delivered some surprises that complicate standard assumptions. Contrary to expectations, the researchers observed a reversed lexicality effect on response times, with non-lexical items answered faster than words, likely driven by the ceiling-level performance on highly implausible nonwords and by flexible response strategies participants adopt depending on list composition. Reliability analyses showed that response times and accuracy for simple and complex pseudowords reached levels comparable to those reported for words in landmark megastudies such as the French Lexicon Project, but reliability for nonwords was substantially low, apparently because near-perfect performance leaves almost no systematic variance to measure. External validation was successful: response times for the 300 simple pseudowords shared with the British Lexicon Project correlated significantly with the earlier data, confirming that the new measurements are trustworthy.
Methodologically, the project also offers a cautionary tale about measurement. The orthographic Levenshtein distance metric known as OLD20, which quantifies how similar a string is to its twenty closest real-word neighbors, behaved in this dataset much like a proxy for stimulus length, with the two variables highly correlated. The authors advise future users of the resource to interpret OLD20 effects cautiously and to rely on length when analyzing the data. Such transparency about the limits of standard psycholinguistic measures is itself a contribution, and it illustrates the value of megastudy datasets for stress-testing the tools the field takes for granted.
The applications of EPLeP extend well beyond theoretical psycholinguistics. Researchers building computational models of visual word recognition can now evaluate their simulations against a rich behavioral benchmark that systematically varies orthographic legality and morphological structure. Experimentalists can use the database to select stimuli with known performance profiles for priming studies or new lexical decision experiments, and clinicians can draw on it to construct diagnostic tests of pseudoword processing for populations with reading disorders such as pure alexia or dyslexia. The data, along with stimulus generation code and analysis scripts, are freely available on the Open Science Framework, and the registered report format ensured that the analyses were preregistered before data collection began.
The broader message of the project is that pseudowords deserve a central place in the science of reading. Recent work has shown that even strings never encountered before can carry semantic weight, activating meaning-related patterns and producing frequency-like effects, and EPLeP now provides the large-scale evidence base to pursue those questions systematically. Although the findings come from English, a language whose orthographic quirks make it an outlier among the world’s writing systems, the authors argue that the underlying architecture of lexical-semantic processing is likely language-independent, and they call for replications in more transparent orthographies, agglutinative languages, and writing systems such as Mandarin where word boundaries are less clearly defined. For now, the English Pseudoword Lexicon Project stands as a reminder that some of the most informative things a mind ever reads are the words that were never there at all.
Subject of Research: A large-scale lexical decision database for investigating how English pseudowords and nonwords are processed during visual word recognition.
Article Title: The English Pseudoword Lexicon Project (EPLeP): A resource for investigating pseudoword processing
Article References: Marson, F., Bonandrini, R., Šaban, I., Amenta, S., & Marelli, M. (2026). The English Pseudoword Lexicon Project (EPLeP): A resource for investigating pseudoword processing. Behavior Research Methods, 58(11), Article 297. https://doi.org/10.3758/s13428-026-03151-5
Image Credits: AI Generated
DOI: 10.3758/s13428-026-03151-5
Keywords: pseudowords, lexical decision, psycholinguistics, visual word recognition, megastudy, morphology, nonwords, mental lexicon, orthographic neighborhood, Behavior Research Methods, open science, word recognition
Cite Scienmag News
Cassandra Pierce. (September 22, 2026). Massive new database reveals how the brain handles words that do not exist. Scienmag. https://scienmag.com/massive-new-database-reveals-how-the-brain-handles-words-that-do-not-exist/
Cassandra Pierce. "Massive new database reveals how the brain handles words that do not exist." Scienmag, 22 September 2026, https://scienmag.com/massive-new-database-reveals-how-the-brain-handles-words-that-do-not-exist/. Accessed 22 September 2026.
Cassandra Pierce. "Massive new database reveals how the brain handles words that do not exist." Scienmag. September 22, 2026. https://scienmag.com/massive-new-database-reveals-how-the-brain-handles-words-that-do-not-exist/

