<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>word frequency &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/word-frequency/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 04:22:15 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>word frequency &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Massive New Database Reveals When We Learn Every Word of Russian</title>
		<link>https://scienmag.com/massive-new-database-reveals-when-we-learn-every-word-of-russian/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 04:22:15 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[age of acquisition]]></category>
		<category><![CDATA[Behavior Research Methods]]></category>
		<category><![CDATA[childhood language acquisition]]></category>
		<category><![CDATA[cognitive psychology]]></category>
		<category><![CDATA[cognitive psychology of language]]></category>
		<category><![CDATA[comprehensive language acquisition resources]]></category>
		<category><![CDATA[impact of early word learning]]></category>
		<category><![CDATA[language acquisition]]></category>
		<category><![CDATA[language processing and reaction times]]></category>
		<category><![CDATA[large-scale Russian vocabulary database]]></category>
		<category><![CDATA[lexical decision task research]]></category>
		<category><![CDATA[lexical norms]]></category>
		<category><![CDATA[lexical processing]]></category>
		<category><![CDATA[longitudinal language learning studies]]></category>
		<category><![CDATA[open data]]></category>
		<category><![CDATA[psycholinguistics]]></category>
		<category><![CDATA[retrospective memory accuracy]]></category>
		<category><![CDATA[Russian language]]></category>
		<category><![CDATA[Russian language vocabulary development]]></category>
		<category><![CDATA[vocabulary development]]></category>
		<category><![CDATA[vocabulary validation in children]]></category>
		<category><![CDATA[word frequency]]></category>
		<category><![CDATA[word recognition]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=225650</guid>

					<description><![CDATA[Researchers have collected and validated age-of-acquisition ratings for 30,849 Russian words from over 2,200 adults, creating the largest such database for any Slavic language and showing that adults' memories of when they learned words closely match children's actual vocabulary development.]]></description>
										<content:encoded><![CDATA[<p>When did you first understand the word &#8220;dragon&#8221;? For most people the answer comes easily, even decades later, and it turns out that these retrospective memories are far more accurate than they might seem. A team of Russian psychologists has now harnessed this phenomenon on an unprecedented scale, collecting age-of-acquisition estimates for 30,849 Russian words from more than 2,200 adult speakers and validating them against the actual vocabulary knowledge of schoolchildren. The result, published in Behavior Research Methods, is the largest and most comprehensive age-of-acquisition resource ever assembled for the Russian language, and one of the largest for any language in the world.</p>
<p>Age of acquisition, usually abbreviated AoA, refers to the age at which a person first learns a word in their native language. For half a century, cognitive psychologists have documented that this seemingly simple variable exerts a powerful influence on how the mind processes language. Pictures of objects whose names were learned early in life are named faster, words acquired early are read faster, and in the lexical decision task, where participants judge whether a letter string is a real word, both reaction times and error rates depend on when the word entered a person&#8217;s vocabulary. The effect even extends to semantic tasks: words learned early generate more associates in word-association experiments and are categorized more quickly, although interestingly the effect does not appear in picture categorization. Because AoA shapes performance across such a broad range of tasks, researchers treat it as one of the essential psycholinguistic variables, alongside word frequency, length, and concreteness, that must be controlled when designing experiments on reading and lexical processing.</p>
<p>The standard way to measure AoA is simply to ask adults. Participants estimate, in years, the age at which they believe they learned each word, producing what researchers call subjective AoA ratings. Such norms now exist for Chinese, Croatian, Dutch, English, French, German, Icelandic, Italian, Japanese, Portuguese, Spanish, Turkish, and many other languages. Methodological preferences have shifted over time: early studies used coarse rating scales, such as a seven-point scale in which the lowest point covered ages zero to two and the highest covered thirteen and older. More recent work favors asking respondents to report a specific age in years, a technique that avoids artificially restricting the response range and is easier for participants to use. Cross-language comparisons show that the order in which words are learned is remarkably consistent across cultures, with reliability-adjusted correlations between languages reaching as high as .96 between Polish and Slovak.</p>
<p>Until now, Russian researchers have worked with a patchwork of small datasets. Existing norms covered only a few hundred items each: 260 nouns denoting pictured objects, 190 and 375 verbs in separate studies, 696 nouns, 414 verbs, 475 adjectives, and 506 concrete and abstract nouns. Most of these ratings were never validated against children&#8217;s actual word knowledge. The new dataset changes that picture dramatically, expanding the lexical coverage of Russian AoA norms by two orders of magnitude and providing, for the first time, a resource comparable in scope to the major English and Dutch norms that each cover roughly 30,000 words.</p>
<p>The scale of the data collection effort was considerable. Approximately 3,000 participants took part, recruited both in person, primarily first-year university students and faculty members, and online through the Yandex.Toloka crowdsourcing platform, where respondents received about $2.50 in compensation. After a battery of quality checks, 2,201 protocols were retained, roughly 73 percent of the total. The respondents, whose mean age was 27.2 years, rated words organized into 103 lists of 300 target words each. Each list also contained ten calibrator words, included to show respondents the possible range of ratings, and thirty control words with previously established ratings, used to verify the validity of each person&#8217;s responses. Data collection proceeded in three annual phases between 2022 and 2024, with ratings gathered for 7,500 words in the first year, 16,500 in the second, and 6,900 in the third.</p>
<p>The stimulus set was drawn mainly from a standard frequency dictionary of contemporary Russian, supplemented by words from category norms and other items familiar to modern speakers. The researchers deliberately sampled across grammatical categories: within each list, 134 items were nouns, 76 verbs, 70 adjectives, 17 adverbs, and 3 belonged to other parts of speech, mirroring the distribution in the source dictionary. Words from the same derivational family, such as &#8220;cunning&#8221; and &#8220;craftiness,&#8221; were generally kept in different lists to avoid inflating similarity between ratings. The final set ranged from one to 24 letters in length, with words presented in their lemmatized form except where a plural is more common in actual usage.</p>
<p>Ensuring data quality required elaborate screening. Each protocol was checked for autocorrelation, meaning suspiciously smooth sequences of ratings, and for its correlation with the known ratings of the control words. Protocols were excluded if a respondent&#8217;s mean rating deviated more than two standard deviations from the list average, if the respondent claimed to have learned every word after age four, or if more than sixty words were marked as unknown. In the online sessions, individual responses with latencies shorter than 800 milliseconds were also discarded, and all data were additionally inspected manually. The result was remarkably consistent: a bootstrap split-half analysis, in which respondents were repeatedly divided into two random halves and their mean ratings correlated, yielded a mean raw correlation of .895 and a Spearman-Brown-corrected reliability of .944, indicating that adults provide highly stable estimates even across thousands of unfamiliar-seeming items.</p>
<p>The most striking aspect of the study is its validation strategy. Rather than relying solely on internal consistency, the researchers tested whether the adult ratings actually predict when children learn words. They constructed multiple-choice vocabulary tests in which students had to select, from five options, the word or phrase best matching a target word, following careful design principles such as ensuring the correct answer was simpler than the target and that distractors were semantically balanced. In the grade-level validation, a 300-item test was administered to students in grades 2, 4, 6, 8, and 10 in schools in the Moscow region and the city of Kursk. A word was considered known at a given grade if at least 75 percent of students answered correctly and students in all higher grades did as well. The correlation between these objective, grade-based estimates and the adult subjective ratings was .805, a remarkably strong correspondence showing that adults&#8217; retrospective judgments closely track the real developmental timeline of word learning.</p>
<p>A second procedure, within-grade validation, examined whether the ratings predict accuracy among children of the same age. Two rounds of testing, eight to ten months apart, involved 700 schoolchildren in total. Across every grade level, the correlations between subjective AoA and the proportion of correct answers were consistently negative: the later a word was estimated to be acquired, the lower the accuracy on its test item, both among second-graders and among eighth-graders. After statistical corrections for measurement error and range restriction, these correlations increased further. The ratings also correlated .698 with objective AoA values obtained by asking children of different ages to name pictured objects, and correlations with the seven previous Russian datasets ranged from .68 to .91, providing converging evidence from every available angle.</p>
<p>The norms also reproduce the expected relationships with other psycholinguistic variables. Later-acquired words tend to be longer, with correlations of about .28 with word length, and less frequent, with correlations between -.38 and -.45 depending on the frequency source. They are also less concrete, less imageable, and less familiar, with the strongest association observed for imageability. Intriguingly, the study also revealed respondent-level patterns: older adults gave slightly later AoA estimates and marked far fewer words as unknown, suggesting that vocabulary knowledge continues to expand across the lifespan and that our memories of word learning may shift with our own age. With the full dataset now freely available on the Open Science Framework, researchers in psycholinguistics, reading development, computational modeling, and language education gain a powerful new tool, and cross-linguistic comparisons of how children build their lexicons become possible for Russian at a scale never before achievable.</p>
<p><strong>Subject of Research:</strong> Subjective age-of-acquisition norms for 30,849 Russian words and their validation against schoolchildren&#x27;s vocabulary knowledge</p>
<p><strong>Article Title:</strong> Subjective age of acquisition norms for 30,849 Russian words</p>
<p><strong>Article References:</strong> Subjective age of acquisition norms for 30,849 Russian words. (n.d.). <a href="https://doi.org/10.3758/s13428-026-03138-2" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03138-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03138-2" rel="noopener noreferrer">10.3758/s13428-026-03138-2</a></p>
<p><strong>Keywords:</strong> age of acquisition, psycholinguistics, Russian language, lexical norms, vocabulary development, word frequency, lexical processing, language acquisition, Behavior Research Methods, open data, word recognition, cognitive psychology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">225650</post-id>	</item>
	</channel>
</rss>
