<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>attention checks &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/attention-checks/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 10 Oct 2026 03:44:22 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>attention checks &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Study Reveals When Statistical Models Can Catch Careless Survey Responders</title>
		<link>https://scienmag.com/new-study-reveals-when-statistical-models-can-catch-careless-survey-responders/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sat, 10 Oct 2026 03:44:22 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[attention checks]]></category>
		<category><![CDATA[Big Five]]></category>
		<category><![CDATA[careless responding]]></category>
		<category><![CDATA[Careless survey responses detection]]></category>
		<category><![CDATA[challenges in detecting careless survey behavior]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[effects of careless responding on correlation accuracy]]></category>
		<category><![CDATA[impact of insufficient effort responding on research validity]]></category>
		<category><![CDATA[implications of careless responses in psychological research]]></category>
		<category><![CDATA[influence of inattentive respondents on effect size inflation]]></category>
		<category><![CDATA[item response theory]]></category>
		<category><![CDATA[large-scale simulation studies in survey methodology]]></category>
		<category><![CDATA[latent class analysis]]></category>
		<category><![CDATA[latent variable mixture modeling in psychology]]></category>
		<category><![CDATA[methodological advancements in survey response analysis]]></category>
		<category><![CDATA[mixture models]]></category>
		<category><![CDATA[online platforms]]></category>
		<category><![CDATA[psychometrics]]></category>
		<category><![CDATA[reliability of latent variable models for data cleaning]]></category>
		<category><![CDATA[simulation study]]></category>
		<category><![CDATA[statistical techniques for identifying inattentive survey takers]]></category>
		<category><![CDATA[straightlining]]></category>
		<category><![CDATA[survey data quality assurance methods]]></category>
		<category><![CDATA[survey methodology]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=257290</guid>

					<description><![CDATA[A comprehensive simulation and reanalysis of Big Five data shows that latent mixture models reliably detect careless survey responders when scales contain at least ten items, mixed item wording, and heterogeneous thresholds, with classifications tracking attention checks and platform quality more sensitively than traditional indices.]]></description>
										<content:encoded><![CDATA[<p>Every year, millions of people fill out psychological questionnaires, and a substantial fraction of them are not really paying attention. They click through items in seconds, pick the same answer over and over, or respond essentially at random. This behavior, known as careless and insufficient effort responding (C/IER), is one of the quiet plagues of survey research, quietly corrupting correlations, inflating effect sizes, and in extreme cases even reversing the sign of scientific findings. Now, a large-scale simulation and empirical study published in Behavior Research Methods by Irina Uglanova and Gabriel Nagy of the Leibniz Institute for Science and Mathematics Education and Esther Ulitzsch of the University of Oslo has mapped out precisely when one of the most promising detection tools—latent variable mixture modeling—can be trusted to separate the attentive from the careless, and when it fails.</p>
<p>The stakes are higher than many researchers realize. Previous work cited in the study shows that a careless responding rate of roughly 14 percent inflated correlations in published psychological studies by about .07, while simulations have demonstrated that even 5 percent careless responding can produce sign reversals in estimated correlations. Other research has found that removing as little as 4 to 10 percent of careless responders improved model fit in empirical datasets, and that just 10 percent contamination can substantially bias estimates of item discrimination parameters. Because researchers never know in advance how contaminated their data are, the authors argue, evaluating results with and without adjustments for presumed carelessness is essential for trustworthy science.</p>
<p>Traditionally, researchers have hunted for careless responders using screening indices: the long-string index that counts how many identical answers appear in a row, the intraindividual response variability index, the even–odd consistency index that correlates halves of a scale, Mahalanobis distance for multivariate outliers, and attention check items planted in the questionnaire. The trouble is that each index catches only one flavor of carelessness, so researchers typically stack several of them into a multihurdle approach. Worse, every hurdle requires a threshold—a cutoff above which a respondent is declared careless—and there are no firmly grounded rules for setting those thresholds. Small changes in cutoff choices have repeatedly been shown to produce substantially different classification rates, handing researchers enormous and often invisible degrees of freedom.</p>
<p>Mixture models offer an alternative that sidesteps the threshold problem entirely. The idea is elegant: assume that the sample contains two latent groups, one of attentive respondents whose answers reflect the traits being measured, and one of careless respondents whose answers do not. A standard measurement model—in this study, a multidimensional graded response model appropriate for categorical Likert items—links responses to latent traits in the attentive class. In the careless class, the model assumes responses do not depend on item content at all: threshold parameters are constrained to be equal across items, and a single latent variable representing each respondent&#8217;s preference for certain response categories replaces the substantive traits. The model then estimates, for every person, a posterior probability of belonging to the attentive class, conveying not just a verdict but a degree of certainty. A posterior probability of .60 signals that the response pattern says little either way; a .90 signals strong evidence of attentiveness.</p>
<p>The exemplar model used in the study was deliberately designed to capture the core features shared by most existing mixture approaches for C/IER while avoiding their more restrictive assumptions. Unlike many predecessors that treat item responses as continuous, it acknowledges their categorical nature, allowing the shape of response distributions to differ between classes—for instance, roughly normal for attentive respondents versus uniform for random responders. It also accommodates dependencies among careless responses arising from individual differences in category preferences. Because the specification retains the two-class structure common to the wider family of mixture models, the authors expect their findings to generalize to closely related approaches.</p>
<p>The heart of the study is a simulation comprising 240 crossed conditions, each replicated 200 times, with samples of 500 respondents and 10 percent careless responding—a realistic baseline given reported rates of roughly 5 to 12 percent in offline studies and 6 to 10 percent in large-scale assessments, but also an extreme class imbalance that makes detection genuinely hard. The researchers varied the number of five-point Likert items (5, 10, or 20), the proportion of negatively worded items (0 or 40 percent), the heterogeneity of item thresholds, the strength of item discriminations, and the type of careless behavior: straightlining, respondent-specific category preference, random responding, or a blend of the last two. Three criteria were evaluated: whether model comparisons correctly detected the presence of C/IER using the Bayesian information criterion, whether the estimated proportion of careless respondents was accurate, and whether individuals were correctly classified.</p>
<p>The results reveal a clear recipe for success. When the data contained no careless responders at all, the simpler model was always correctly preferred—no false alarms. When 10 percent careless responding was present, straightlining was detected consistently under every condition tested. For the subtler patterns, however, questionnaire design mattered enormously. With only five items, reliable detection generally required scales mixing positively and negatively worded items together with high item heterogeneity, meaning items differ substantially in how easy they are to endorse. With ten items, either mixed wording or high heterogeneity sufficed. With twenty items, the mixture model was preferred in nearly every condition regardless of wording or heterogeneity. The logic is intuitive: on a short, uniformly worded scale of easy-to-endorse items, an attentive respondent who consistently picks the category that best describes them produces a pattern nearly indistinguishable from a straightliner. Paradoxically, stronger item discriminations can make this problem worse, sharpening the tendency of attentive respondents to answer uniformly, while under random responding the effect reverses and weak discriminations become the hazard.</p>
<p>Individual-level classification told a similar story. Specificity—the probability of correctly assigning simulated attentive respondents to the attentive class—exceeded .90 in every single condition, partly because with only 10 percent contamination, a classifier defaulting to attentive already achieves 90 percent specificity. Sensitivity, the ability to catch true careless responders, depended heavily on scale length, item heterogeneity, and wording, rising with all three for most careless response types. For estimating the overall proportion of careless respondents, mixed-worded scales with high heterogeneity kept bias within acceptable bounds even at five items, while unidirectional, low-heterogeneity, short scales inflated estimates toward 14 percent. Straightlining proved uniquely stubborn: the proportion of careless respondents was systematically underestimated, dropping as low as 6 percent, and longer scales actually made the underestimation worse rather than better.</p>
<p>To show that the latent classes really mean what researchers hope they mean, the team reanalyzed a publicly available Big Five dataset collected across four online crowdsourcing platforms—MTurk, CloudResearch, Prolific, and Qualtrics—with roughly 500 respondents each, using the 50-item NEO-PI-R inventory in which each trait is measured by ten items, half negatively worded. By the simulation&#8217;s own criteria, these scales offered favorable conditions for detection. The model classified between 76.5 and 83.3 percent of respondents as attentive depending on the scale, and classifications were strikingly consistent: 62.4 percent of respondents were classified as attentive on all five scales, while only 8.4 percent failed to reach the attentive class on any scale, with posterior probabilities across scales correlating between .509 and .666. This cross-scale stability suggests the model captures a stable disposition toward careless responding rather than scale-specific quirks.</p>
<p>Crucially, the model-based classifications converged with evidence that does not rely on response patterns at all. Respondents who passed more of the five embedded attention check items received higher posterior probabilities of being attentive, with virtually all respondents passing four or five checks classified as attentive with near-certainty across all scales. Even more striking, the latent class variable tracked the data-collection platform more strongly than the attention check index did: averaged across scales, Nagelkerke&#8217;s pseudo-R-squared was 0.224 for the model-based classification versus 0.071 for the index. Both methods identified MTurk as by far the worst platform, but the model painted a starker picture—only 47 to 59 percent of MTurk respondents were classified as attentive on some scales, compared with 85.6 percent flagged by attention checks, while Prolific fared best at 88 to 95 percent attentive. The authors conclude that with scales of at least ten items, mixed wording, and sufficient item heterogeneity, mixture modeling is a viable, comparatively easy-to-implement tool for addressing careless responding—available in standard software such as the R package mirt—while cautioning that short, uniformly worded, low-heterogeneity scales remain its blind spot, and that questions about trait-score accuracy and misclassification consequences await future research.</p>
<p><strong>Subject of Research:</strong> Detection of careless and insufficient effort responding in self-report questionnaires using latent variable mixture models</p>
<p><strong>Article Title:</strong> Understanding when mixture models effectively identify careless responding: Simulation and validity evidence</p>
<p><strong>Article References:</strong> Uglanova, I., Nagy, G., &amp; Ulitzsch, E. (2026). Understanding when mixture models effectively identify careless responding: Simulation and validity evidence. <em>Behavior Research Methods, 58</em>(11), Article 313. <a href="https://doi.org/10.3758/s13428-026-03165-z" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03165-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03165-z" rel="noopener noreferrer">10.3758/s13428-026-03165-z</a></p>
<p><strong>Keywords:</strong> careless responding, mixture models, item response theory, survey methodology, psychometrics, latent class analysis, attention checks, straightlining, Big Five, data quality, online platforms, simulation study</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">257290</post-id>	</item>
	</channel>
</rss>
