<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>data contamination &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/data-contamination/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 24 Sep 2026 22:33:51 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>data contamination &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Agents Are Being Graded Wrong: Landmark Audit Finds No Benchmark Controls All Key Threats</title>
		<link>https://scienmag.com/ai-agents-are-being-graded-wrong-landmark-audit-finds-no-benchmark-controls-all-key-threats/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 22:33:51 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[agent evaluation]]></category>
		<category><![CDATA[AI agent evaluation]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[Artificial Intelligence Review]]></category>
		<category><![CDATA[assessment of AI threat detection and control]]></category>
		<category><![CDATA[autonomous AI decision-making]]></category>
		<category><![CDATA[benchmarking]]></category>
		<category><![CDATA[benchmarking limitations in artificial intelligence]]></category>
		<category><![CDATA[challenges in interpreting AI benchmark scores]]></category>
		<category><![CDATA[critique of current AI performance metrics]]></category>
		<category><![CDATA[data contamination]]></category>
		<category><![CDATA[evaluation infrastructure]]></category>
		<category><![CDATA[execution cost]]></category>
		<category><![CDATA[external tool integration in AI systems]]></category>
		<category><![CDATA[impact of benchmark design on AI scoring]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[long-chain reasoning in AI agents]]></category>
		<category><![CDATA[meta-taxonomy]]></category>
		<category><![CDATA[multi-step task planning in language models]]></category>
		<category><![CDATA[non-determinism]]></category>
		<category><![CDATA[PRISMA-ScR]]></category>
		<category><![CDATA[PRISMA-ScR systematic review methodology]]></category>
		<category><![CDATA[systematic review]]></category>
		<category><![CDATA[systematic review of AI benchmarks]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212879</guid>

					<description><![CDATA[A systematic survey of 259 studies finds that none of seventeen prominent AI agent benchmarks jointly controls data contamination, non-determinism, and execution cost, prompting a roadmap for reliability-first evaluation.]]></description>
										<content:encoded><![CDATA[<p>Large language models no longer simply answer questions. They plan multi-step tasks, call external tools, browse environments, and act autonomously across long chains of decisions. Yet according to a sweeping new systematic survey published in Artificial Intelligence Review, the benchmarks used to grade these AI agents produce scores that are far less meaningful than the field assumes. The study, led by Vinoth Nageshwaran of the University of the Cumberlands with colleagues at Indiana University of Pennsylvania, Ton Duc Thang University, and Elizabeth City State University, argues that a benchmark number is not self-interpreting: what it means depends entirely on which capability it measures, how that capability is scored, and where the agent is tested when it is measured.</p>
<p>The research team built their analysis on a reproducible systematic mapping review conducted under the PRISMA-ScR standard, the accepted protocol for systematic reviews in health and computer science research. After a rigorous screening pipeline, 259 primary studies formed the core analytical corpus, with 294 studies in the released living-review corpus. The authors are unusually candid about the limits of their own method: primary-arm records were screened by a single audited automated pass, while human double-screening with an inter-rater agreement of kappa equal to 0.65 was applied only to a supplementary arm. They explicitly describe broad-map attribute shares as provisional heuristic estimates and stress that the corpus is a carefully constructed mapping sample, not a census of the entire field. That kind of methodological honesty is rare in a literature often criticized for overclaiming.</p>
<p>The survey&#8217;s first major contribution is a crisp conceptual boundary called the dependent-step test. The test demarcates genuine autonomous agentic evaluation from ordinary static natural-language-processing evaluation and from prompt engineering. In essence, an evaluation qualifies as agentic only when later steps in a task depend on the outcomes of earlier steps, so that errors compound and the agent must recover, replan, or abandon a strategy mid-trajectory. A model that answers a thousand independent trivia questions is being tested very differently from one that must book a flight, notice the payment failed, diagnose why, and retry with a corrected form. The dependent-step test gives researchers a principled way to decide whether a benchmark is actually measuring agency or merely repackaging static question answering.</p>
<p>The second contribution is a meta-taxonomy that places every agent benchmark in a three-pillar coordinate system: capability, meaning what the benchmark measures; scoring paradigm, meaning how performance is judged; and environment topology, meaning where the agent operates. This framework matters because two benchmarks with identical headline scores may probe entirely different faculties. One may reward factual recall in a sandboxed text environment, while another measures tool orchestration in a live operating system where a single mistyped command can cascade into failure. By locating each benchmark along these three axes, the taxonomy lets researchers compare evaluations like coordinates on a map rather than as isolated leaderboard entries, exposing blind spots where entire capability regions remain untested.</p>
<p>Third, the authors paired an evidence map over the full corpus with a purposive, hand-verified, venue-verified landscape matrix of seventeen prominent benchmarks, achieving an inter-rater reliability of kappa equal to 0.77, a level conventionally regarded as substantial agreement. This matrix functions as a deep-dive companion to the broad map: where the corpus-wide analysis offers breadth with provisional estimates, the seventeen-benchmark matrix offers verified depth on the evaluations that most actively shape the field&#8217;s self-image. The contrast between the two instruments is itself instructive, showing how much confidence a reader should place in each kind of claim.</p>
<p>The survey&#8217;s most striking finding emerges from its critical analysis of that verified matrix. The authors identify a structural trilemma facing agent evaluation: data contamination, non-determinism, and execution cost. Data contamination occurs when benchmark tasks leak into training data, inflating scores without genuine capability. Non-determinism arises because agentic trajectories involve stochastic model behavior and environment interactions, so the same agent can score differently across runs. Execution cost reflects the computational expense of running agents through long, tool-using trajectories, which discourages the repeated trials needed to tame non-determinism. Among the seventeen prominent, verified benchmarks, the study found that none reports evidence that all three threats are jointly controlled, and none reports a complete standardized run-cost record, a 0 out of 17 result that should unsettle anyone quoting leaderboard numbers.</p>
<p>The implications of that 0 out of 17 finding ripple outward. When contamination is uncontrolled, a high score may reflect memorization rather than reasoning. When non-determinism is unquantified, a single reported run may be a lucky draw rather than a stable estimate of ability. When cost is unrecorded, results cannot be reproduced at reasonable expense, and the community cannot even assess whether an evaluation is practical to repeat. Together these gaps mean that many celebrated comparisons between competing agents may be statistically fragile, and that the field&#8217;s rapid narrative of progress rests on measurements whose error bars are largely unknown. The trilemma is structural because addressing any one threat tends to worsen another: more repetitions to handle non-determinism raise cost, and larger task sets that resist contamination raise it further.</p>
<p>The transparency of the review process itself deserves attention as a model for the field. The authors released a companion repository containing their pipeline, including a build script that reproduces every reported count against a committed HTTP response cache, yielding byte-identical, MD5-verified outputs with zero live API calls. They document per-source query status, flagging databases where only counts could be verified and one major index that remained unresolved for lack of credentials. They even ran a capture-recapture diagnostic between OpenAlex and Crossref, then declined to treat its output as a recall estimate because the two indices are heterogeneous targeted searches rather than independent random samples, a violation of the estimator&#8217;s core assumptions. A probe-set validation instead confirmed that canonical benchmarks were captured completely, with Scopus corroborating the result. This is bibliometrics done with unusual care.</p>
<p>Looking forward, the survey proposes a five-direction roadmap that reframes progress not as the invention of yet more individual benchmarks but as a shift toward standardized, reliability-first, cost-aware evaluation infrastructure. In practice that means community-agreed protocols for reporting contamination checks, variance across repeated runs, and the compute cost of every evaluation, so that a benchmark score arrives with the same kind of measurement metadata that a physics experiment or clinical trial would demand. The roadmap effectively asks the field to grow up: to treat agent evaluation as a measurement science with known instruments, calibrated error, and reproducible procedures, rather than as a proliferation of leaderboards each with its own unstated assumptions.</p>
<p>For an industry pouring billions into agentic AI products, the stakes could hardly be higher. Autonomous agents are beginning to write code, manage workflows, and interact with real systems on behalf of users, and deployment decisions increasingly cite benchmark performance as evidence of safety and competence. If, as this survey shows, no prominent benchmark currently demonstrates joint control over contamination, non-determinism, and cost, then the numbers guiding those decisions are weaker than they appear. The study does not claim agents are failing; it claims we cannot yet reliably tell. Turning evaluation from a competitive spectacle into trustworthy infrastructure, the authors argue, is the prerequisite for knowing what today&#8217;s agents can actually do, and the open repository accompanying the paper offers the community a concrete place to start.</p>
<p><strong>Subject of Research:</strong> Evaluation and benchmarking of large language model agents</p>
<p><strong>Article Title:</strong> Large language model agent evaluation and benchmarking: a systematic survey, meta-taxonomy, and critical research roadmap</p>
<p><strong>Article References:</strong> Nageshwaran, V., Ezekiel, S., Tran, T. T., &amp; Lakshmi Narasimhan, V. (2026). Large language model agent evaluation and benchmarking: a systematic survey, meta-taxonomy, and critical research roadmap. <em>Artificial Intelligence Review</em>. <a href="https://doi.org/10.1007/s10462-026-11678-4" rel="noopener noreferrer">https://doi.org/10.1007/s10462-026-11678-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10462-026-11678-4" rel="noopener noreferrer">10.1007/s10462-026-11678-4</a></p>
<p><strong>Keywords:</strong> large language models, AI agents, benchmarking, agent evaluation, meta-taxonomy, systematic review, data contamination, non-determinism, execution cost, PRISMA-ScR, evaluation infrastructure, Artificial Intelligence Review</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212879</post-id>	</item>
		<item>
		<title>Bayesian Reasoning Problems Could Expose AI Bots Hiding in Online Surveys</title>
		<link>https://scienmag.com/bayesian-reasoning-problems-could-expose-ai-bots-hiding-in-online-surveys/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 21:18:36 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI detection in crowdsourced research]]></category>
		<category><![CDATA[Bayesian reasoning]]></category>
		<category><![CDATA[Bayesian reasoning in online survey validation]]></category>
		<category><![CDATA[Bayesian reasoning problems revealing AI bots]]></category>
		<category><![CDATA[behavioral research methods]]></category>
		<category><![CDATA[capability-gap test]]></category>
		<category><![CDATA[capability-gap testing for AI identification]]></category>
		<category><![CDATA[chatbot identification in research]]></category>
		<category><![CDATA[cognitive psychology methods for AI detection]]></category>
		<category><![CDATA[crowdsourcing]]></category>
		<category><![CDATA[data contamination]]></category>
		<category><![CDATA[data quality]]></category>
		<category><![CDATA[human versus AI performance in Bayesian tasks]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[large language models in behavioral studies]]></category>
		<category><![CDATA[large language models influencing survey responses]]></category>
		<category><![CDATA[natural frequencies]]></category>
		<category><![CDATA[online behavioral research integrity]]></category>
		<category><![CDATA[online research]]></category>
		<category><![CDATA[online survey data contamination]]></category>
		<category><![CDATA[positive predictive value]]></category>
		<category><![CDATA[predictive value calculation in survey validation]]></category>
		<category><![CDATA[Prolific]]></category>
		<category><![CDATA[signal detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212579</guid>

					<description><![CDATA[A new study proposes using Bayesian reasoning problems, with their well-established human performance ceilings, as calibrated detectors of large language model contamination in online research samples.]]></description>
										<content:encoded><![CDATA[<p>Online behavioral research is facing a quiet crisis. As large language models become woven into everyday life, researchers who recruit participants through crowdsourcing platforms increasingly suspect that some of their respondents are not human at all — or are humans outsourcing their answers to chatbots. A new study published in Behavior Research Methods proposes an elegant solution to this problem, and it comes from an unexpected corner of cognitive psychology: the humble Bayesian reasoning problem, a puzzle that humans have famously struggled with for fifty years.</p>
<p>The study, conducted by independent researcher Vera Wilde, introduces what she calls a capability-gap test. The logic is deceptively simple. Certain problems, such as calculating a positive predictive value from base-rate information, have been studied so extensively in humans that scientists know, with meta-analytic precision, exactly how well people can perform. When participants in an online study dramatically exceed those well-established human ceilings, the most plausible explanation is not superhuman cognition but machine assistance — a canary in the data coalmine, signaling that the sample may be contaminated by large language models.</p>
<p>The empirical basis for the proposal comes from two preregistered pilot studies of a Bayesian reasoning training tool, with a combined sample of 148 participants recruited through the Prolific platform. The tool was designed to teach people how to solve Bayesian inference problems, the kind of task exemplified by medical diagnosis questions: given a disease with a certain prevalence, a test with a certain sensitivity and false-positive rate, what is the probability that a person who tests positive actually has the disease? Decades of research, dating back to classic work by Daniel Kahneman and Amos Tversky and extended by Gerd Gigerenzer and Ulrich Hoffrage, have shown that most people fail such problems, even when the numbers are presented in natural frequency formats that make the underlying logic easier to grasp.</p>
<p>That failure is precisely what makes the task useful as a detector. A meta-analysis by McDowell and Jacobs found that only about 24 percent of people can solve a single Bayesian reasoning problem presented in natural frequency format — and that figure represents a ceiling, a level at which achieving a perfect score on a battery of five such problems is effectively unattainable for genuine human respondents. Yet in Wilde&#8217;s pilots, participants&#8217; accuracy on positive predictive value calculation problems reached roughly three times the established human performance ceiling. In the second pilot, 57 percent of participants achieved perfect 5-for-5 scores, a result that should be extraordinarily rare in an uncontaminated human sample.</p>
<p>The technical heart of the approach lies in distinguishing two outcome measures: accuracy and algorithm use. Accuracy refers simply to whether the participant produced the correct numerical answer. Algorithm use, by contrast, refers to evidence in the participant&#8217;s response that they actually followed the Bayesian reasoning process — for example, constructing a frequency tree, counting cases, or showing the intermediate steps of the calculation. This distinction matters because a training intervention designed to improve Bayesian reasoning should, if it works, change both measures in tandem. A large language model, however, can produce correct answers without any visible reasoning process, or with a reasoning process that does not respond to the training manipulation in the way human learning does.</p>
<p>By tracking both measures simultaneously, researchers can separate two rival explanations for suspiciously high performance. If accuracy spikes but algorithm use does not, contamination is the likelier culprit, because the model supplies answers without the participant acquiring the underlying skill. If both accuracy and algorithm use rise together, the pattern is consistent with authentic learning effects, and the treatment signal can be preserved rather than discarded. In this way, the capability-gap test does not merely flag bad data; it helps researchers decide which parts of their dataset reflect genuine psychological phenomena and which parts reflect machine-generated noise.</p>
<p>Wilde frames the detection problem itself as a signal detection problem, structurally analogous to the mass screenings for low-prevalence conditions — such as disease screening — around which the Bayesian reasoning literature was originally developed. Just as a medical test must balance hits against false alarms, a contamination detector must catch bot-driven responses without wrongly excluding honest participants who happen to be statistically savvy. The known reference distributions from the Bayesian reasoning literature make this calibration possible in a way that ad hoc attention checks cannot. Rather than relying on generic screening questions, researchers can compare observed performance against quantified human benchmarks and estimate the probability that a given response pattern arose from machine assistance.</p>
<p>The approach offers four practical advantages over existing data-quality tools. First, the human performance ceilings are grounded in meta-analyses rather than informal intuition, giving researchers a defensible threshold for suspicion. Second, the human–large language model performance gap on these problems is large, which increases the sensitivity of the test. Third, the known reference distributions allow nuanced assessment rather than crude pass–fail judgments. Fourth, Bayesian reasoning problems are easy to embed in existing surveys, requiring no special software or platform cooperation. Together, these properties make the method deployable at scale across the many fields — psychology, marketing, political science, epidemiology — that increasingly depend on online samples.</p>
<p>The stakes are considerable. Prior research on crowd work has documented substantial and growing use of large language models by online workers, and studies of data contamination in machine learning itself show how memorized content can masquerade as genuine capability. If a meaningful fraction of respondents in an online study are completing tasks with chatbot help, effect sizes may be distorted, replication attempts may fail for reasons that have nothing to do with the underlying science, and the credibility of entire literatures built on crowdsourced data could be undermined. The problem echoes an older statistical concern: John Tukey&#8217;s foundational work on sampling from contaminated distributions warned that even small amounts of contamination can seriously mislead inference drawn from nominally clean data.</p>
<p>Wilde is careful to note the provenance of the idea: neither pilot study was originally designed to validate a contamination detection method, and the proposal emerged from post hoc analysis of unexpectedly strong results. That origin makes the capability-gap test a promising hypothesis rather than a fully validated diagnostic, and the author provides practical recommendations for researchers who wish to use these problems as data-quality diagnostics while the validation literature matures. Both studies were preregistered on the Open Science Framework, and all data, materials, and analysis code are publicly available, allowing other teams to scrutinize and extend the approach. If the method holds up under broader testing, Bayesian reasoning problems — long a symbol of human statistical frailty — may find a second career as guardians of scientific integrity, ensuring that the data feeding behavioral science come from human minds rather than the machines trained on those minds&#8217; collective output.</p>
<p><strong>Subject of Research:</strong> Detecting large language model contamination in online behavioral research samples using Bayesian reasoning problems as capability-gap tests</p>
<p><strong>Article Title:</strong> An LLM canary in the online data coalmine: Bayesian reasoning problems as a capability-gap test for LLM contamination in online samples</p>
<p><strong>Article References:</strong> Wilde, V. (2026). An LLM canary in the online data coalmine: Bayesian reasoning problems as a capability-gap test for LLM contamination in online samples. <em>Behavior Research Methods, 58</em>(11), Article 301. <a href="https://doi.org/10.3758/s13428-026-03184-w" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03184-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03184-w" rel="noopener noreferrer">10.3758/s13428-026-03184-w</a></p>
<p><strong>Keywords:</strong> large language models, Bayesian reasoning, data contamination, online research, data quality, signal detection, natural frequencies, positive predictive value, crowdsourcing, behavioral research methods, capability-gap test, Prolific</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212579</post-id>	</item>
	</channel>
</rss>
