<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>implicit bias &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/implicit-bias/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 20 Sep 2026 21:13:50 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>implicit bias &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI reasoning models show human-like implicit bias in how hard they think</title>
		<link>https://scienmag.com/ai-reasoning-models-show-human-like-implicit-bias-in-how-hard-they-think/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 21:13:50 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI fairness]]></category>
		<category><![CDATA[AI reasoning models]]></category>
		<category><![CDATA[algorithmic bias]]></category>
		<category><![CDATA[bias detection in AI systems]]></category>
		<category><![CDATA[bias in artificial intelligence]]></category>
		<category><![CDATA[chain-of-thought reasoning]]></category>
		<category><![CDATA[computational effort]]></category>
		<category><![CDATA[human implicit bias]]></category>
		<category><![CDATA[implicit association test]]></category>
		<category><![CDATA[implicit bias]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[measuring AI effort]]></category>
		<category><![CDATA[mental associations in AI]]></category>
		<category><![CDATA[Nature Machine Intelligence]]></category>
		<category><![CDATA[Nature Machine Intelligence study]]></category>
		<category><![CDATA[psychology of AI]]></category>
		<category><![CDATA[reasoning models]]></category>
		<category><![CDATA[reasoning process analysis]]></category>
		<category><![CDATA[reasoning-token counts]]></category>
		<category><![CDATA[RM-IAT]]></category>
		<category><![CDATA[step-by-step reasoning]]></category>
		<category><![CDATA[stereotypes]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202676</guid>

					<description><![CDATA[Researchers adapted the human implicit association test for reasoning AI models and found that stereotype-inconsistent tasks demand measurably more computational effort, revealing bias-like patterns that predict downstream behaviour.]]></description>
										<content:encoded><![CDATA[<p>For more than two decades, psychologists have probed the hidden associations inside the human mind with a deceptively simple tool: the Implicit Association Test, or IAT. By measuring how much slower people are when they must pair concepts that conflict with their automatic mental associations, researchers have uncovered patterns of bias that people themselves often do not consciously endorse. Now, a team of researchers has borrowed that logic and pointed it at a new kind of mind — the reasoning model, a class of large language models that generates explicit, step-by-step chains of thought before answering. The result is a striking demonstration that artificial systems may carry bias-like signatures not in what they say, but in how much effort it takes them to think.</p>
<p>The study, published in Nature Machine Intelligence by Messi H. J. Lee of Washington University in St. Louis and Calvin K. Lai of Rutgers University, introduces what the authors call the reasoning-model implicit association test, or RM-IAT. Rather than measuring reaction times in milliseconds, as the human IAT does, the RM-IAT measures reasoning-token counts — the number of tokens a model produces in its internal reasoning process before arriving at an answer. The central insight is an analogy of computational effort: just as humans take longer to respond when a task conflicts with their implicit associations, a reasoning model may expend more computational effort — more thinking tokens — when a task is incompatible with the associations embedded in its learned representations.</p>
<p>The technical setup mirrors the classic two-block design of the human IAT. In association-compatible blocks, the pairing of a social category with a stereotypical attribute aligns with the associations the model is presumed to have absorbed from its training data. In association-incompatible blocks, those pairings are crossed, requiring the model to reason against the grain of its learned statistics. The researchers then compared the reasoning-token counts across the two conditions. If the model&#8217;s internal associations lean in a stereotypical direction, the incompatible condition should demand more elaborate reasoning, and therefore more tokens, to reach an equally acceptable answer.</p>
<p>Across four reasoning models — OpenAI&#8217;s o3-mini, DeepSeek-R1, gpt-oss-20b and Qwen3-8B — the researchers found consistent evidence for exactly that pattern. Association-incompatible tasks reliably required greater computational effort than association-compatible tasks, producing effect sizes analogous to the latency differences observed in human IAT studies. In other words, when these models were forced to reason in ways that cut across stereotypical associations, their chain-of-thought processes grew measurably longer, as though the models needed to work harder to override a default tendency. The consistency of the effect across four architecturally different models, trained by different organisations with different pipelines, suggests the phenomenon is not an idiosyncrasy of a single system but a general property of reasoning models trained on human language.</p>
<p>One model, however, broke the pattern in a fascinating way. Claude 3.7 Sonnet exhibited reversed effects: it expended more computational effort on association-compatible tasks rather than incompatible ones. A thematic analysis of its reasoning traces revealed why. Unlike the other models, Claude 3.7 Sonnet&#8217;s internal reasoning frequently turned inward, scrutinising the possibility of bias and stereotypes in the task itself before answering. This self-monitoring — an apparent product of its safety training — meant that stereotype-consistent pairings triggered extra deliberation about whether responding in line with a stereotype was appropriate, inflating token counts precisely where the other models were fastest. The finding offers a vivid illustration that the RM-IAT does not merely detect raw statistical associations; it detects the shape of a model&#8217;s processing dynamics, including the compensatory reflexes instilled by alignment training.</p>
<p>A crucial question for any such measure is whether it captures something meaningful about behaviour or is merely a statistical curiosity. The researchers addressed this by testing convergent validity: they examined whether RM-IAT effects predicted biases in downstream model outputs on tasks known to elicit biases in large language models. They found evidence that models showing stronger RM-IAT effects also displayed measurable biases in word association tasks and in decision-making scenarios — two domains where LLM bias has been extensively documented in prior research. This predictive relationship echoes the meta-analytic literature on the human IAT, where implicit measures show modest but reliable correlations with judgement and behaviour. In the machine context, the analogy suggests that the effortful override seen in reasoning tokens is not decoupled from what the models ultimately produce.</p>
<p>The theoretical framing draws heavily from the psychology of automaticity. Implicit biases are characterised as automatic mental processes that shape perception, judgement and behaviour — processes that are efficient, unintentional and often outside awareness. The human IAT operationalises this through the principle that compatible associations facilitate processing while incompatible ones impede it. Lee and Lai&#8217;s contribution is to argue that the same logic can be applied to reasoning models, where chain-of-thought generation provides a visible, countable trace of processing. Previous studies of bias in language models have focused almost exclusively on outputs — whether the model&#8217;s answers are discriminatory, whether generated text contains stereotypes, whether word embeddings encode prejudicial associations. The RM-IAT shifts the analytical lens upstream, to the reasoning process itself, opening a window on bias-like dynamics that output-level audits can miss.</p>
<p>The findings also connect to a growing body of evidence that surface-level debiasing does not eliminate deeper associations. Prior work has shown that explicitly unbiased language models can still form biased internal associations, and that alignment techniques such as reinforcement learning from human feedback may suppress stereotypical outputs without eradicating the underlying statistical tendencies learned from training corpora. The RM-IAT results are consistent with this picture: the models studied presumably produce carefully hedged, often non-stereotypical answers on sensitive topics, yet their reasoning processes still show the signature of effortful override when tasks run against stereotypical associations. Bias, in this sense, appears to persist as a property of the models&#8217; computational dynamics even when it is masked at the output layer.</p>
<p>For AI developers and auditors, the practical implications are significant. Reasoning-token counts are cheap to measure, require no special access to model weights, and can be collected at scale, making the RM-IAT a potential screening tool for bias-like processing in deployed systems. Because the measure is sensitive to training interventions — as the Claude 3.7 Sonnet reversal demonstrates — it could serve as a diagnostic for whether safety training changes not just what models say but how they think. At the same time, the authors are careful with the interpretation: the measure captures bias-like processing differences, not proof of subjective experience or attitudes in the human sense. The analogy is structural, drawing on a shared signature of processing efficiency, rather than a claim that models possess minds with unconscious prejudice.</p>
<p>The broader significance of the work lies in the deepening dialogue between cognitive science and machine learning. Tools built to illuminate the automatic workings of the human mind are proving unexpectedly informative about statistical systems trained on human-produced text, and the patterns they reveal — effortful override, association-driven processing costs, dissociations between expression and underlying tendency — are recognisably familiar. As reasoning models become the dominant interface through which people interact with artificial intelligence, understanding the hidden dynamics of their chains of thought becomes as important as auditing their answers. The RM-IAT offers a rigorous, replicable method for doing so, and its central finding — that the hardest thinking for a reasoning model is often the thinking that goes against its associations — is likely to shape how researchers study machine bias for years to come. The authors have released their data openly on Figshare and their code on GitHub, with archived versions on Zenodo and a reproducible capsule on Code Ocean, ensuring that other laboratories can immediately extend this new bridge between the psychology of implicit bias and the engineering of reasoning machines.</p>
<p><strong>Subject of Research:</strong> Measurement of implicit-bias-like processing patterns in large language model reasoning systems</p>
<p><strong>Article Title:</strong> Implicit-bias-like patterns in reasoning models</p>
<p><strong>Article References:</strong> Lee, M. H. J., &amp; Lai, C. K. (2026). Implicit-bias-like patterns in reasoning models. <em>Nature Machine Intelligence</em>. <a href="https://doi.org/10.1038/s42256-026-01300-1" rel="noopener noreferrer">https://doi.org/10.1038/s42256-026-01300-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s42256-026-01300-1" rel="noopener noreferrer">10.1038/s42256-026-01300-1</a></p>
<p><strong>Keywords:</strong> implicit bias, reasoning models, large language models, RM-IAT, computational effort, stereotypes, AI fairness, chain-of-thought reasoning, psychology of AI, algorithmic bias, Nature Machine Intelligence, implicit association test</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202676</post-id>	</item>
		<item>
		<title>Neonatal Transfers Differ by Race and Ethnicity in Very Low Birth Weight Infants</title>
		<link>https://scienmag.com/neonatal-transfers-differ-by-race-and-ethnicity-in-very-low-birth-weight-infants/</link>
		
		<dc:creator><![CDATA[Harold Sullivan]]></dc:creator>
		<pubDate>Fri, 11 Sep 2026 23:43:49 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Pediatry]]></category>
		<category><![CDATA[AANHPI]]></category>
		<category><![CDATA[birth weight and neonatal mortality]]></category>
		<category><![CDATA[California]]></category>
		<category><![CDATA[California neonatal health disparities]]></category>
		<category><![CDATA[Health disparities]]></category>
		<category><![CDATA[healthcare equity in neonatal intensive care]]></category>
		<category><![CDATA[impact of hospital level on neonatal survival]]></category>
		<category><![CDATA[implicit bias]]></category>
		<category><![CDATA[inter-hospital neonatal transport]]></category>
		<category><![CDATA[Journal of Perinatology]]></category>
		<category><![CDATA[neonatal intensive care]]></category>
		<category><![CDATA[neonatal transfer disparities]]></category>
		<category><![CDATA[neonatal transfer policies and race]]></category>
		<category><![CDATA[neonatal transport]]></category>
		<category><![CDATA[NICU level]]></category>
		<category><![CDATA[odds ratio]]></category>
		<category><![CDATA[perinatal regionalization]]></category>
		<category><![CDATA[race and ethnicity]]></category>
		<category><![CDATA[race and ethnicity in neonatal care]]></category>
		<category><![CDATA[racial and ethnic differences in neonatal treatment]]></category>
		<category><![CDATA[racial disparities in preterm infant care]]></category>
		<category><![CDATA[sociodemographic factors in neonatal transfers]]></category>
		<category><![CDATA[very low birth weight]]></category>
		<category><![CDATA[very low birth weight infant outcomes]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=193138</guid>

					<description><![CDATA[A large California cohort study finds that very low birth weight infants of Asian American, Native Hawaiian, and Pacific Islander background are significantly less likely to undergo acute inter-hospital transport than non-Hispanic White infants after risk adjustment.]]></description>
										<content:encoded><![CDATA[<p>When a baby is born weighing very little, the hospital where that baby first receives care can shape the entire course of their life. Very low birth weight infants, defined as those weighing less than 1,500 grams at birth, are among the most medically fragile patients in any health system, and decades of research have shown that delivery at a hospital equipped with a high-level neonatal intensive care unit substantially improves their chances of survival without disability. A new study published in the Journal of Perinatology now adds a sobering dimension to this picture, revealing that the likelihood of an acute inter-hospital transport for these vulnerable newborns varies significantly by race and ethnicity, even after accounting for the hospitals where they are born, the clinical severity of their conditions, and the sociodemographic characteristics of their mothers.</p>
<p>The research, led by Sarah N. Kunz of Harvard Medical School and Beth Israel Deaconess Medical Center together with colleagues at Stanford University School of Medicine and the California Perinatal Quality Care Collaborative, examined data from California on very low birth weight, preterm infants born before 37 weeks of gestation who were less than 28 days old. The study period spanned 2012 through 2018, a window that captures a mature era of regionalized perinatal care in the nation&#8217;s most populous state. California offers a particularly valuable setting for this kind of analysis because of its sheer scale and the richness of its linked birth cohort records, which allow researchers to follow infants from birth through any subsequent acute transport between hospitals with unusual precision.</p>
<p>The design of the study was a retrospective cohort analysis, meaning the investigators looked backward at records that had already been generated by routine clinical care. Their outcome of interest was acute inter-hospital transport, the urgent movement of a newborn from one hospital to another, typically because the receiving institution can provide a level of neonatal intensive care that the discharging hospital cannot. These transports are among the highest-stakes events in neonatal medicine. A tiny infant, often weighing less than a carton of milk, is placed in a portable incubator, connected to a transport ventilator, monitored continuously, and driven or flown across traffic and distance to a destination intensive care unit. Every minute of that journey carries physiologic risk, and the quality of the stabilization before departure can determine whether the infant arrives in stable condition or in crisis.</p>
<p>To isolate the effect of race and ethnicity on transport likelihood, the team calculated odds ratios comparing each racial and ethnic group against non-Hispanic White infants. The comparison groups included non-Hispanic Black infants, infants classified as Asian American, Native Hawaiian, and Pacific Islander, often abbreviated AANHPI, and Hispanic infants. Critically, the models did not stop at this simple comparison. The investigators controlled for a battery of confounding factors organized into four domains: the characteristics of the hospital network in which the birth occurred, hospital-level factors such as the level of the neonatal intensive care unit, maternal sociodemographic characteristics, and infant clinical characteristics that reflect how sick the baby was at the time a transport decision would be made.</p>
<p>The central finding was striking in its specificity. Asian American, Native Hawaiian, and Pacific Islander infants were significantly less likely to be acutely transported than non-Hispanic White infants, with an odds ratio of 0.85 and a p-value of 0.01 after full risk adjustment. An odds ratio of 0.85 translates to roughly a 15 percent lower odds of transport for this group compared with their White counterparts, once all measured sources of confounding have been removed from the equation. In other words, this was not a difference explained by where these infants happened to be born, by the capabilities of their birth hospitals, by their gestational age or birth weight, or by the illness severity recorded in their charts. Something in the system itself appeared to be operating differently for these infants.</p>
<p>Equally instructive was what the analysis revealed about the strongest predictors of transport overall. Hospital-level factors, most notably the level of the neonatal intensive care unit at the birth hospital, were the variables most significantly associated with whether a transport occurred. This makes biological and organizational sense. An infant born at a community hospital without a Level III or Level IV neonatal intensive care unit is far more likely to require transfer to a regional center than an infant born already inside such a center. This mechanism is the entire logic of perinatal regionalization, the organized system through which states route high-risk mothers and infants toward hospitals with the resources to care for them. Studies stretching back to the 1980s have consistently demonstrated that very low birth weight infants delivered at appropriate levels of care experience lower mortality, and meta-analyses have confirmed the survival advantage of regionalized systems.</p>
<p>Yet the system does not always function as designed. Prior work by some of the same investigators has documented the phenomenon of deregionalization, in which an increasing share of very low birth weight infants are born at hospitals that lack the highest levels of neonatal capability, eroding the protective effect of the regional model. Other research from the California Perinatal Quality Care Collaborative has shown that racial and ethnic disparities extend deep into the quality of care itself, with infants from minoritized groups receiving care in neonatal intensive care units that systematically deliver lower-quality services, a pattern described as racial segregation and inequality within the neonatal intensive care landscape. The new transport findings fit into this larger mosaic, suggesting that inequities are not confined to what happens inside intensive care units but also shape the very pathways by which infants move between them.</p>
<p>The authors&#8217; interpretation of their findings is deliberately measured. They conclude that the differing likelihood of acute transport by race and ethnicity may reflect underlying inequities and implicit biases in the system of care. This framing is important because it locates the problem not in the decisions of any single clinician but in the accumulated, often invisible patterns of how referrals are initiated, how transport teams are dispatched, and how risk is assessed across different patient populations. Implicit bias in clinical decision-making is well documented across medicine, and neonatal transport involves rapid judgments under time pressure, exactly the conditions in which unexamined assumptions about patients are most likely to influence behavior. Whether the lower transport rate among AANHPI infants reflects under-triage, differences in referral relationships, communication barriers, or other mechanisms is a question the study was not designed to answer, but the statistically robust association demands investigation.</p>
<p>The implications for policy and practice are concrete. Because hospital-level factors dominate the transport equation, strengthening adherence to regionalization principles, ensuring that high-risk deliveries occur at appropriately equipped hospitals, and auditing transport decisions for racial and ethnic equity are all actionable levers. The study also underscores the value of the granular data infrastructure maintained by quality collaboratives, which made it possible to detect a disparity that would be invisible in national aggregates. For the AANHPI category in particular, the finding adds urgency to calls for disaggregated data, since this grouping bundles together populations with widely divergent risk profiles and outcomes. What the study ultimately delivers is a measurable signal that the ambulance is not the same ambulance for every baby, a finding that should resonate far beyond California and into every perinatal system that claims, as its founding promise, that the sickest infants will reach the highest level of care regardless of who they are.</p>
<p>The dataset underpinning the analysis came from the California Perinatal Quality Care Collaborative, a statewide initiative that collects detailed clinical information from neonatal intensive care units across California. Because these records are gathered under a data use agreement rather than released publicly, the authors note that the underlying data are available only upon reasonable request with the collaborative&#8217;s permission. This governance model is common among quality collaboratives, which balance the research value of granular clinical data against the privacy protections owed to patients and participating hospitals.</p>
<p>The study&#8217;s statistical approach deserves emphasis. By adjusting simultaneously for network, hospital, maternal, and infant characteristics, the investigators sought to ensure that the observed difference in transport odds was not an artifact of clustering, since infants born at the same hospital share equipment, staffing, and referral practices. This kind of multilevel risk adjustment is essential in perinatal research, where the hospital an infant is born in is itself shaped by maternal residence, insurance, and patterns of segregation that precede any clinical decision.</p>
<p>The findings also connect to a long line of evidence on neonatal outcomes by race and ethnicity. Prior studies have documented disparities in very preterm neonatal morbidities, differences in mortality among very preterm infants across hospitals in large cities, and variation in outcomes for infants born at less than 30 weeks of gestation over time. Systematic reviews of neonatal intensive care have catalogued racial and ethnic differences spanning access, quality, and outcomes, indicating that no single point in the care pathway is immune.</p>
<p>Acute transport occupies a distinctive position in this pathway because it sits at the junction of multiple institutions. A transport decision requires coordination between the referring hospital, the transport team, and the receiving center, each with its own protocols and thresholds for escalation. Disparities arising at this junction may be harder to detect than disparities within a single unit, which makes the statistical signal reported here particularly valuable for quality improvement efforts aimed at ensuring equitable access to regionalized care.</p>
<p><strong>Subject of Research:</strong> Racial and ethnic disparities in acute inter-hospital neonatal transport among very low birth weight infants in California</p>
<p><strong>Article Title:</strong> Differential transport patterns by race and ethnicity in very low birth weight infants</p>
<p><strong>Article References:</strong> Kunz, S. N., Zitnik, M., Helkey, D., Razdan, S., Gould, J. B., Profit, J., &amp; Zupancic, J. A. F. (2026). Differential transport patterns by race and ethnicity in very low birth weight infants. <em>Journal of Perinatology</em>. <a href="https://doi.org/10.1038/s41372-026-02896-3" rel="noopener noreferrer">https://doi.org/10.1038/s41372-026-02896-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41372-026-02896-3" rel="noopener noreferrer">10.1038/s41372-026-02896-3</a></p>
<p><strong>Keywords:</strong> very low birth weight, neonatal transport, race and ethnicity, health disparities, neonatal intensive care, perinatal regionalization, odds ratio, California, Journal of Perinatology, AANHPI, NICU level, implicit bias</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">193138</post-id>	</item>
	</channel>
</rss>
