<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>RLWM paradigm &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/rlwm-paradigm/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 23:58:02 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>RLWM paradigm &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Response Times Expose Hidden Brain Mechanisms Behind Learning and Mental Illness</title>
		<link>https://scienmag.com/response-times-expose-hidden-brain-mechanisms-behind-learning-and-mental-illness/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 23:58:02 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[cognitive control]]></category>
		<category><![CDATA[computational modeling]]></category>
		<category><![CDATA[computational psychiatry]]></category>
		<category><![CDATA[decision-making]]></category>
		<category><![CDATA[especially in schizophrenia.]]></category>
		<category><![CDATA[evidence accumulation]]></category>
		<category><![CDATA[hierarchical Bayesian estimation]]></category>
		<category><![CDATA[involving gradual learning over time. Incorporating response times into computational models reveals hidden brain mechanisms underlying learning and mental illness]]></category>
		<category><![CDATA[parameter recovery]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[response times]]></category>
		<category><![CDATA[RLWM paradigm]]></category>
		<category><![CDATA[schizophrenia]]></category>
		<category><![CDATA[working memory]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=250561</guid>

					<description><![CDATA[By jointly modeling response times alongside choices, researchers disentangled working memory, cognitive control and reinforcement learning, and uncovered schizophrenia deficits that choice-only models had masked.]]></description>
										<content:encoded><![CDATA[<p>Every decision you make carries two kinds of information: what you chose, and how long it took you to choose it. For decades, computational models of learning have focused almost exclusively on the first signal, treating choices as the currency of the mind. A new study published in PLOS Computational Biology argues that this habit has been quietly distorting science. By adding response times to the standard toolkit of reinforcement learning models, a team led by Michael J. Frank of Brown University, working with collaborators across one of the largest schizophrenia research consortia, shows that the timing of decisions is not a footnote to cognition. It is the key that unlocks which mental machinery actually produced the behavior, and it exposes clinical deficits that choice-only models had rendered invisible.</p>
<p>The problem the researchers tackled is one of the most stubborn ambiguities in cognitive science. When a person learns which option in a task pays off most often, at least two distinct systems could be driving the performance. The first is working memory, the capacity-limited mental workspace that holds recent outcomes in mind and allows rapid, flexible adaptation after just one or two experiences. The second is the slower, incremental process of reinforcement learning, in which prediction errors gradually nudge the values of options up or down over many trials, supporting robust long-term retention. A well-known experimental paradigm called RLWM was designed to pull these systems apart by manipulating how much information participants must hold in mind during learning. Yet even with clever task design, a fundamental ambiguity remains: any single choice can often be explained equally well by a fast working memory strategy or by a slow reinforcement learning process with the right parameters.</p>
<p>The consequences of that ambiguity are not merely philosophical. When models are fit only to choices, the researchers show, the fitting procedure tends to compensate for the missing timing information by inflating reinforcement learning rates. In other words, the model concludes that people learned faster from feedback than they actually did, because it has no other way to explain the pattern of responses. This parameter inflation is not a small technical artifact. It changes the scientific interpretation of the data, potentially misattributing to dopamine-like learning processes what is actually the signature of working memory or cognitive control. The team demonstrated that models fit to choices alone failed a critical out-of-sample test: they could not accurately predict participants&#8217; choices in a held-out test phase after learning had ended.</p>
<p>The solution was to jointly model decision dynamics. Rather than fitting only which button a participant pressed, the new approach accounts for the full distribution of response times alongside the choices, using hierarchical Bayesian parameter estimation to pool information across participants while respecting individual differences. The mathematical framework builds on evidence accumulation models of decision making, in which noisy evidence for each option accumulates over time until it crosses a decision boundary, triggering a response. In this architecture, the learning processes set the quality of the evidence, while the decision process determines how that evidence is converted into a choice and a latency. The two layers constrain each other: a model that gets the learning wrong will also get the timing wrong, and the joint fit punishes such errors.</p>
<p>The payoff was substantial. Despite the added complexity, models that incorporated decision dynamics showed markedly improved parameter recovery, meaning that when the researchers simulated data with known parameters, the fitted models could accurately recover the true values of learning rates, working memory capacities, and decision thresholds. The joint models also produced accurate out-of-sample predictions for choices in the held-out test phase, passing the test that choice-only models failed. This matters because out-of-sample prediction is one of the strictest standards a computational model can meet: it demonstrates that the model has captured a generalizable mechanism rather than merely memorizing the idiosyncrasies of a particular dataset. For a field increasingly reliant on model-based measures of cognition, the difference between a model that recovers parameters faithfully and one that inflates them is the difference between measuring the mind and measuring the model&#8217;s own assumptions.</p>
<p>Beyond correcting old estimates, the joint modeling approach revealed something genuinely new. The analysis uncovered a previously unidentified neurocognitive process: participants proactively widen their decision boundaries when working memory load increases. In evidence accumulation terms, a wider boundary means the accumulator must gather more evidence before committing to a response, which increases response caution and slows decisions. This is a form of proactive cognitive control, an anticipatory adjustment of the decision threshold rather than a reactive correction after an error. Participants appear to sense when a task will strain their working memory and strategically slow down to protect accuracy. No choice-only model could ever detect this, because a wider boundary leaves the choices unchanged; it is written exclusively in the timing of responses. The finding illustrates the study&#8217;s central thesis in miniature: a whole dimension of cognitive strategy was hiding in plain sight, invisible to anyone who ignored the clock.</p>
<p>The clinical implications emerged when the team applied the joint model to patients with schizophrenia, drawing on data collected through a large multi-site research consortium. Previous choice-only analyses of similar tasks had established that patients show working memory deficits, and the new model replicated that finding. But the richer model went further, revealing two additional deficits that had been masked before. Patients showed impairments in proactive control mechanisms, meaning they failed to adaptively widen their decision boundaries under increased working memory load in the way healthy participants did. They also exhibited slowed incremental reinforcement learning, a deficit in the gradual, feedback-driven value updating system itself. Each of these three impairments, in working memory, proactive control, and reinforcement learning, points to a different underlying neural circuit and a different potential treatment target, yet choice-only models had collapsed them into a blur.</p>
<p>The masking effect deserves particular attention, because it illustrates how methodological choices can shape clinical conclusions. A model that attributes all behavioral differences to a single learning parameter will miss deficits that express themselves only in response times or in the coupling between control processes and task demands. For psychiatry, where computational psychiatry aims to derive mechanistic, individual-level measures of psychopathology, this is a cautionary tale and a promising roadmap. The authors describe their approach as enabling accurate, model-based, mechanism-oriented computational phenotyping. Instead of asking whether a patient group differs from controls on some aggregate behavioral score, the joint model asks which specific computational mechanisms differ, quantified at the level of individual participants and validated by out-of-sample prediction. Such phenotypes could eventually serve as intermediate markers in treatment trials, sensitive to interventions that act on one cognitive system but not another.</p>
<p>The study also carries a broader lesson for cognitive science at large. Response times have been central to mathematical psychology since the earliest diffusion models of perceptual decision making, yet much of the modern reinforcement learning literature drifted toward choice-only fitting, partly for simplicity. This work demonstrates that the two traditions are stronger in combination than either is alone. Learning models supply the changing evidence that drives decisions; decision models supply the transform that converts evidence into behavior; and hierarchical Bayesian estimation ties the whole hierarchy together across individuals and groups. As the authors show, the extra complexity is not a luxury but a necessity: it is precisely what allows the model to distinguish working memory from reinforcement learning, strategy from capacity, and caution from competence. In a task where a single button press could have sprung from either of two minds within the same skull, the milliseconds turn out to be the tell.</p>
<p>For researchers designing the next generation of learning experiments, the message is concrete. Record response times, model them jointly with choices, and validate models on held-out data before trusting their parameters. For clinicians, the message is that the behavioral signatures of schizophrenia are richer and more specific than aggregate accuracy scores suggest, spanning deficits in memory, control, and learning that can now be measured separately within a single task. And for anyone who has wondered what a split-second hesitation reveals about the mind, the answer from this study is: potentially everything. The clock, long treated as a side channel of behavior, is now firmly established as a main line of evidence about how we learn, decide, and, when illness strikes, how those processes come apart.</p>
<p><strong>Subject of Research:</strong> Joint computational modeling of choice and response time dynamics to separate working memory, cognitive control and reinforcement learning, and to identify cognitive deficits in schizophrenia.</p>
<p><strong>Article Title:</strong> Modeling decision dynamics disentangles working memory, cognitive control and reinforcement learning and reveals clinical differences</p>
<p><strong>Article References:</strong> Bera, K., Fengler, A., Boudewyn, M. A., Carter, C. S., Erickson, M. A., Gold, J. M., Luck, S. J., Ragland, J. D., Yonelinas, A. P., MacDonald III, A. W., Barch, D. M., &amp; Frank, M. J. (2026). Modeling decision dynamics disentangles working memory, cognitive control and reinforcement learning and reveals clinical differences. <em>PLOS Computational Biology, 22</em>(9), e1014796. <a href="https://doi.org/10.1371/journal.pcbi.1014796" rel="noopener noreferrer">https://doi.org/10.1371/journal.pcbi.1014796</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1371/journal.pcbi.1014796" rel="noopener noreferrer">10.1371/journal.pcbi.1014796</a></p>
<p><strong>Keywords:</strong> reinforcement learning, working memory, cognitive control, decision making, response times, computational modeling, hierarchical Bayesian estimation, schizophrenia, computational psychiatry, parameter recovery, evidence accumulation, RLWM paradigm</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">250561</post-id>	</item>
	</channel>
</rss>
