<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>reward-driven modulation of sensory encoding &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/reward-driven-modulation-of-sensory-encoding/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 14:25:13 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>reward-driven modulation of sensory encoding &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Feeling Right Is Rewarding: How Inner Confidence Steals Your Brain&#8217;s Visual Resources</title>
		<link>https://scienmag.com/feeling-right-is-rewarding-how-inner-confidence-steals-your-brains-visual-resources/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 14:25:13 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[attentional biases without external incentives]]></category>
		<category><![CDATA[brain mechanisms of self-confidence]]></category>
		<category><![CDATA[computational modelling]]></category>
		<category><![CDATA[confidence]]></category>
		<category><![CDATA[impact of internal rewards on memory and perception]]></category>
		<category><![CDATA[influence of internal feedback on visual processing]]></category>
		<category><![CDATA[internal reward signals]]></category>
		<category><![CDATA[intrinsic reward]]></category>
		<category><![CDATA[metacognition]]></category>
		<category><![CDATA[motivation and visual attention]]></category>
		<category><![CDATA[neural basis of feeling right and confidence]]></category>
		<category><![CDATA[neural resource allocation]]></category>
		<category><![CDATA[perceptual decision-making]]></category>
		<category><![CDATA[population coding]]></category>
		<category><![CDATA[psychophysics]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[reinforcement learning in the brain]]></category>
		<category><![CDATA[reward learning]]></category>
		<category><![CDATA[reward-driven modulation of sensory encoding]]></category>
		<category><![CDATA[subjective confidence and perception]]></category>
		<category><![CDATA[visual attention]]></category>
		<category><![CDATA[working memory]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=248134</guid>

					<description><![CDATA[New research shows that the brain treats internal confidence as a reward, using reinforcement learning to divert limited visual processing resources toward stimuli that make us feel accurate, often at the cost of overall performance.]]></description>
										<content:encoded><![CDATA[<p>Every moment, the brain confronts far more visual information than it can process. With a limited pool of neural resources at its disposal, it must decide which sights deserve rich, detailed encoding and which can be relegated to the fuzzy periphery of awareness. A large body of research has shown that external incentives, such as money or food, bias this allocation: stimuli that predict tangible rewards are processed preferentially, remembered more precisely and even capture attention involuntarily long after the reward contingencies have disappeared. But much of everyday life offers no external payoff at all. When you rehearse a phone number or try to catch a fleeting expression on someone&#8217;s face, the only feedback you receive comes from within. A new study published in Nature Human Behaviour by Ivan Tomić, Rodrigo Raimundo and Paul M. Bays of the University of Cambridge argues that these internal signals, particularly the subjective sense of being accurate, act as genuine rewards that are learned about through the same reinforcement machinery that governs responses to money, and that they can redirect the brain&#8217;s visual resources in ways that are simultaneously adaptive and self-defeating.</p>
<p>The team&#8217;s experimental platform was deceptively simple. Observers viewed two coloured patches of moving dots, and after a brief delay a colour cue told them which of the two motion directions they had to reproduce by rotating an arrow. Because responses were continuous rather than categorical, the precision of each reproduction could be measured to a fraction of a degree, and the distribution of errors across hundreds of trials could be fed into a well-established mathematical framework known as the neural resource model. In that model, a visual stimulus is encoded by a population of idealized direction-selective neurons with bell-shaped tuning curves, and the total activity of the population is held fixed by divisive normalization, a canonical neural computation. Crucially, the model allows a gain parameter to be fitted freely, revealing how the fixed pool of spiking activity was divided between the two competing stimuli. A narrower hill of activity means more resources and more precise reports; a broader, heavier-tailed error distribution betrays a starved representation.</p>
<p>In the first experiment, the researchers replicated the classical external reward effect with a twist. Accurate recall of one colour earned observers fifteen points, while accurate recall of the other earned only five, and the accumulated points translated into a bonus payment. Reproduction errors were reliably smaller for the high-reward colour, and the fitted gain parameter showed that roughly sixty-five per cent of the neural resources were channelled towards the privileged stimulus. Yet when the team computed the allocation that would truly maximize the expected point total, a strategy that would have demanded starving the low-reward item far more aggressively, the observers fell short. They distributed resources more equally than a reward-maximizing policy required, earned fewer points per trial than the optimum predicted, and showed essentially no correlation between their individual allocations and their individual optima. Something beyond the pursuit of points, the authors reasoned, was holding the allocation in check, and the most obvious candidate was that being accurate across the board is itself satisfying.</p>
<p>To test that idea directly, the researchers turned to deception. In a second series of experiments, they manipulated the feedback displayed after each response, silently magnifying the apparent error for stimuli of one colour and shrinking it for the other. Observers had no idea the feedback was distorted; none suspected the manipulation when debriefed, yet eighty-four per cent came to describe the error-magnified colour as harder to remember. The consequences were striking. Reproductions of the error-minified stimulus, the one that made observers feel competent, became genuinely more precise, even though the manipulation had changed nothing about the stimulus itself or the observers&#8217; true accuracy. A control experiment in which the two stimuli appeared sequentially, eliminating competition at the moment of encoding, abolished the effect entirely, pinpointing attentional competition during perception rather than memory maintenance as the locus of the bias. A further refinement showed that the asymmetry was driven by the pleasant side of the manipulation: shrinking errors boosted precision, but inflating them did not reliably impair it.</p>
<p>The third series of experiments replaced fabricated feedback with honest variation in stimulus quality. On most trials, one colour carried dots moving with eighty-five per cent coherence, an easy, low-noise motion signal, while the other carried only forty-five per cent coherence, a noisy, difficult one. On interleaved probe trials, both stimuli were shown at the same intermediate coherence, so any difference in precision had to reflect learned associations rather than the stimuli themselves. Once again, the colour previously paired with the easy, confidence-inspiring signal was reproduced more accurately even when it was objectively no easier than its rival. The fitted resource model estimated that observers allocated nearly twice as much neural activity to the high-coherence colour, and the effect survived strict fixation control in a laboratory version of the task with eye tracking, ruling out overt shifts of gaze as an explanation. When stimuli were presented sequentially, the advantage dissolved, echoing the conclusion that the bias operates through attentional competition at encoding.</p>
<p>Perhaps the most provocative finding is that these internally driven allocations were not merely suboptimal; they were actively counterproductive. The researchers simulated the allocation policy that would minimize overall response error in the difficulty experiments and found it was essentially equal division of resources, because pouring resources into a noisy stimulus cannot rescue an estimate drowning in perceptual noise without bankrupting the good one. Observers did the opposite, favouring the already-easy item, and their overall error exceeded the optimum as a result. In the feedback experiment, the error-minimizing strategy would have required shifting twice the resources towards the magnified, difficult item; instead observers favoured the minified one, producing feedback error significantly worse than optimal. The brain, it seems, was not solving the task it was given. It was chasing the feeling of getting things right.</p>
<p>To explain this seemingly irrational behaviour mechanistically, the authors augmented the neural resource model with a simple reinforcement learning rule. On each trial, three candidate reward signals are combined into a composite: the external points earned, an exponential transform of the feedback error, and an estimate of internal confidence derived from the model itself. Because spiking activity is generated by a Poisson process, the precision of the decoded estimate fluctuates from trial to trial, and the width of the likelihood function over stimulus values provides a natural, report-free measure of how confident the observer&#8217;s brain is on that particular trial. This composite reward updates a running value associated with each stimulus colour, with a leak parameter determining how quickly old value decays, and the accumulated value is mapped through an exponential function onto the neural gain that sets resource allocation on the next trial. Rewards, whether external or internal, thus become attached to the identifying features of objects, and those features inherit the priority of the rewards they have predicted.</p>
<p>The model&#8217;s vindication came from a comparison of two independent estimates. The freely fitted gain parameter, extracted from error distributions alone with no knowledge of rewards, correlated almost perfectly with the allocation trajectory computed purely from the history of accumulated rewards, with correlation coefficients approaching 0.98 in the external reward experiment and remaining above 0.83 in the confidence-driven experiments. Within individual observers, trials on which the model predicted below-median allocation to the probed stimulus produced reliably larger errors than trials with above-median allocation, across all four experiments. In other words, a learning rule that treats the private sensation of confidence as a reward, and nothing more, is sufficient to reconstruct how the brain divided its visual budget, trial by trial, without any explicit strategy or awareness.</p>
<p>The findings resonate with a growing literature suggesting that the striatum and its dopaminergic inputs, long celebrated for encoding prediction errors about money and juice, also encode prediction errors about confidence, and that perceptual learning can proceed without any external feedback at all, driven by internal confidence signals. They also connect to the economics of mental effort: tasks perceived as difficult are avoided and discounted, and the present results suggest that difficulty discounts the subjective value of a stimulus in much the same way, diverting resources towards whatever feels achievable. The authors are careful to note the limits of the work. Confidence was never measured directly but manipulated through validated proxies, and the tasks were demanding enough that even easy stimuli required genuine effort, leaving open whether extremely trivial stimuli might trigger the opposite, curiosity-driven reallocation towards harder items. Still, the broader implication is hard to escape. The allocation of perception is not a dispassionate optimization of accuracy but a market shaped by reward, and some of the most powerful currency in that market is manufactured entirely inside our own heads, rewarding us for being right and quietly reshaping what we are able to see.</p>
<p><strong>Subject of Research:</strong> How intrinsic reward signals such as confidence guide the allocation of limited visual processing resources through reinforcement learning</p>
<p><strong>Article Title:</strong> Intrinsic rewards guide visual resource allocation via reinforcement learning</p>
<p><strong>Article References:</strong> Intrinsic rewards guide visual resource allocation via reinforcement learning. (n.d.). <a href="https://doi.org/10.1038/s41562-026-02573-7" rel="noopener noreferrer">https://doi.org/10.1038/s41562-026-02573-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41562-026-02573-7" rel="noopener noreferrer">10.1038/s41562-026-02573-7</a></p>
<p><strong>Keywords:</strong> reinforcement learning, intrinsic reward, visual attention, working memory, confidence, neural resource allocation, population coding, perceptual decision-making, metacognition, reward learning, psychophysics, computational modelling</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">248134</post-id>	</item>
	</channel>
</rss>
