<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>reward prediction error &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/reward-prediction-error/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 13:22:06 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>reward prediction error &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Brainwaves That Say &#8216;Not What I Meant&#8217;: AI Learns to Read Unspoken Corrections</title>
		<link>https://scienmag.com/brainwaves-that-say-not-what-i-meant-ai-learns-to-read-unspoken-corrections/</link>
		
		<dc:creator><![CDATA[Cassandra Pierce]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 13:22:06 +0000</pubDate>
				<category><![CDATA[Science Education]]></category>
		<category><![CDATA[AI behavior adjustment through brain signals]]></category>
		<category><![CDATA[AI understanding unspoken human cues]]></category>
		<category><![CDATA[Brain-Computer Interface]]></category>
		<category><![CDATA[brain-computer interface for AI]]></category>
		<category><![CDATA[Brainwave-based AI correction]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[EEG]]></category>
		<category><![CDATA[goal–action ambiguity resolution in AI systems]]></category>
		<category><![CDATA[Human-AI Collaboration.]]></category>
		<category><![CDATA[IEEE Transactions on Cybernetics]]></category>
		<category><![CDATA[improving human–AI collaboration]]></category>
		<category><![CDATA[innovations in AI alignment and error correction]]></category>
		<category><![CDATA[intent inference]]></category>
		<category><![CDATA[KAIST]]></category>
		<category><![CDATA[machine learning from neural feedback]]></category>
		<category><![CDATA[Microsoft Research Asia]]></category>
		<category><![CDATA[Neural Value Alignment]]></category>
		<category><![CDATA[Neural Value Alignment technology]]></category>
		<category><![CDATA[physical AI]]></category>
		<category><![CDATA[real-time human intent detection]]></category>
		<category><![CDATA[reward prediction error]]></category>
		<category><![CDATA[state prediction error]]></category>
		<category><![CDATA[subconscious error detection in AI training]]></category>
		<category><![CDATA[unspoken correction in human–AI interaction]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=247946</guid>

					<description><![CDATA[KAIST and Microsoft Research Asia have developed Neural Value Alignment, a brain–computer interface technology that decodes EEG signals to detect when AI misunderstands a person's goal or method and corrects its behavior in real time.]]></description>
										<content:encoded><![CDATA[<p>Anyone who has watched a robot assistant perform exactly the wrong task while technically following instructions knows the quiet frustration of the unspoken correction — the moment the brain registers &#8220;that&#8217;s not what I meant&#8221; seconds before the mouth catches up. A research team at the Korea Advanced Institute of Science and Technology (KAIST), working with Microsoft Research Asia (MSRA), has now built a technology designed to catch precisely that moment. The system, called Neural Value Alignment (NVA), reads human brainwaves in real time to detect when an AI&#8217;s behavior diverges from a person&#8217;s actual intent, and then feeds that information back to the AI so it can adjust itself without ever being told what went wrong. The work, published online in August 2026 in IEEE Transactions on Cybernetics, points toward a fundamentally different model of human–AI collaboration: one in which machines no longer wait for explicit commands but instead sense, through the brain&#8217;s own error signals, when they have misunderstood the person they are meant to serve.</p>
<p>The core problem the researchers set out to solve is what they describe as two-sided &#8220;goal–action ambiguity.&#8221; Conventional AI systems infer human intent from externally observable information — speech, gestures, actions — but the same observable behavior can reflect very different goals, and the same goal can be pursued through very different actions. When a person picks up a cup, an observer cannot know whether they intend to drink from it or to hand it to someone else. Conversely, if the goal is simply to quench a thirst, that goal could be satisfied by reaching for the cup, grabbing a water bottle, or asking another person to bring a drink. Because of this ambiguity, when an AI misreads a human&#8217;s intent, the user is typically forced to correct it through additional commands or demonstrations, a slow and often frustrating loop that undermines the promise of natural collaboration between people and machines.</p>
<p>The KAIST and MSRA team&#8217;s insight was to bypass observable behavior entirely and tap into the brain&#8217;s own error-detection machinery. When a person encounters an unexpected situation, the brain produces rapid, unconscious &#8220;prediction error&#8221; signals — an instantaneous internal alarm that something does not match expectations. In the context of human–AI interaction, these signals amount to the brain&#8217;s immediate &#8220;that&#8217;s not what I meant&#8221; response when an AI performs the wrong action or pursues the wrong goal. Crucially, the researchers found that these signals are not monolithic. They distinguished two distinct types of brain responses, each carrying different information about what exactly went wrong in the collaboration.</p>
<p>The first type is reward prediction error, or RPE, which appears when the AI misunderstands the person&#8217;s ultimate goal — when the machine has grasped the wrong endpoint entirely. The second is state prediction error, or SPE, which appears when the goal is correct but the process or method of action differs from what the person expected. This distinction matters enormously for how an AI should respond. If the brain is broadcasting an SPE signal, the appropriate correction is to keep the goal and change the strategy. If it is broadcasting an RPE signal, the AI must discard its assumption about what the person wants and search again for the true objective. By separating these two error classes, the technology gives AI systems a structured way to diagnose the nature of their own misunderstanding rather than simply registering that something is amiss.</p>
<p>To capture these signals, the team measured real-time electroencephalography (EEG) from people as they observed AI systems performing tasks. The recordings revealed that the brain produced measurably different patterns depending on whether the AI had misunderstood the goal itself or had chosen the wrong method while pursuing the correct goal. The researchers also identified distinctive brainwave signatures that emerged when both types of errors occurred simultaneously — a doubly ambiguous situation in which neither the destination nor the route matched the person&#8217;s expectations. These findings establish, at the level of neural measurement, that the human brain encodes goal-level and method-level disagreement with an AI as separable, decodable phenomena.</p>
<p>The decoding itself was accomplished with deep learning. By training neural networks on the recorded EEG signals, the team developed a system that can determine, from brain activity alone, how a person is interpreting the AI&#8217;s behavior at any given moment. In practical terms, even without a person saying &#8220;that&#8217;s wrong,&#8221; the AI can recognize whether the human brain is signaling that &#8220;the goal is wrong&#8221; or &#8220;the method is wrong.&#8221; The team then wrapped this decoding capability into a Neural Value Alignment–based human–AI synergy algorithm, which feeds the decoded brain signals back to the AI in real time. When the AI detects an SPE signal, it interprets the situation as &#8220;the desired goal is correct, but the method is wrong&#8221; and adjusts its action strategy accordingly. When it detects an RPE signal, it understands that &#8220;the goal itself was misunderstood&#8221; and re-searches for what the person truly intended. In both cases, the correction loop closes without a single spoken word.</p>
<p>Simulation results offered early evidence that the approach works under realistic stress. The proposed method adapted more quickly than existing approaches even in uncertain conditions, such as when a person&#8217;s goal changed suddenly mid-task or when some of the human neural feedback was missing — a scenario that reflects the noisy, incomplete data typical of real-world brain–computer interfaces. This robustness matters because laboratory EEG is far cleaner than what wearable sensors will deliver in homes, factories, or vehicles. The researchers emphasize that the significance of the work lies in demonstrating that AI systems can correct themselves by reading a person&#8217;s unconscious &#8220;that&#8217;s not what I meant&#8221; brain response, without requiring the user to repeatedly say &#8220;do it this way&#8221; or &#8220;that&#8217;s not right.&#8221;</p>
<p>The potential applications stretch across the emerging landscape of embodied and interactive AI. As the technology matures, the researchers say, it could be applied to physical AI robots in homes and industrial settings, allowing them to understand user intent more naturally and adjust their actions accordingly. Autonomous vehicles could quickly reflect driver judgment in situations where a split-second correction matters. Medical and rehabilitation robots could serve patients who have difficulty speaking or moving, for whom conventional command-based interaction is impractical and whose needs are most acute. Educational AI systems could adapt to a student&#8217;s cognitive state in real time, sensing confusion or disagreement that the student never articulates. In each case, the common thread is a tighter, faster coupling between human judgment and machine behavior than explicit interfaces allow.</p>
<p>Professor Sang Wan Lee of KAIST&#8217;s Department of Brain and Cognitive Sciences, who directs the Center for Neuroscience-Inspired Artificial Intelligence and led the international collaboration, framed the work as a shift in the source of intent information. &#8220;This research is meaningful because it shows that AI can move beyond inferring human intent only from visible behavioral outcomes and instead directly use cognitive signals generated in the brain during collaboration with AI,&#8221; he said. He added that the technology can be expanded to a wide range of fields where human judgment and AI behavior must be closely connected, including physical AI, brain–computer interfaces, autonomous driving, precision personalized education, medical robotics, and human–computer interaction. From MSRA, Miran Lee, Director of the Microsoft Research Accelerator, described the achievement as the result of the ongoing international collaboration between KAIST and Microsoft Research Asia and expressed the intention to continue the partnership to develop world-class BCI technologies that enable humans and AI to communicate and collaborate more naturally.</p>
<p>The study&#8217;s first author is Xin Xu, a Ph.D. student in KAIST&#8217;s Department of Brain and Cognitive Sciences; researchers from Microsoft Research Asia, including Yansen Wang, Dongqi Han, and Dongsheng Li, also participated. The paper, titled &#8220;Neural Value Alignment: Human–AI Collaboration Under Goal–Action Ambiguity,&#8221; appeared in IEEE Transactions on Cybernetics under DOI 10.1109/TCYB.2026.3722605. The work builds on a broader research program: a companion KAIST–MSRA study on helping AI adapt rapidly to continuously changing environments, led by first author Niklas Koeppe, was presented in June at ICML 2026 under the title &#8220;Mitigating Plasticity Loss through Architectural Design in Continual Learning.&#8221; The research was supported by the Institute of Information &amp; Communications Technology Planning &amp; Evaluation (IITP), funded by the Ministry of Science and ICT. Taken together, the results sketch a future in which the most important interface between humans and machines is not the keyboard, the touchscreen, or even the voice assistant, but the silent, instantaneous feedback of the brain itself — a future in which AI finally hears the corrections we never say out loud.</p>
<p><strong>Subject of Research:</strong> Brainwave-based AI technology that detects human–AI cognitive mismatch and realigns AI behavior with human intent</p>
<p><strong>Article Title:</strong> KAIST develops AI that can sense “that’s not what I meant” without being told</p>
<p><strong>Article References:</strong> KAIST develops AI that can sense “that’s not what I meant” without being told. (n.d.). <a href="https://www.eurekalert.org/news-releases/1143282" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> KAIST, Microsoft Research Asia, Neural Value Alignment, brain–computer interface, EEG, reward prediction error, state prediction error, human–AI collaboration, deep learning, intent inference, IEEE Transactions on Cybernetics, physical AI</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">247946</post-id>	</item>
		<item>
		<title>Virtual Reality Could Rewire Reward Circuits to Treat Anhedonia</title>
		<link>https://scienmag.com/virtual-reality-could-rewire-reward-circuits-to-treat-anhedonia/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:54:20 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[anhedonia]]></category>
		<category><![CDATA[behavioral activation]]></category>
		<category><![CDATA[clinical applications of virtual reality in mental health]]></category>
		<category><![CDATA[Depression]]></category>
		<category><![CDATA[dopamine]]></category>
		<category><![CDATA[immersive technology]]></category>
		<category><![CDATA[immersive virtual reality in psychiatry]]></category>
		<category><![CDATA[innovative treatments for anhedonia]]></category>
		<category><![CDATA[Mental health]]></category>
		<category><![CDATA[neuroscience of virtual reality therapy]]></category>
		<category><![CDATA[personalized VR experiences for reward system]]></category>
		<category><![CDATA[positive affect treatment]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[reward circuit rewiring with VR]]></category>
		<category><![CDATA[reward prediction error]]></category>
		<category><![CDATA[reward sensitivity]]></category>
		<category><![CDATA[targeting reward sensitivity with virtual reality]]></category>
		<category><![CDATA[ventral striatum]]></category>
		<category><![CDATA[virtual reality]]></category>
		<category><![CDATA[virtual reality for anhedonia treatment]]></category>
		<category><![CDATA[virtual reality for depression and anxiety]]></category>
		<category><![CDATA[virtual reality in schizophrenia treatment]]></category>
		<category><![CDATA[virtual reality-based mental health interventions]]></category>
		<category><![CDATA[VR for PTSD and substance use disorders]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201944</guid>

					<description><![CDATA[Researchers propose that interactive virtual reality experiences engineered to generate positive reward prediction errors could retrain the brain's reward system and treat anhedonia.]]></description>
										<content:encoded><![CDATA[<p>Anhedonia, the persistent inability to feel pleasure or interest in activities that once brought joy, is one of the most stubborn and disabling symptoms in psychiatry. It cuts across diagnoses, appearing in major depression, anxiety disorders, schizophrenia, PTSD and substance use conditions, and it often lingers even when other symptoms respond to treatment. Now, a new perspective published in Nature Mental Health argues that immersive, interactive virtual reality could become a powerful clinical tool for reigniting the brain&#8217;s reward machinery, and it lays out a concrete, neuroscience-driven blueprint for how to design such experiences.</p>
<p>The perspective, authored by Mehmet Kosa of Marshall University, Nora M. Barnes-Horowitz of the University of Colorado Boulder, and Michelle G. Craske of the University of California, Los Angeles, builds on a growing body of evidence that treatments specifically targeting reward sensitivity can meaningfully reduce anhedonia. Rather than treating low mood as a single undifferentiated problem, the authors argue that clinicians should think of anhedonia as a disorder of the reward system, one that can be systematically probed and retrained. Virtual reality, they contend, offers an unprecedented medium for doing exactly that, because it can deliver experiences that are specific, fine-tuned, immersive and personalized in ways that traditional talk therapy or flat-screen interventions cannot match.</p>
<p>The theoretical foundation of the proposal rests on decades of work dissecting how the brain processes reward. Neuroscientists have long distinguished between separable components of reward: &#8216;liking&#8217;, the hedonic pleasure experienced in the moment; &#8216;wanting&#8217;, the motivational drive that pushes an organism to seek out rewards; and learning, the process by which the brain updates its expectations about what is worth pursuing. These components rely on partially distinct neural circuits, and they can be impaired independently of one another. A person with anhedonia may be able to enjoy a pleasurable experience once they are fully engaged in it, yet lack the motivation to begin it, or they may fail to learn from positive outcomes that would normally encourage repetition. This fractionation matters clinically, because a treatment that boosts one component may do nothing for another.</p>
<p>At the heart of the new argument is a deceptively simple computational concept: the reward prediction error. When an outcome turns out better than expected, dopamine neurons in the ventral striatum and midbrain fire in a characteristic burst, signaling a positive prediction error. This signal is not merely a marker of pleasure; it is the engine of reinforcement learning, teaching the brain which actions and contexts are worth repeating. Decades of research, from classic animal studies to modern human neuroimaging, have established that these prediction error signals drive learning, memory encoding, attention and motivation. Crucially, work by Wolfram Schultz and others has shown that the magnitude of the dopamine response scales with the discrepancy between expected and obtained reward, not with the reward itself.</p>
<p>What makes this relevant to anhedonia is a growing literature suggesting that prediction error signaling is blunted in people with the symptom. Studies of reinforcement learning in depression have found flattened responses to unexpected rewards, and related work on internet gaming disorder has documented dampened prediction error signals as well. Meanwhile, computational models of momentary well-being, developed by Robb Rutledge and colleagues, indicate that fleeting happiness depends more on recent prediction errors than on the absolute size of rewards received. In other words, the emotional lift we feel in daily life may be driven less by what we get and more by how much better it was than we anticipated. If anhedonic individuals experience smaller positive prediction errors, the world may feel flat not because nothing good happens, but because nothing good surprises them.</p>
<p>This is where virtual reality enters the picture. The authors propose that carefully engineered VR interactions could be designed specifically to generate positive reward prediction errors, delivering outcomes that reliably and pleasantly exceed expectations. Unlike real life, where reward statistics are messy and uncontrollable, a virtual environment gives designers precise control over timing, probability, magnitude and surprise. A user might reach for a glowing object expecting a modest sparkle and instead trigger a cascading burst of color, sound and narrative payoff. Each such moment, the argument goes, delivers a clean, well-timed dopaminergic teaching signal to a reward system that has grown underresponsive. Over repeated exposures, these signals could gradually recalibrate reward sensitivity, strengthening the learning and motivational components of reward processing that are most impaired in anhedonia.</p>
<p>The second pillar of the proposal is equally ambitious: rather than treating VR as a single undifferentiated pleasure machine, the authors argue that distinct activities within a virtual experience should be mapped onto the distinct phases of reward processing. Anticipatory reward, the &#8216;wanting&#8217; phase, could be targeted with activities that build expectation and motivation, such as quests, exploration and goal pursuit that require effortful engagement before any payoff arrives. Consummatory reward, the &#8216;liking&#8217; phase, could be targeted with immersive savoring experiences, rich sensory environments and moments designed purely for in-the-moment enjoyment. Reward learning, meanwhile, could be targeted with probabilistic tasks and feedback structures that require the user to discover, through trial and error, which choices lead to better outcomes. This phase-specific design philosophy mirrors the structure of Positive Affect Treatment, the neuroscience-informed psychotherapy developed by Craske and colleagues, which has shown in multiple randomized controlled trials that directly targeting reward sensitivity can reduce anhedonia, depression and anxiety symptoms.</p>
<p>The evidence base for this approach is still young but encouraging. A pilot study of VR-based reward training for anhedonia demonstrated feasibility, and randomized trials of virtual reality-enhanced behavioral activation for major depressive disorder have shown promising results. Meta-analyses of VR exposure therapy for anxiety disorders have established that immersive technologies can produce clinically meaningful change, and systematic reviews of positive mood induction confirm that virtual environments can reliably elicit positive emotions. Research on presence, the subjective sense of &#8216;being there&#8217; in a virtual world, suggests that immersion amplifies emotional responses, and studies of interactive media indicate that agency, the feeling that one&#8217;s actions matter, adds further emotional impact on top of immersion alone. Interactivity, the authors emphasize, is not a cosmetic feature: it is what allows the user to generate their own prediction errors through action, rather than passively receiving them.</p>
<p>The perspective also confronts the practical and ethical challenges honestly. VR sickness remains a barrier for some users, and validated questionnaires now exist to monitor it. Privacy and security concerns in immersive platforms are real, since these systems can collect intimate behavioral and physiological data. Ethical questions about children and adolescents, and about the risk that highly rewarding virtual experiences could shade into problematic use, particularly given what is known about gaming disorder, must be taken seriously. The authors call for adaptive designs that personalize difficulty and reward statistics to each user, rigorous measurement of presence and side effects, and careful attention to who benefits and who might be harmed.</p>
<p>If the vision succeeds, the implications could extend well beyond anhedonia. Objective measures of reward sensitivity are already being used to track treatment response, and neuroimaging work has linked ventral striatal reward activation to lasting improvements in life satisfaction. A VR platform that reliably elicits positive prediction errors and targets each reward phase could serve simultaneously as a clinical intervention, a research instrument and a personalized medicine tool, allowing clinicians to identify which component of reward processing is impaired in a given patient and to tune the virtual experience accordingly. The authors frame their contribution as a roadmap rather than a finished treatment, but the destination is clear: a future in which the brain&#8217;s pleasure circuits, dulled by illness, can be systematically and safely reawakened, one well-timed surprise at a time, inside a headset.</p>
<p><strong>Subject of Research:</strong> Using interactive virtual reality to generate reward prediction errors and target anticipatory, consummatory and learning reward phases as a treatment approach for anhedonia.</p>
<p><strong>Article Title:</strong> Reward prediction error and reward components in interactive virtual reality for anhedonia</p>
<p><strong>Article References:</strong> Kosa, M., Barnes-Horowitz, N. M., &amp; Craske, M. G. (2026). Reward prediction error and reward components in interactive virtual reality for anhedonia. <em>Nature Mental Health</em>. <a href="https://doi.org/10.1038/s44220-026-00731-4" rel="noopener noreferrer">https://doi.org/10.1038/s44220-026-00731-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s44220-026-00731-4" rel="noopener noreferrer">10.1038/s44220-026-00731-4</a></p>
<p><strong>Keywords:</strong> anhedonia, virtual reality, reward prediction error, dopamine, reward sensitivity, positive affect treatment, reinforcement learning, ventral striatum, behavioral activation, mental health, immersive technology, depression</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201944</post-id>	</item>
	</channel>
</rss>
