<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>sensor noise and hardware wear &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/sensor-noise-and-hardware-wear/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 11 Oct 2026 08:01:45 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>sensor noise and hardware wear &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Robots Learn to Reuse Old Experience Without Being Fooled by It</title>
		<link>https://scienmag.com/robots-learn-to-reuse-old-experience-without-being-fooled-by-it/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 11 Oct 2026 08:01:45 +0000</pubDate>
				<category><![CDATA[Policy]]></category>
		<category><![CDATA[adaptive control policies]]></category>
		<category><![CDATA[continuous environment variation]]></category>
		<category><![CDATA[distribution shift]]></category>
		<category><![CDATA[embodied systems machine learning]]></category>
		<category><![CDATA[learning from changing environments]]></category>
		<category><![CDATA[locomotion]]></category>
		<category><![CDATA[manipulation]]></category>
		<category><![CDATA[National Science Review]]></category>
		<category><![CDATA[non-stationary dynamics]]></category>
		<category><![CDATA[off-policy learning]]></category>
		<category><![CDATA[off-policy learning challenges]]></category>
		<category><![CDATA[OMPO]]></category>
		<category><![CDATA[real-world robot experience]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[reinforcement learning data reuse]]></category>
		<category><![CDATA[replay buffer]]></category>
		<category><![CDATA[replay buffer limitations]]></category>
		<category><![CDATA[robot policy adaptation]]></category>
		<category><![CDATA[robotics]]></category>
		<category><![CDATA[Robotics reinforcement learning]]></category>
		<category><![CDATA[sensor noise and hardware wear]]></category>
		<category><![CDATA[transfer learning in robotics]]></category>
		<category><![CDATA[transition occupancy]]></category>
		<category><![CDATA[Tsinghua University]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=261626</guid>

					<description><![CDATA[Researchers at Tsinghua University and collaborators developed Occupancy-Matching Policy Optimization, a reinforcement learning method that aligns historical robot experience with recent interactions to maintain stable learning as policies, hardware and physical conditions change.]]></description>
										<content:encoded><![CDATA[<p>Robots that operate in the physical world face an uncomfortable truth: everything they learn is based on a version of reality that no longer exists. A control policy changes continuously as training proceeds, hardware wears down, payloads shift, and friction, gravity, wind, sensor noise and contact conditions refuse to stay constant. Yet real-world interaction is expensive, sometimes dangerously slow, and often impractical to repeat at scale. An agent cannot simply throw away its accumulated experience and collect a fresh dataset every time the ground beneath its wheels or the arm on its shoulder changes. The promise of reinforcement learning for embodied systems has always depended on reusing the past; the problem is that the past can quietly become a lie.</p>
<p>This tension sits at the heart of a fundamental weakness in standard off-policy reinforcement learning. In the off-policy paradigm, an agent stores its transitions — the states it observed, the actions it took and the next states that resulted — in a large memory structure known as a replay buffer, and it draws on that buffer repeatedly to update its value estimates and improve its behavior. The approach is enormously sample-efficient precisely because old data remains useful. But those historical transitions were generated by older policies operating under older physical dynamics. When a learning algorithm treats them as if they still describe the current environment, the resulting value estimates become biased. Worse, the policy may converge toward actions whose predicted outcomes are no longer physically possible in the world it actually inhabits, producing confident behavior built on stale assumptions.</p>
<p>A research team led by Tsinghua University, working with Huawei Technologies&#8217; 2012 Laboratories and the Shanghai Research Institute for Intelligent Autonomous Systems at Tongji University, has now proposed a way to resolve this conflict with a single mathematical idea. Their insight is that two seemingly different sources of change — a shift in the agent&#8217;s own policy and a shift in the environment&#8217;s dynamics — can be viewed through one common lens. A policy shift changes which action is selected in a given state, while a dynamics shift changes which next state follows from a given action. In both cases, what actually changes is the joint distribution over states, actions and next states. The researchers call this joint distribution transition occupancy, and they argue that if an agent can detect which of its stored transitions no longer match the current transition occupancy, it can correct for policy drift and physical change at the same time.</p>
<p>That principle has been turned into a practical algorithm called Occupancy-Matching Policy Optimization, or OMPO. Rather than treating every replayed transition as equally trustworthy, OMPO maintains two buffers of very different character. The first is a large global buffer holding the accumulated, heterogeneous history of the agent&#8217;s experience — data collected under earlier controllers, older hardware conditions and previous task variants. The second is a small first-in, first-out local buffer containing only the most recent interactions, which by construction reflect the current policy and the current physical regime. A discriminator network then compares transitions drawn from the two buffers and learns to estimate which historical samples remain compatible with the present. Compatible experience is retained and used for learning; stale or mismatched transitions are downweighted so they cannot poison the value estimates.</p>
<p>Several additional design choices make the method robust in practice. OMPO employs a sign-free logarithmic link that reformulates the matching objective as a stable min-max optimization, allowing the algorithm to handle reward signals that mix bonuses and penalties without requiring a task-specific reward transformation. A distributional critic models the full distribution of possible returns rather than only their mean, which helps the agent account for randomness arising from action noise, imperfect sensing and the inherently stochastic physics of contact. For tasks that operate on images, the policy, the critic and the discriminator share a co-trained visual encoder, while the action selection is carried out by an ODE-based flow actor capable of representing multimodal action distributions — a meaningful advantage in manipulation, where several distinct motor strategies may all lead to success.</p>
<p>The evaluation was deliberately broad, spanning three forms of distribution shift: policy shifts under stationary dynamics, transfer between tasks or domains, and the hardest combination of policy shifts together with non-stationary dynamics. The benchmark suite covered eight DeepMind Control locomotion tasks, a quadruped dog walk-to-run transfer experiment, four MuJoCo environments in which body dimensions, gravity and wind changed during training, manipulation suites including Panda-Gym and Meta-World, and finally a proof-of-concept test on a real physical robot. Across these settings, OMPO was compared against strong off-policy baselines such as SAC and TD7, as well as context-aware and non-stationary methods such as CaDM and CEMRL.</p>
<p>The results were consistent. In DeepMind Control, OMPO learned faster and more stably than SAC and TD7 on both state-based and visual variants of the tasks. In the dog transfer experiment, the agent entered the running objective with a replay buffer dominated by walking experience — exactly the situation that derails conventional off-policy learning. Context-aware baselines struggled to adapt, whereas OMPO reweighted the old walking data according to its compatibility with the new regime and acquired the faster gait with substantially less negative transfer. Under the non-stationary MuJoCo conditions, where torso and foot lengths changed across episodes and gravity and wind shifted stochastically during training, OMPO consistently outperformed CaDM and CEMRL. The authors read these outcomes as support for the method&#8217;s central premise: an agent does not need to explicitly identify every changed physical parameter if it can simply detect which transitions no longer match the current occupancy.</p>
<p>Manipulation results reinforced the picture. With action noise injected into Panda-Gym, OMPO achieved a 98.4 percent success rate on Panda-Reach-Dense and 94.3 percent on Panda-Reach-Sparse. Across five reported contact-rich Meta-World tasks it recorded the highest success rate in each, with the largest margins appearing in coffee_push, where success reached 76.8 percent against 26.4 percent for the strongest baseline, and in hammer, where it reached 78.3 percent against 32.8 percent. The team then moved to hardware. On a TianJi robot performing pick-and-place, a disturbed condition moved the object before grasping, invalidating the outcome of the robot&#8217;s nominal approach. Anchored by its recent interactions, the robot re-approached the displaced object, grasped it and completed the placement. The authors are careful to describe this as a proof of concept under a support-preserving disturbance rather than a comprehensive hardware benchmark, but it demonstrates that the occupancy-matching correction survives contact with real physics.</p>
<p>The broader significance of the work lies in what it makes possible. By providing one correction mechanism that covers both policy and dynamics shifts, OMPO offers a route to reusing costly historical experience without assuming a stationary world — an assumption that virtually all classical reinforcement learning theory quietly makes and that the physical world routinely violates. The framework could complement the current generation of vision-language-action models and world-model-based agents by serving as an online post-training mechanism for continual adaptation, allowing systems pretrained on massive datasets to keep adjusting safely as their bodies and surroundings evolve. The researchers are candid about the limits: OMPO still requires sufficient overlap between historical and current experience, and a local buffer that refreshes quickly enough to track the active regime. Support-breaking failures, extremely rapid shifts, actuator loss and severe hardware damage remain outside the current validation, and larger real-robot studies together with stronger convergence and stability guarantees are the stated next steps. The research, conducted by Yu Luo, Lei Lv, Fuchun Sun and Huaping Liu, was supported by the National Key Research and Development Program of China, with the implementation publicly released on GitHub and the accepted manuscript published in National Science Review on 19 August 2026.</p>
<p><strong>Subject of Research:</strong> Occupancy-matching reinforcement learning for robot adaptation to policy and dynamics shifts</p>
<p><strong>Article Title:</strong> New AI method helps robots learn from the past without being trapped by it</p>
<p><strong>Article References:</strong> New AI method helps robots learn from the past without being trapped by it. (n.d.). <a href="https://www.eurekalert.org/news-releases/1142246" rel="noopener noreferrer">Original publication</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> Not provided</p>
<p><strong>Keywords:</strong> reinforcement learning, robotics, OMPO, replay buffer, distribution shift, off-policy learning, transition occupancy, non-stationary dynamics, manipulation, locomotion, Tsinghua University, National Science Review</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">261626</post-id>	</item>
	</channel>
</rss>
