Sunday, October 11, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Policy

Robots Learn to Reuse Old Experience Without Being Fooled by It

October 11, 2026
in Policy
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
Robots Learn to Reuse Old Experience Without Being Fooled by It

Robots Learn to Reuse Old Experience Without Being Fooled by It

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Robots that operate in the physical world face an uncomfortable truth: everything they learn is based on a version of reality that no longer exists. A control policy changes continuously as training proceeds, hardware wears down, payloads shift, and friction, gravity, wind, sensor noise and contact conditions refuse to stay constant. Yet real-world interaction is expensive, sometimes dangerously slow, and often impractical to repeat at scale. An agent cannot simply throw away its accumulated experience and collect a fresh dataset every time the ground beneath its wheels or the arm on its shoulder changes. The promise of reinforcement learning for embodied systems has always depended on reusing the past; the problem is that the past can quietly become a lie.

This tension sits at the heart of a fundamental weakness in standard off-policy reinforcement learning. In the off-policy paradigm, an agent stores its transitions — the states it observed, the actions it took and the next states that resulted — in a large memory structure known as a replay buffer, and it draws on that buffer repeatedly to update its value estimates and improve its behavior. The approach is enormously sample-efficient precisely because old data remains useful. But those historical transitions were generated by older policies operating under older physical dynamics. When a learning algorithm treats them as if they still describe the current environment, the resulting value estimates become biased. Worse, the policy may converge toward actions whose predicted outcomes are no longer physically possible in the world it actually inhabits, producing confident behavior built on stale assumptions.

A research team led by Tsinghua University, working with Huawei Technologies’ 2012 Laboratories and the Shanghai Research Institute for Intelligent Autonomous Systems at Tongji University, has now proposed a way to resolve this conflict with a single mathematical idea. Their insight is that two seemingly different sources of change — a shift in the agent’s own policy and a shift in the environment’s dynamics — can be viewed through one common lens. A policy shift changes which action is selected in a given state, while a dynamics shift changes which next state follows from a given action. In both cases, what actually changes is the joint distribution over states, actions and next states. The researchers call this joint distribution transition occupancy, and they argue that if an agent can detect which of its stored transitions no longer match the current transition occupancy, it can correct for policy drift and physical change at the same time.

That principle has been turned into a practical algorithm called Occupancy-Matching Policy Optimization, or OMPO. Rather than treating every replayed transition as equally trustworthy, OMPO maintains two buffers of very different character. The first is a large global buffer holding the accumulated, heterogeneous history of the agent’s experience — data collected under earlier controllers, older hardware conditions and previous task variants. The second is a small first-in, first-out local buffer containing only the most recent interactions, which by construction reflect the current policy and the current physical regime. A discriminator network then compares transitions drawn from the two buffers and learns to estimate which historical samples remain compatible with the present. Compatible experience is retained and used for learning; stale or mismatched transitions are downweighted so they cannot poison the value estimates.

Several additional design choices make the method robust in practice. OMPO employs a sign-free logarithmic link that reformulates the matching objective as a stable min-max optimization, allowing the algorithm to handle reward signals that mix bonuses and penalties without requiring a task-specific reward transformation. A distributional critic models the full distribution of possible returns rather than only their mean, which helps the agent account for randomness arising from action noise, imperfect sensing and the inherently stochastic physics of contact. For tasks that operate on images, the policy, the critic and the discriminator share a co-trained visual encoder, while the action selection is carried out by an ODE-based flow actor capable of representing multimodal action distributions — a meaningful advantage in manipulation, where several distinct motor strategies may all lead to success.

The evaluation was deliberately broad, spanning three forms of distribution shift: policy shifts under stationary dynamics, transfer between tasks or domains, and the hardest combination of policy shifts together with non-stationary dynamics. The benchmark suite covered eight DeepMind Control locomotion tasks, a quadruped dog walk-to-run transfer experiment, four MuJoCo environments in which body dimensions, gravity and wind changed during training, manipulation suites including Panda-Gym and Meta-World, and finally a proof-of-concept test on a real physical robot. Across these settings, OMPO was compared against strong off-policy baselines such as SAC and TD7, as well as context-aware and non-stationary methods such as CaDM and CEMRL.

The results were consistent. In DeepMind Control, OMPO learned faster and more stably than SAC and TD7 on both state-based and visual variants of the tasks. In the dog transfer experiment, the agent entered the running objective with a replay buffer dominated by walking experience — exactly the situation that derails conventional off-policy learning. Context-aware baselines struggled to adapt, whereas OMPO reweighted the old walking data according to its compatibility with the new regime and acquired the faster gait with substantially less negative transfer. Under the non-stationary MuJoCo conditions, where torso and foot lengths changed across episodes and gravity and wind shifted stochastically during training, OMPO consistently outperformed CaDM and CEMRL. The authors read these outcomes as support for the method’s central premise: an agent does not need to explicitly identify every changed physical parameter if it can simply detect which transitions no longer match the current occupancy.

Manipulation results reinforced the picture. With action noise injected into Panda-Gym, OMPO achieved a 98.4 percent success rate on Panda-Reach-Dense and 94.3 percent on Panda-Reach-Sparse. Across five reported contact-rich Meta-World tasks it recorded the highest success rate in each, with the largest margins appearing in coffee_push, where success reached 76.8 percent against 26.4 percent for the strongest baseline, and in hammer, where it reached 78.3 percent against 32.8 percent. The team then moved to hardware. On a TianJi robot performing pick-and-place, a disturbed condition moved the object before grasping, invalidating the outcome of the robot’s nominal approach. Anchored by its recent interactions, the robot re-approached the displaced object, grasped it and completed the placement. The authors are careful to describe this as a proof of concept under a support-preserving disturbance rather than a comprehensive hardware benchmark, but it demonstrates that the occupancy-matching correction survives contact with real physics.

The broader significance of the work lies in what it makes possible. By providing one correction mechanism that covers both policy and dynamics shifts, OMPO offers a route to reusing costly historical experience without assuming a stationary world — an assumption that virtually all classical reinforcement learning theory quietly makes and that the physical world routinely violates. The framework could complement the current generation of vision-language-action models and world-model-based agents by serving as an online post-training mechanism for continual adaptation, allowing systems pretrained on massive datasets to keep adjusting safely as their bodies and surroundings evolve. The researchers are candid about the limits: OMPO still requires sufficient overlap between historical and current experience, and a local buffer that refreshes quickly enough to track the active regime. Support-breaking failures, extremely rapid shifts, actuator loss and severe hardware damage remain outside the current validation, and larger real-robot studies together with stronger convergence and stability guarantees are the stated next steps. The research, conducted by Yu Luo, Lei Lv, Fuchun Sun and Huaping Liu, was supported by the National Key Research and Development Program of China, with the implementation publicly released on GitHub and the accepted manuscript published in National Science Review on 19 August 2026.

Subject of Research: Occupancy-matching reinforcement learning for robot adaptation to policy and dynamics shifts

Article Title: New AI method helps robots learn from the past without being trapped by it

Article References: New AI method helps robots learn from the past without being trapped by it. (n.d.). Original publication

Image Credits: AI Generated

DOI: Not provided

Keywords: reinforcement learning, robotics, OMPO, replay buffer, distribution shift, off-policy learning, transition occupancy, non-stationary dynamics, manipulation, locomotion, Tsinghua University, National Science Review

Cite Scienmag News

Denise Maddox. (October 11, 2026). Robots Learn to Reuse Old Experience Without Being Fooled by It. Scienmag. https://scienmag.com/robots-learn-to-reuse-old-experience-without-being-fooled-by-it/

Denise Maddox. "Robots Learn to Reuse Old Experience Without Being Fooled by It." Scienmag, 11 October 2026, https://scienmag.com/robots-learn-to-reuse-old-experience-without-being-fooled-by-it/. Accessed 11 October 2026.

Denise Maddox. "Robots Learn to Reuse Old Experience Without Being Fooled by It." Scienmag. October 11, 2026. https://scienmag.com/robots-learn-to-reuse-old-experience-without-being-fooled-by-it/

Tags: adaptive control policiescontinuous environment variationdistribution shiftembodied systems machine learninglearning from changing environmentslocomotionmanipulationNational Science Reviewnon-stationary dynamicsoff-policy learningoff-policy learning challengesOMPOreal-world robot experiencereinforcement learningreinforcement learning data reusereplay bufferreplay buffer limitationsrobot policy adaptationroboticsRobotics reinforcement learningsensor noise and hardware weartransfer learning in roboticstransition occupancyTsinghua University
Share26Tweet16
Previous Post

Turning Toxic Mine Water Into Metal Treasure With Electricity

Next Post

Satellite Records Reveal Rising Rainfall in One of Earth’s Driest Corners of Iraq

Related Posts

Older Haitians Living With HIV Face a Rising Tide of Chronic Disease
Medicine

Older Haitians Living With HIV Face a Rising Tide of Chronic Disease

October 11, 2026
Insurance mandates, not just notifications, drive supplemental breast screening
Policy

Insurance mandates, not just notifications, drive supplemental breast screening

October 11, 2026
Antibodies Fade Fast: Malawi Study Maps Hidden Waves of COVID-19 Infection
Medicine

Antibodies Fade Fast: Malawi Study Maps Hidden Waves of COVID-19 Infection

October 11, 2026
Medical AI in radiology may copy and amplify bias, major review finds
Policy

Medical AI in radiology may copy and amplify bias, major review finds

October 11, 2026
Smartphone Cough Monitoring Shows Promise but Falls Short for Tuberculosis Screening in Uganda
Medicine

Smartphone Cough Monitoring Shows Promise but Falls Short for Tuberculosis Screening in Uganda

October 11, 2026
Vietnam’s Climate-Smart Farming Revolution Stalls Where Gender Meets the Ground
Policy

Vietnam’s Climate-Smart Farming Revolution Stalls Where Gender Meets the Ground

October 11, 2026
Next Post
Satellite Records Reveal Rising Rainfall in One of Earth’s Driest Corners of Iraq

Satellite Records Reveal Rising Rainfall in One of Earth's Driest Corners of Iraq

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Mapping the Totality of Childhood Exposures: How Exposome Science Is Rewriting Pediatric Health
  • Satellite Records Reveal Rising Rainfall in One of Earth’s Driest Corners of Iraq
  • Robots Learn to Reuse Old Experience Without Being Fooled by It
  • Turning Toxic Mine Water Into Metal Treasure With Electricity

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading