Sunday, October 4, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

New AI Framework Teaches Robots to Choose Their Own Skills and When to Use Them

October 4, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
New AI Framework Teaches Robots to Choose Their Own Skills and When to Use Them

New AI Framework Teaches Robots to Choose Their Own Skills and When to Use Them

New AI Framework Teaches Robots to Choose Their Own Skills and When to Use Them

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

One of the most stubborn problems in artificial intelligence is teaching an agent to carry out tasks that stretch far into the future, where rewards arrive rarely and the path to success is measured in hundreds of small decisions. A robot asked to stack blocks into a pyramid or open a microwave must chain together dozens of primitive motions, and a learning algorithm that treats every one of those motions as an independent choice quickly drowns in the sheer scale of the search. Researchers at Shanghai Jiao Tong University and collaborating institutions have now unveiled a framework that tackles this problem head-on by letting an agent learn not only which skills to deploy, but also how long each skill should last, adapting that timing on the fly as circumstances change.

The framework, described in the journal Applied Intelligence, is called HALSI, short for Hierarchical Adaptive Latent Skill Inference. Its central insight is that the temporal structure of behavior matters as much as the behavior itself. Earlier skill-based hierarchical methods decompose complicated tasks into reusable behavior primitives, which eases the burden of long-horizon decision-making, but they typically fix the length of each skill in advance and compress behavior into oversimplified latent representations. That rigidity limits both temporal flexibility and the diversity of behaviors an agent can express. HALSI removes the fixed horizon entirely, learning skills of variable duration from data and adjusting them online as the task unfolds.

At the heart of the system sits a Transformer encoder that reads variable-length trajectories of state-action pairs and distills them into a compact latent embedding. The architecture is causal, meaning each token corresponds to a state-action tuple, and positional encodings preserve the sequential order so that self-attention can capture long-range dependencies across time. A special classification token at the output summarizes the overall temporal and behavioral pattern of the skill into a single vector. Crucially, during pretraining the system preserves the natural lengths of trajectory segments rather than chopping demonstrations into fixed windows, which allows the encoder to learn duration-agnostic representations that retain the genuine variability of different behaviors.

To turn those latent skills into actual motion, HALSI pairs the encoder with a conditional diffusion policy. Denoising diffusion probabilistic models, which have transformed image generation in recent years, work here by iteratively refining Gaussian noise into a coherent action sequence, conditioned on both a timestep embedding and the concatenation of the latent skill vector with the current state. This generative approach lets the decoder model complex, multimodal action distributions and produce temporally consistent behaviors that a simpler policy class would struggle to capture. Once pretrained on offline demonstration data, the encoder and decoder remain frozen during online reinforcement learning, serving as a stable library of skills that the higher levels of the hierarchy can draw upon.

The first major innovation addresses the fixed-horizon problem directly: an adaptive duration policy that predicts a discrete execution length for each skill it selects. Rather than committing to a predetermined number of steps, the high-level controller outputs both a latent direction and a duration, and the diffusion decoder generates a matching action sequence of exactly that length before a new skill is sampled. The researchers found that different tasks induce strikingly different patterns of skill durations. In a slippery pushing task, durations cluster tightly at short values, reflecting the need for frequent fine-grained corrections on low-friction surfaces. In a tool-use task involving a hook, the distribution becomes broad and multimodal, as the agent alternates between brief reactive skills and longer sustained motions for aligning, inserting, and pulling. A pick-and-place task showed a bimodal pattern corresponding to its two distinct sub-goals.

The second innovation is a multi-scale optimization scheme built on Proximal Policy Optimization. A high-level modulation policy refines latent skills with task-aware residuals, adding a learned direction to the encoded latent mean so that the skill embedding itself is adjusted rather than the raw actions. Because directly perturbing behavior early in training would destabilize learning, the team introduced a curriculum: a scheduling factor follows a logistic curve over training steps, starting near zero so the agent initially relies almost entirely on the pretrained skill encoder, then gradually rising to give the residual policy increasing influence. A low-level soft-blending controller complements this by mixing the offline-decoded skill action with an online-learned adaptive actor at every step, allowing moment-to-moment corrections while preserving the structure of the skill prior.

The ablation studies reveal how sensitive this balance is. When the blending coefficient was set low, the agent largely ignored its pretrained skills and fell back on the reactive actor alone, producing unstable and inefficient behavior in long-horizon tasks such as pyramid stacking, where structured skill priors are essential for guidance. When the coefficient was pushed very high, the agent over-trusted the offline decoder and lost the ability to respond to distribution shifts, a weakness most visible in the slippery pushing environment where precise reactive adjustments are frequently needed. The best overall performance emerged consistently at a coefficient of 0.8, a setting that retains the semantic consistency of the offline skills while leaving room for online refinement.

Evaluated on two demanding benchmark suites, the framework delivered substantial gains. On four long-horizon manipulation tasks from the Reskill benchmark, built on MuJoCo-based Fetch environments, the agent had to transfer skills learned from roughly 40,000 trajectories collected with deliberately suboptimal scripted controllers in simplified settings, then apply them in downstream environments featuring slippery surfaces, cluttered distractor objects, multi-stage stacking goals, and contact-rich tool use. On the D4RL Kitchen benchmark, where a Franka arm operates household appliances across episodes of roughly 280 steps with rewards granted only when subtasks are completed, the sparse-reward structure makes credit assignment notoriously difficult. Across these tests, HALSI surpassed state-of-the-art hierarchical baselines, achieving up to 30 percent higher returns and converging 25 percent faster.

Beyond the headline numbers, the work offers a broader lesson about what makes hierarchical learning succeed. The duration analysis showed that a fixed-duration scheme would inevitably fragment some behaviors while leaving others insufficiently abstract, whereas a learned duration model allocates temporal granularity according to context and task stage. The authors suggest that jointly modeling latent skills and their adaptive temporal scope is the key ingredient, rather than improving either component in isolation. The source code has been released publicly, and the datasets are available from the corresponding author on reasonable request, lowering the barrier for other groups to build on the approach.

The implications reach well beyond simulated kitchens and block-stacking. Long-horizon, sparse-reward problems pervade robotics, autonomous driving, and industrial automation, and any method that lets agents reuse skills flexibly across tasks with differing dynamics could accelerate progress in all of them. By combining the sequence-modeling power of Transformers, the expressive generative capacity of diffusion models, and a principled curriculum for handing control from imitation to online learning, HALSI points toward a future in which robots do not merely execute preprogrammed routines, but decide for themselves which skill fits the moment and how long to commit to it.

Subject of Research: Hierarchical reinforcement learning with adaptive latent skill inference for temporal abstraction in long-horizon robotic manipulation

Article Title: HALSI: Hierarchical adaptive latent skill inference for temporal abstraction in reinforcement learning

Article References: Si, X., Ping, Y., Wang, T., Chen, K., & Gao, Y. (2026). HALSI: Hierarchical adaptive latent skill inference for temporal abstraction in reinforcement learning. Applied Intelligence, 56(15), Article 444. https://doi.org/10.1007/s10489-026-07485-7

Image Credits: AI Generated

DOI: 10.1007/s10489-026-07485-7

Keywords: reinforcement learning, hierarchical RL, skill abstraction, temporal abstraction, diffusion policy, transformer encoder, robotic manipulation, PPO, D4RL Kitchen, Fetch benchmark, machine learning, latent skills

Cite Scienmag News

Denise Maddox. (October 4, 2026). New AI Framework Teaches Robots to Choose Their Own Skills and When to Use Them. Scienmag. https://scienmag.com/new-ai-framework-teaches-robots-to-choose-their-own-skills-and-when-to-use-them/

Denise Maddox. "New AI Framework Teaches Robots to Choose Their Own Skills and When to Use Them." Scienmag, 4 October 2026, https://scienmag.com/new-ai-framework-teaches-robots-to-choose-their-own-skills-and-when-to-use-them/. Accessed 4 October 2026.

Denise Maddox. "New AI Framework Teaches Robots to Choose Their Own Skills and When to Use Them." Scienmag. October 4, 2026. https://scienmag.com/new-ai-framework-teaches-robots-to-choose-their-own-skills-and-when-to-use-them/

Tags: adaptive timing in robotic skillsAI framework for complex robotic tasksbehavior decomposition in AID4RL Kitchendiffusion policyFetch benchmarkHierarchical adaptive skill inferencehierarchical RLlatent skillslong-term decision-making in AIMachine learningmulti-step task executionPPOprimitive motion chainingreinforcement learningreusable behavior primitivesreward-based reinforcement learningrobotic manipulationrobotics task planningskill abstractionskill duration learning in robotstackling sparse rewards in reinforcement learningtemporal abstractiontransformer encoder
Share26Tweet16
Previous Post

Two Therapy Styles, One Result: Easing the Burden of Cancer Caregivers

Next Post

Extra Feed on the Roof of the World Makes Yak Meat Dramatically More Tender

Related Posts

Before the First Breath: How Teams Prepare for Birth at the Edge of Viability
Technology and Engineering

Before the First Breath: How Teams Prepare for Birth at the Edge of Viability

October 4, 2026
Bare Gold Electrode Detects Trace Mercury in Water Without Nanocoatings
Technology and Engineering

Bare Gold Electrode Detects Trace Mercury in Water Without Nanocoatings

October 4, 2026
Steel and Carbon Fibers Turn Concrete Into a Self-Sensing Structural Material
Technology and Engineering

Steel and Carbon Fibers Turn Concrete Into a Self-Sensing Structural Material

October 4, 2026
New Rendering Method Brings Realistic Tree Canopies to Real Time on a Memory Budget
Technology and Engineering

New Rendering Method Brings Realistic Tree Canopies to Real Time on a Memory Budget

October 4, 2026
From ELIZA to RAG: How Chatbots Learned to Look Up Facts Before They Speak
Technology and Engineering

From ELIZA to RAG: How Chatbots Learned to Look Up Facts Before They Speak

October 4, 2026
AI Simulator Predicts Bacteria in Wastewater From Simple Measurements, Beating GANs by 35 Percent
Technology and Engineering

AI Simulator Predicts Bacteria in Wastewater From Simple Measurements, Beating GANs by 35 Percent

October 4, 2026
Next Post
Extra Feed on the Roof of the World Makes Yak Meat Dramatically More Tender

Extra Feed on the Roof of the World Makes Yak Meat Dramatically More Tender

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • When Schools Are Shells: Displaced Children in Ethiopia Expose a Hollow Right to Education
  • Robotic Floats Reveal a Hidden September Collapse in Arabian Sea Plankton
  • Cosmic Horizon Flux and Matter Creation Reshape the Thermodynamic Story of Dark Energy
  • Deep Neural Network Predicts Air Quality in Indian City With Striking Accuracy

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,149 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading