<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>robotics task planning &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/robotics-task-planning/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 06:53:57 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>robotics task planning &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New AI Framework Teaches Robots to Choose Their Own Skills and When to Use Them</title>
		<link>https://scienmag.com/new-ai-framework-teaches-robots-to-choose-their-own-skills-and-when-to-use-them/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 06:53:57 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive timing in robotic skills]]></category>
		<category><![CDATA[AI framework for complex robotic tasks]]></category>
		<category><![CDATA[behavior decomposition in AI]]></category>
		<category><![CDATA[D4RL Kitchen]]></category>
		<category><![CDATA[diffusion policy]]></category>
		<category><![CDATA[Fetch benchmark]]></category>
		<category><![CDATA[Hierarchical adaptive skill inference]]></category>
		<category><![CDATA[hierarchical RL]]></category>
		<category><![CDATA[latent skills]]></category>
		<category><![CDATA[long-term decision-making in AI]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[multi-step task execution]]></category>
		<category><![CDATA[PPO]]></category>
		<category><![CDATA[primitive motion chaining]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[reusable behavior primitives]]></category>
		<category><![CDATA[reward-based reinforcement learning]]></category>
		<category><![CDATA[robotic manipulation]]></category>
		<category><![CDATA[robotics task planning]]></category>
		<category><![CDATA[skill abstraction]]></category>
		<category><![CDATA[skill duration learning in robots]]></category>
		<category><![CDATA[tackling sparse rewards in reinforcement learning]]></category>
		<category><![CDATA[temporal abstraction]]></category>
		<category><![CDATA[transformer encoder]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233982</guid>

					<description><![CDATA[A new hierarchical reinforcement learning framework called HALSI learns latent skills of flexible duration, achieving up to 30% higher returns and 25% faster convergence on long-horizon robotic manipulation benchmarks.]]></description>
										<content:encoded><![CDATA[<p>One of the most stubborn problems in artificial intelligence is teaching an agent to carry out tasks that stretch far into the future, where rewards arrive rarely and the path to success is measured in hundreds of small decisions. A robot asked to stack blocks into a pyramid or open a microwave must chain together dozens of primitive motions, and a learning algorithm that treats every one of those motions as an independent choice quickly drowns in the sheer scale of the search. Researchers at Shanghai Jiao Tong University and collaborating institutions have now unveiled a framework that tackles this problem head-on by letting an agent learn not only which skills to deploy, but also how long each skill should last, adapting that timing on the fly as circumstances change.</p>
<p>The framework, described in the journal Applied Intelligence, is called HALSI, short for Hierarchical Adaptive Latent Skill Inference. Its central insight is that the temporal structure of behavior matters as much as the behavior itself. Earlier skill-based hierarchical methods decompose complicated tasks into reusable behavior primitives, which eases the burden of long-horizon decision-making, but they typically fix the length of each skill in advance and compress behavior into oversimplified latent representations. That rigidity limits both temporal flexibility and the diversity of behaviors an agent can express. HALSI removes the fixed horizon entirely, learning skills of variable duration from data and adjusting them online as the task unfolds.</p>
<p>At the heart of the system sits a Transformer encoder that reads variable-length trajectories of state-action pairs and distills them into a compact latent embedding. The architecture is causal, meaning each token corresponds to a state-action tuple, and positional encodings preserve the sequential order so that self-attention can capture long-range dependencies across time. A special classification token at the output summarizes the overall temporal and behavioral pattern of the skill into a single vector. Crucially, during pretraining the system preserves the natural lengths of trajectory segments rather than chopping demonstrations into fixed windows, which allows the encoder to learn duration-agnostic representations that retain the genuine variability of different behaviors.</p>
<p>To turn those latent skills into actual motion, HALSI pairs the encoder with a conditional diffusion policy. Denoising diffusion probabilistic models, which have transformed image generation in recent years, work here by iteratively refining Gaussian noise into a coherent action sequence, conditioned on both a timestep embedding and the concatenation of the latent skill vector with the current state. This generative approach lets the decoder model complex, multimodal action distributions and produce temporally consistent behaviors that a simpler policy class would struggle to capture. Once pretrained on offline demonstration data, the encoder and decoder remain frozen during online reinforcement learning, serving as a stable library of skills that the higher levels of the hierarchy can draw upon.</p>
<p>The first major innovation addresses the fixed-horizon problem directly: an adaptive duration policy that predicts a discrete execution length for each skill it selects. Rather than committing to a predetermined number of steps, the high-level controller outputs both a latent direction and a duration, and the diffusion decoder generates a matching action sequence of exactly that length before a new skill is sampled. The researchers found that different tasks induce strikingly different patterns of skill durations. In a slippery pushing task, durations cluster tightly at short values, reflecting the need for frequent fine-grained corrections on low-friction surfaces. In a tool-use task involving a hook, the distribution becomes broad and multimodal, as the agent alternates between brief reactive skills and longer sustained motions for aligning, inserting, and pulling. A pick-and-place task showed a bimodal pattern corresponding to its two distinct sub-goals.</p>
<p>The second innovation is a multi-scale optimization scheme built on Proximal Policy Optimization. A high-level modulation policy refines latent skills with task-aware residuals, adding a learned direction to the encoded latent mean so that the skill embedding itself is adjusted rather than the raw actions. Because directly perturbing behavior early in training would destabilize learning, the team introduced a curriculum: a scheduling factor follows a logistic curve over training steps, starting near zero so the agent initially relies almost entirely on the pretrained skill encoder, then gradually rising to give the residual policy increasing influence. A low-level soft-blending controller complements this by mixing the offline-decoded skill action with an online-learned adaptive actor at every step, allowing moment-to-moment corrections while preserving the structure of the skill prior.</p>
<p>The ablation studies reveal how sensitive this balance is. When the blending coefficient was set low, the agent largely ignored its pretrained skills and fell back on the reactive actor alone, producing unstable and inefficient behavior in long-horizon tasks such as pyramid stacking, where structured skill priors are essential for guidance. When the coefficient was pushed very high, the agent over-trusted the offline decoder and lost the ability to respond to distribution shifts, a weakness most visible in the slippery pushing environment where precise reactive adjustments are frequently needed. The best overall performance emerged consistently at a coefficient of 0.8, a setting that retains the semantic consistency of the offline skills while leaving room for online refinement.</p>
<p>Evaluated on two demanding benchmark suites, the framework delivered substantial gains. On four long-horizon manipulation tasks from the Reskill benchmark, built on MuJoCo-based Fetch environments, the agent had to transfer skills learned from roughly 40,000 trajectories collected with deliberately suboptimal scripted controllers in simplified settings, then apply them in downstream environments featuring slippery surfaces, cluttered distractor objects, multi-stage stacking goals, and contact-rich tool use. On the D4RL Kitchen benchmark, where a Franka arm operates household appliances across episodes of roughly 280 steps with rewards granted only when subtasks are completed, the sparse-reward structure makes credit assignment notoriously difficult. Across these tests, HALSI surpassed state-of-the-art hierarchical baselines, achieving up to 30 percent higher returns and converging 25 percent faster.</p>
<p>Beyond the headline numbers, the work offers a broader lesson about what makes hierarchical learning succeed. The duration analysis showed that a fixed-duration scheme would inevitably fragment some behaviors while leaving others insufficiently abstract, whereas a learned duration model allocates temporal granularity according to context and task stage. The authors suggest that jointly modeling latent skills and their adaptive temporal scope is the key ingredient, rather than improving either component in isolation. The source code has been released publicly, and the datasets are available from the corresponding author on reasonable request, lowering the barrier for other groups to build on the approach.</p>
<p>The implications reach well beyond simulated kitchens and block-stacking. Long-horizon, sparse-reward problems pervade robotics, autonomous driving, and industrial automation, and any method that lets agents reuse skills flexibly across tasks with differing dynamics could accelerate progress in all of them. By combining the sequence-modeling power of Transformers, the expressive generative capacity of diffusion models, and a principled curriculum for handing control from imitation to online learning, HALSI points toward a future in which robots do not merely execute preprogrammed routines, but decide for themselves which skill fits the moment and how long to commit to it.</p>
<p><strong>Subject of Research:</strong> Hierarchical reinforcement learning with adaptive latent skill inference for temporal abstraction in long-horizon robotic manipulation</p>
<p><strong>Article Title:</strong> HALSI: Hierarchical adaptive latent skill inference for temporal abstraction in reinforcement learning</p>
<p><strong>Article References:</strong> Si, X., Ping, Y., Wang, T., Chen, K., &amp; Gao, Y. (2026). HALSI: Hierarchical adaptive latent skill inference for temporal abstraction in reinforcement learning. <em>Applied Intelligence, 56</em>(15), Article 444. <a href="https://doi.org/10.1007/s10489-026-07485-7" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07485-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07485-7" rel="noopener noreferrer">10.1007/s10489-026-07485-7</a></p>
<p><strong>Keywords:</strong> reinforcement learning, hierarchical RL, skill abstraction, temporal abstraction, diffusion policy, transformer encoder, robotic manipulation, PPO, D4RL Kitchen, Fetch benchmark, machine learning, latent skills</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233982</post-id>	</item>
	</channel>
</rss>
