<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>embodied intelligence in robotic manipulation &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/embodied-intelligence-in-robotic-manipulation/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 19:51:40 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>embodied intelligence in robotic manipulation &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Robots That Learn by Touch: How Embodied Intelligence Is Rewriting Machine Manipulation</title>
		<link>https://scienmag.com/robots-that-learn-by-touch-how-embodied-intelligence-is-rewriting-machine-manipulation/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 19:51:40 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[advancements in robotic touch and perception]]></category>
		<category><![CDATA[artificial general intelligence]]></category>
		<category><![CDATA[diffusion policy]]></category>
		<category><![CDATA[embodied AI in unstructured environments]]></category>
		<category><![CDATA[embodied intelligence]]></category>
		<category><![CDATA[embodied intelligence in robotic manipulation]]></category>
		<category><![CDATA[history of embodied intelligence]]></category>
		<category><![CDATA[humanoid robots]]></category>
		<category><![CDATA[imitation learning]]></category>
		<category><![CDATA[machine learning for physical tasks]]></category>
		<category><![CDATA[multimodal perception]]></category>
		<category><![CDATA[neural networks for robotic manipulation]]></category>
		<category><![CDATA[physical cognition in robotics]]></category>
		<category><![CDATA[physical interaction in robotics]]></category>
		<category><![CDATA[real-world robotic applications]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[robot datasets]]></category>
		<category><![CDATA[robot manipulation]]></category>
		<category><![CDATA[robotic sensory-motor integration]]></category>
		<category><![CDATA[robots grasping and manipulating objects]]></category>
		<category><![CDATA[sim-to-real transfer]]></category>
		<category><![CDATA[tactile sensing in robots]]></category>
		<category><![CDATA[vision-language-action models]]></category>
		<category><![CDATA[world models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=218654</guid>

					<description><![CDATA[A new systematic review maps how embodied intelligence, from vision-language-action models to world models, is transforming robotic manipulation and what obstacles remain before robots achieve human-like dexterity.]]></description>
										<content:encoded><![CDATA[<p>For more than seventy years, artificial intelligence has lived mostly in the abstract. Machines could beat grandmasters at chess and generate fluent prose, yet they struggled to pick up a wine glass without crushing it. A sweeping new review published in the journal Vicinagearth argues that this gap between digital brilliance and physical clumsiness is now closing, and that the key is a concept researchers call embodied intelligence: the idea that genuine understanding emerges only when a mind is wired directly into a body that senses and acts in the real world. The paper, led by Honghao Song and Zhe Sun of Northwestern Polytechnical University and colleagues, offers the first systematic map of how embodied intelligence is transforming robotic manipulation, the ability of machines to grasp, twist, fold, and assemble objects in unstructured environments.</p>
<p>The intellectual roots of the field trace back to Alan Turing, who in 1950 suggested that intelligence is not merely abstract computation but something demonstrated through dynamic interaction between a body and its environment. The review frames that insight as a theoretical foundation: a physical carrier is a necessary prerequisite for intelligence to operate in the physical world. In practice, the authors define embodied manipulation as a closed-loop process that takes embodied cognition as its engine and a physical robot as its carrier. Unlike traditional robotics, where perception and execution are decoupled modules stitched together, an embodied agent must simultaneously interpret human instructions, fuse low-level sensory data from multimodal sensors, and generate strategies adapted to its own mechanical constraints, correcting itself in real time through feedback.</p>
<p>Mathematically, the authors formalize this as a partially observable Markov decision process, a framework that acknowledges a sobering truth: a real robot almost never knows the true state of the world. Instead, it must maintain a probability distribution over possible states, updated with every observation and action. Where classical robot learning relied on tidy, low-dimensional state vectors, embodied manipulation confronts megapixel images, open-vocabulary language commands, tactile readings, and force signals all at once, forming a vast composite state space. Policies are therefore parameterized as deep neural networks trained to map this torrent of perception directly onto motor commands, maximizing expected cumulative reward over time.</p>
<p>What would an ideal embodied manipulator look like? The review identifies six hallmarks. It must achieve consistent multimodal perception, aligning vision, language, and haptics in a unified semantic space. It needs comprehensive multimodal understanding, the way large language models absorb web-scale knowledge. It must generalize across tasks and adapt zero-shot to unseen objects and scenes. It requires spatial intelligence, reasoning in three dimensions about geometry and the temporal consequences of its own actions. It needs basic physical commonsense, an intuitive grasp that objects fall, slide, and deform. And ultimately it must evolve, self-planning and self-correcting as environments shift.</p>
<p>Two competing technical philosophies currently dominate the field. The data-driven route treats manipulation as imitation at scale: collect enormous datasets of expert demonstrations, then train end-to-end vision-language-action models, or VLAs, that map what the robot sees and hears directly into what it does. Landmark systems such as RT-1, RT-2, PaLM-E, and the open-source OpenVLA exemplify this approach, and datasets like X-Embodiment, which aggregates a million-scale corpus spanning 22 robot morphologies and 527 skills, provide the fuel. Diffusion models, borrowed from image generation, have proven surprisingly effective here: the Diffusion Policy framework refines random noise into smooth, continuous action sequences, and successors like RDT-1B extend the idea to bimanual, cross-dataset learning.</p>
<p>The model-driven camp counters that data alone cannot carry robots through the open world. Researchers in this tradition build world models, internal simulations that let a robot imagine the consequences of its actions before committing to them. DayDreamer demonstrated that real robots can learn manipulation skills entirely through such imagined rollouts, updating their world model from live experience. Others inject the vast knowledge of multimodal large language models into the control loop: SayCan uses a language model to score which atomic skills actually serve a spoken instruction, while systems like ReKep generate spatial constraint functions from keypoints to solve manipulation trajectories without any task-specific training. Reinforcement learning adds a third pillar, fine-tuning imitation-trained policies through real-world trial and error, with frameworks like ConRFT balancing sample efficiency against safe execution.</p>
<p>Behind both paradigms lies a rapidly maturing infrastructure. Low-cost teleoperation rigs such as ALOHA and the handheld UMI gripper have slashed the price of collecting dexterous demonstration data, while Stanford&#8217;s HumanPlus tracks whole-body human motion to teach humanoids by shadowing. On the simulation side, GPU-accelerated engines like Isaac Gym and MuJoCo allow thousands of training environments to run in parallel, and newer frameworks such as Genesis and ManiSkill3 push toward photorealistic, four-dimensional worlds where policies can be trained before ever touching hardware. Generative techniques are now synthesizing data outright: RoboGen proposes its own skills and builds simulation scenes to practice them, while DemoGen mathematically replans existing trajectories to multiply datasets without a single new demonstration.</p>
<p>Yet the review is candid about how far robots remain from human dexterity. Vision-based policies depend on dense camera arrays that turn real workplaces into idealized laboratories, and a single modality can be catastrophically fooled; a robot wiping a table cannot feel whether it is pressing down or merely sliding the cloth. Contact-rich tasks like turning keys or opening valves expose the weakness of position-servo control that ignores force dynamics, and force sensors remain too expensive for mass deployment. Generalization is fragile, with minor lighting changes able to collapse performance, and full autonomy remains elusive. There is also a computational squeeze: the large models that perform best are precisely the ones too heavy for the edge processors a robot can physically carry, motivating compact architectures like SmolVLA, which achieves tenfold faster inference, and 1-bit compressed models like BitVLA.</p>
<p>The authors close with a plea that models and data should be treated as symbiotic rather than rival paradigms: data supplies the empirical knowledge that covers the world&#8217;s unpredictability, while models supply the interpretable framework that makes that knowledge usable. They also flag safety and ethics as unsolved frontiers, from concealed sim-to-real risks and adversarial attacks on embodied decision-makers to questions of job displacement and liability when autonomous machines cause harm. If the field&#8217;s trajectory holds, the milestone that matters may not be another benchmark score but the first robot that folds laundry in a stranger&#8217;s home as confidently as in its training lab, a quiet proof that intelligence, as Turing suspected, was always something a body does.</p>
<p><strong>Subject of Research:</strong> Embodied intelligence for robotic manipulation, covering data-driven and model-driven approaches, their supporting infrastructure, and open challenges</p>
<p><strong>Article Title:</strong> Embodied intelligence for robot manipulation: development and challenges</p>
<p><strong>Article References:</strong> Song, H., Wang, L., Qiao, X., Chen, Y., Sun, D., &amp; Sun, Z. (2025). Embodied intelligence for robot manipulation: development and challenges. <em>Vicinagearth, 2</em>(1), Article 8. <a href="https://doi.org/10.1007/s44336-025-00020-1" rel="noopener noreferrer">https://doi.org/10.1007/s44336-025-00020-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-025-00020-1" rel="noopener noreferrer">10.1007/s44336-025-00020-1</a></p>
<p><strong>Keywords:</strong> embodied intelligence, robot manipulation, vision-language-action models, world models, reinforcement learning, imitation learning, diffusion policy, multimodal perception, sim-to-real transfer, artificial general intelligence, humanoid robots, robot datasets</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">218654</post-id>	</item>
	</channel>
</rss>
