<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>imitation learning &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/imitation-learning/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 19:51:40 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>imitation learning &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Robots That Learn by Touch: How Embodied Intelligence Is Rewriting Machine Manipulation</title>
		<link>https://scienmag.com/robots-that-learn-by-touch-how-embodied-intelligence-is-rewriting-machine-manipulation/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 19:51:40 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[advancements in robotic touch and perception]]></category>
		<category><![CDATA[artificial general intelligence]]></category>
		<category><![CDATA[diffusion policy]]></category>
		<category><![CDATA[embodied AI in unstructured environments]]></category>
		<category><![CDATA[embodied intelligence]]></category>
		<category><![CDATA[embodied intelligence in robotic manipulation]]></category>
		<category><![CDATA[history of embodied intelligence]]></category>
		<category><![CDATA[humanoid robots]]></category>
		<category><![CDATA[imitation learning]]></category>
		<category><![CDATA[machine learning for physical tasks]]></category>
		<category><![CDATA[multimodal perception]]></category>
		<category><![CDATA[neural networks for robotic manipulation]]></category>
		<category><![CDATA[physical cognition in robotics]]></category>
		<category><![CDATA[physical interaction in robotics]]></category>
		<category><![CDATA[real-world robotic applications]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[robot datasets]]></category>
		<category><![CDATA[robot manipulation]]></category>
		<category><![CDATA[robotic sensory-motor integration]]></category>
		<category><![CDATA[robots grasping and manipulating objects]]></category>
		<category><![CDATA[sim-to-real transfer]]></category>
		<category><![CDATA[tactile sensing in robots]]></category>
		<category><![CDATA[vision-language-action models]]></category>
		<category><![CDATA[world models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=218654</guid>

					<description><![CDATA[A new systematic review maps how embodied intelligence, from vision-language-action models to world models, is transforming robotic manipulation and what obstacles remain before robots achieve human-like dexterity.]]></description>
										<content:encoded><![CDATA[<p>For more than seventy years, artificial intelligence has lived mostly in the abstract. Machines could beat grandmasters at chess and generate fluent prose, yet they struggled to pick up a wine glass without crushing it. A sweeping new review published in the journal Vicinagearth argues that this gap between digital brilliance and physical clumsiness is now closing, and that the key is a concept researchers call embodied intelligence: the idea that genuine understanding emerges only when a mind is wired directly into a body that senses and acts in the real world. The paper, led by Honghao Song and Zhe Sun of Northwestern Polytechnical University and colleagues, offers the first systematic map of how embodied intelligence is transforming robotic manipulation, the ability of machines to grasp, twist, fold, and assemble objects in unstructured environments.</p>
<p>The intellectual roots of the field trace back to Alan Turing, who in 1950 suggested that intelligence is not merely abstract computation but something demonstrated through dynamic interaction between a body and its environment. The review frames that insight as a theoretical foundation: a physical carrier is a necessary prerequisite for intelligence to operate in the physical world. In practice, the authors define embodied manipulation as a closed-loop process that takes embodied cognition as its engine and a physical robot as its carrier. Unlike traditional robotics, where perception and execution are decoupled modules stitched together, an embodied agent must simultaneously interpret human instructions, fuse low-level sensory data from multimodal sensors, and generate strategies adapted to its own mechanical constraints, correcting itself in real time through feedback.</p>
<p>Mathematically, the authors formalize this as a partially observable Markov decision process, a framework that acknowledges a sobering truth: a real robot almost never knows the true state of the world. Instead, it must maintain a probability distribution over possible states, updated with every observation and action. Where classical robot learning relied on tidy, low-dimensional state vectors, embodied manipulation confronts megapixel images, open-vocabulary language commands, tactile readings, and force signals all at once, forming a vast composite state space. Policies are therefore parameterized as deep neural networks trained to map this torrent of perception directly onto motor commands, maximizing expected cumulative reward over time.</p>
<p>What would an ideal embodied manipulator look like? The review identifies six hallmarks. It must achieve consistent multimodal perception, aligning vision, language, and haptics in a unified semantic space. It needs comprehensive multimodal understanding, the way large language models absorb web-scale knowledge. It must generalize across tasks and adapt zero-shot to unseen objects and scenes. It requires spatial intelligence, reasoning in three dimensions about geometry and the temporal consequences of its own actions. It needs basic physical commonsense, an intuitive grasp that objects fall, slide, and deform. And ultimately it must evolve, self-planning and self-correcting as environments shift.</p>
<p>Two competing technical philosophies currently dominate the field. The data-driven route treats manipulation as imitation at scale: collect enormous datasets of expert demonstrations, then train end-to-end vision-language-action models, or VLAs, that map what the robot sees and hears directly into what it does. Landmark systems such as RT-1, RT-2, PaLM-E, and the open-source OpenVLA exemplify this approach, and datasets like X-Embodiment, which aggregates a million-scale corpus spanning 22 robot morphologies and 527 skills, provide the fuel. Diffusion models, borrowed from image generation, have proven surprisingly effective here: the Diffusion Policy framework refines random noise into smooth, continuous action sequences, and successors like RDT-1B extend the idea to bimanual, cross-dataset learning.</p>
<p>The model-driven camp counters that data alone cannot carry robots through the open world. Researchers in this tradition build world models, internal simulations that let a robot imagine the consequences of its actions before committing to them. DayDreamer demonstrated that real robots can learn manipulation skills entirely through such imagined rollouts, updating their world model from live experience. Others inject the vast knowledge of multimodal large language models into the control loop: SayCan uses a language model to score which atomic skills actually serve a spoken instruction, while systems like ReKep generate spatial constraint functions from keypoints to solve manipulation trajectories without any task-specific training. Reinforcement learning adds a third pillar, fine-tuning imitation-trained policies through real-world trial and error, with frameworks like ConRFT balancing sample efficiency against safe execution.</p>
<p>Behind both paradigms lies a rapidly maturing infrastructure. Low-cost teleoperation rigs such as ALOHA and the handheld UMI gripper have slashed the price of collecting dexterous demonstration data, while Stanford&#8217;s HumanPlus tracks whole-body human motion to teach humanoids by shadowing. On the simulation side, GPU-accelerated engines like Isaac Gym and MuJoCo allow thousands of training environments to run in parallel, and newer frameworks such as Genesis and ManiSkill3 push toward photorealistic, four-dimensional worlds where policies can be trained before ever touching hardware. Generative techniques are now synthesizing data outright: RoboGen proposes its own skills and builds simulation scenes to practice them, while DemoGen mathematically replans existing trajectories to multiply datasets without a single new demonstration.</p>
<p>Yet the review is candid about how far robots remain from human dexterity. Vision-based policies depend on dense camera arrays that turn real workplaces into idealized laboratories, and a single modality can be catastrophically fooled; a robot wiping a table cannot feel whether it is pressing down or merely sliding the cloth. Contact-rich tasks like turning keys or opening valves expose the weakness of position-servo control that ignores force dynamics, and force sensors remain too expensive for mass deployment. Generalization is fragile, with minor lighting changes able to collapse performance, and full autonomy remains elusive. There is also a computational squeeze: the large models that perform best are precisely the ones too heavy for the edge processors a robot can physically carry, motivating compact architectures like SmolVLA, which achieves tenfold faster inference, and 1-bit compressed models like BitVLA.</p>
<p>The authors close with a plea that models and data should be treated as symbiotic rather than rival paradigms: data supplies the empirical knowledge that covers the world&#8217;s unpredictability, while models supply the interpretable framework that makes that knowledge usable. They also flag safety and ethics as unsolved frontiers, from concealed sim-to-real risks and adversarial attacks on embodied decision-makers to questions of job displacement and liability when autonomous machines cause harm. If the field&#8217;s trajectory holds, the milestone that matters may not be another benchmark score but the first robot that folds laundry in a stranger&#8217;s home as confidently as in its training lab, a quiet proof that intelligence, as Turing suspected, was always something a body does.</p>
<p><strong>Subject of Research:</strong> Embodied intelligence for robotic manipulation, covering data-driven and model-driven approaches, their supporting infrastructure, and open challenges</p>
<p><strong>Article Title:</strong> Embodied intelligence for robot manipulation: development and challenges</p>
<p><strong>Article References:</strong> Song, H., Wang, L., Qiao, X., Chen, Y., Sun, D., &amp; Sun, Z. (2025). Embodied intelligence for robot manipulation: development and challenges. <em>Vicinagearth, 2</em>(1), Article 8. <a href="https://doi.org/10.1007/s44336-025-00020-1" rel="noopener noreferrer">https://doi.org/10.1007/s44336-025-00020-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-025-00020-1" rel="noopener noreferrer">10.1007/s44336-025-00020-1</a></p>
<p><strong>Keywords:</strong> embodied intelligence, robot manipulation, vision-language-action models, world models, reinforcement learning, imitation learning, diffusion policy, multimodal perception, sim-to-real transfer, artificial general intelligence, humanoid robots, robot datasets</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">218654</post-id>	</item>
		<item>
		<title>Virtual Expert Teaches AI to Steer Ultrasound Probes for Liver Scans</title>
		<link>https://scienmag.com/virtual-expert-teaches-ai-to-steer-ultrasound-probes-for-liver-scans/</link>
		
		<dc:creator><![CDATA[Ophelia Keating]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 16:53:17 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[advancements in robotic ultrasound technology]]></category>
		<category><![CDATA[AI-guided ultrasound probe navigation for liver volumetric imaging]]></category>
		<category><![CDATA[anatomy-aware representation]]></category>
		<category><![CDATA[automated robotic ultrasound probe control]]></category>
		<category><![CDATA[autonomous liver scan acquisition]]></category>
		<category><![CDATA[computer-assisted radiology and surgery]]></category>
		<category><![CDATA[computer-assisted surgery]]></category>
		<category><![CDATA[data augmentation]]></category>
		<category><![CDATA[deep learning for ultrasound probe navigation]]></category>
		<category><![CDATA[imitation learning]]></category>
		<category><![CDATA[intercostal window targeting in liver scans]]></category>
		<category><![CDATA[liver imaging]]></category>
		<category><![CDATA[machine learning in ultrasound imaging]]></category>
		<category><![CDATA[medical imaging AI]]></category>
		<category><![CDATA[operator-independent ultrasound imaging]]></category>
		<category><![CDATA[probe guidance]]></category>
		<category><![CDATA[robotic ultrasound]]></category>
		<category><![CDATA[sensor-free ultrasound probe guidance]]></category>
		<category><![CDATA[simulated training for ultrasound probe positioning]]></category>
		<category><![CDATA[target view localization]]></category>
		<category><![CDATA[ultrasound simulation]]></category>
		<category><![CDATA[virtual expert]]></category>
		<category><![CDATA[virtual expert-assisted ultrasound scanning]]></category>
		<category><![CDATA[volumetric liver ultrasound]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=196627</guid>

					<description><![CDATA[Researchers have developed an AI framework that learns to automatically guide an ultrasound probe to liver target views by imitating a virtual expert in a realistic simulated scanning environment.]]></description>
										<content:encoded><![CDATA[<p>Volumetric ultrasound of the liver is one of the most demanding routines in clinical imaging. Unlike a snapshot radiograph, a volumetric acquisition requires the sonographer to sweep and hold the probe along precisely chosen intercostal windows, threading the imaging plane between ribs and around bowel gas until the liver is captured in standardized target views. The quality of the resulting three-dimensional data depends heavily on operator expertise, and even experienced sonographers vary in how reliably they can localize target planes. A new study published in the International Journal of Computer Assisted Radiology and Surgery describes an automatic probe guidance framework that learns to perform this task by imitating a virtual expert inside a simulated scanning environment, removing the need for manual trajectory demonstrations or extra sensing hardware.</p>
<p>The research, led by Taiyu Han, Hanying Liang, Guochen Ning of Tsinghua University and Septimiu E. Salcudean of the University of British Columbia, together with colleagues, addresses a core bottleneck in the emerging field of robotic and computer-assisted ultrasound. Existing approaches to autonomous probe navigation typically rely on either large volumes of human demonstration data or additional tracking sensors mounted on the probe and patient. Both requirements are costly and difficult to satisfy in busy clinics. The new framework instead extracts guidance policies from a virtual expert whose demonstrations are generated entirely within simulation, using only the kind of target views that are normally available in clinical practice.</p>
<p>Technically, the pipeline begins with cross-modal medical images, which are segmented and processed to construct a simulated ultrasound scanning environment. A hybrid ultrasound simulator then renders realistic images through a combination of two complementary mechanisms. Physics-based ray casting models how acoustic beams interact with tissue interfaces, capturing the geometric consequences of probe motion, while generation-based image synthesis adds the textural realism of speckle, shadowing and acoustic artifacts that characterize real B-mode images. The result is a stream of anatomically consistent and acoustically plausible ultrasound frames that respond faithfully to changes in probe pose, giving the learning algorithm a faithful proxy for the real imaging task.</p>
<p>Within this simulated environment, optimal scanning trajectories are generated automatically based solely on target views. The virtual expert defines the ideal probe pose and path that connect an arbitrary starting position to the standardized plane needed for volumetric liver acquisition, and the learning system is trained to reproduce these decisions from the images it observes. This is an imitation learning formulation: rather than discovering a policy through slow trial-and-error reinforcement learning, the model directly learns to map observed ultrasound content to the corrective probe movements a skilled operator would make. Because the demonstrations are synthesized, the authors can generate them at scale without ever asking a clinician to annotate trajectories.</p>
<p>Robustness, however, is the central challenge for any image-based guidance system that must eventually operate on real patients. The researchers introduce pose-level and image-level data augmentation during training, exposing the model to systematic variations in probe orientation, anatomical appearance and imaging conditions so that its learned policy does not overfit the particular characteristics of the simulated data. In parallel, they encode the observed ultrasound images into an anatomy-aware state representation tailored to intercostal liver scanning. Rather than treating every pixel pattern as equally informative, this representation emphasizes the anatomical structures that matter for navigation, such as rib shadows, hepatic vessels and the diaphragm, allowing the network to infer where the probe sits relative to the target plane even when the raw image is ambiguous.</p>
<p>The evaluation combined experiments in simulation with tests on real clinical data. Compared with baseline models and ablated variants in which individual components were removed, the proposed framework achieved more accurate and more stable localization of target views for volumetric liver ultrasound acquisition. The ablation studies underline how each design choice contributes: the augmented training data improved generalization across different anatomical conditions, while the anatomy-aware representation reduced rib interference, a persistent failure mode in which the probe drifts behind a rib and loses sight of the liver entirely. The method also increased liver coverage, meaning the guided sweep captured more of the organ in a single acquisition.</p>
<p>The clinical motivation for automating this task is substantial. Volumetric liver ultrasound plays an important role in diagnosis and monitoring, including the assessment of non-alcoholic fatty liver disease and liver fibrosis, conditions with an enormous global burden. Yet target view localization remains highly operator dependent, and variability between sonographers can affect the reproducibility of quantitative measurements such as shear wave speed. A guidance system that reliably steers the probe into standardized planes could make volumetric acquisitions more consistent across operators and centers, shorten examination times, and open the door to screening protocols that do not require a highly specialized sonographer at every station.</p>
<p>What makes the approach particularly practical is its data efficiency and hardware minimalism. Because the policy learns from a virtual expert rather than from recorded human scans, and because it bases its decisions on the ultrasound image stream alone, the framework requires neither manual trajectory annotations nor additional electromagnetic or optical tracking sensors. That combination matters for integration into computer-assisted and robotic ultrasound systems, where the cost and complexity of peripheral hardware often determine whether a laboratory prototype can become a clinical product. The authors note that their results suggest strong potential for such integration, positioning the framework as a step toward intelligent robotic sonographers that can assist or, in some workflows, partially replace manual probe positioning.</p>
<p>The work also reflects a broader trend in medical imaging AI: the shift from learning on scarce, expensive real-world demonstrations toward learning in high-fidelity simulation and transferring to reality. The hybrid simulator strategy, blending physics-based rendering with learned image synthesis, is designed precisely to narrow the gap between synthetic training images and the noisy, artifact-laden images encountered at the bedside. Combined with deliberate augmentation and anatomy-informed representations, the framework demonstrates that simulated expertise can translate into accurate, stable probe control on real clinical data, at least for the structured, well-defined navigation task of intercostal liver scanning.</p>
<p>Challenges remain before such systems reach routine use. Real tissues deform, patients move and breathe, and body habitus varies widely, all of which stress any image-guided controller. Still, the study reports that the framework maintains robustness across different anatomical conditions, and its reliance on standard target views means it can be deployed with the image content clinicians already produce. As autonomous and semi-autonomous ultrasound platforms mature, frameworks like this one, which learn from virtual experts in realistic simulated worlds, may define how the next generation of imaging systems acquires its skills, turning the craft of probe handling into a reproducible computational capability.</p>
<p><strong>Subject of Research:</strong> Automatic probe guidance for volumetric liver ultrasound acquisition using imitation learning from a virtual expert</p>
<p><strong>Article Title:</strong> Automatic probe guidance for volumetric liver ultrasound acquisition via imitation learning from a virtual expert</p>
<p><strong>Article References:</strong> Automatic probe guidance for volumetric liver ultrasound acquisition via imitation learning from a virtual expert. (n.d.). <a href="https://doi.org/10.1007/s11548-026-03787-w" rel="noopener noreferrer">https://doi.org/10.1007/s11548-026-03787-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11548-026-03787-w" rel="noopener noreferrer">10.1007/s11548-026-03787-w</a></p>
<p><strong>Keywords:</strong> volumetric liver ultrasound, probe guidance, imitation learning, virtual expert, ultrasound simulation, target view localization, robotic ultrasound, data augmentation, anatomy-aware representation, computer-assisted surgery, medical imaging AI, liver imaging</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">196627</post-id>	</item>
	</channel>
</rss>
