<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>user study &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/user-study/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 10 Oct 2026 00:32:59 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>user study &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns to Steal the Camera Work of Any Movie Clip and Apply It to Your Photos</title>
		<link>https://scienmag.com/ai-learns-to-steal-the-camera-work-of-any-movie-clip-and-apply-it-to-your-photos/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 10 Oct 2026 00:32:59 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-driven camera motion transfer]]></category>
		<category><![CDATA[applying movie camera techniques to photos]]></category>
		<category><![CDATA[camera motion transfer]]></category>
		<category><![CDATA[CameraScore metric]]></category>
		<category><![CDATA[cinema-inspired image enhancement]]></category>
		<category><![CDATA[cinematic camera movement synthesis]]></category>
		<category><![CDATA[creative tools]]></category>
		<category><![CDATA[diffusion models]]></category>
		<category><![CDATA[generative AI]]></category>
		<category><![CDATA[generative AI for camera work]]></category>
		<category><![CDATA[homography]]></category>
		<category><![CDATA[human-computer interaction]]></category>
		<category><![CDATA[image-to-video synthesis]]></category>
		<category><![CDATA[improving AI video generation with real camera moves]]></category>
		<category><![CDATA[LoRA fine-tuning]]></category>
		<category><![CDATA[motion transfer in AI-generated images]]></category>
		<category><![CDATA[natural cinematic shot replication]]></category>
		<category><![CDATA[non-technical camera motion application]]></category>
		<category><![CDATA[reference video camera motion extraction]]></category>
		<category><![CDATA[user study]]></category>
		<category><![CDATA[user-friendly filmic perspective editing]]></category>
		<category><![CDATA[video generation]]></category>
		<category><![CDATA[zero-shot learning]]></category>
		<category><![CDATA[zero-shot personalized camera control]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=256650</guid>

					<description><![CDATA[Researchers have developed a zero-shot AI framework that transfers the camera motion of a single reference video onto any static image, letting non-expert creators direct cinematic shots without 3D data or technical expertise.]]></description>
										<content:encoded><![CDATA[<p>Anyone who has tried to direct an AI video generator knows the frustration. You type &#8220;zoom in slowly&#8221; or &#8220;pan across the scene,&#8221; and the model delivers something close, but never quite the sweeping dolly shot or the tense, creeping push-in you saw in a favorite film. Camera motion, the invisible hand that guides perspective and emotion in cinema, has remained stubbornly out of reach for casual creators. Now a team of researchers from the University of Maryland and Dolby Laboratories has built a system that lets anyone borrow the exact camera movement from a single reference video and apply it to their own still image, with no 3D reconstruction, no trajectory files, and no technical expertise required.</p>
<p>The work, published in Discover Artificial Intelligence, addresses what the researchers call the &#8220;expressive gap&#8221; between a creator&#8217;s cinematic vision and the blunt instruments available in generative tools. Text prompts are too abstract to capture nuanced motion, while professional motion panels demand calibration skills and patience that most people simply do not have. The new framework, described as zero-shot personalized camera motion control, works by example: users supply a short clip whose camera behavior they admire, along with their own static image, and the system generates a video in which the user&#8217;s scene moves exactly the way the reference clip did.</p>
<p>Technically, the method is a two-phase pipeline built on top of a pretrained text-to-video diffusion model. In the first phase, the system performs inference-time optimization using two sets of Low-Rank Adaptation networks, or LoRAs, inserted into the model&#8217;s UNet architecture. Spatial LoRAs, placed in the spatial self-attention layers, learn the visual appearance of scenes, while temporal LoRAs, placed in the temporal self-attention layers, capture the camera motion itself. Cross-attention layers that bind text tokens to visual features remain frozen throughout. The training proceeds in two stages: first, the temporal LoRAs learn the motion dynamics of the full reference video while the spatial LoRAs learn generic appearance from randomly sampled single frames, discouraging overfitting to any one moment; second, the temporal LoRAs are frozen and the spatial LoRAs are re-tuned to represent the user&#8217;s target image.</p>
<p>The crucial innovation is an orthogonality regularizer that prevents these two sets of learned features from interfering with each other. When the spatial LoRAs are adapted to a new image, they risk overwriting shared low-level representations in a way that suppresses the frozen temporal LoRA&#8217;s motion signal at inference time. To prevent this, the researchers introduced a loss function that penalizes alignment between the spatial and temporal LoRA weight directions, comparing only the most significant singular directions obtained through truncated singular value decomposition. Minimizing this inner product encourages the two updates to occupy near-orthogonal subspaces, so that adapting the appearance of one scene does not erode the camera motion learned from another.</p>
<p>The second phase tackles a subtler problem: even a well-tuned diffusion model is not explicitly constrained by geometry, so the generated video can drift from the intended camera path. The researchers therefore added homography-guided inference, borrowing a classical computer vision technique that computes the planar transformation mapping one image onto another. Using SIFT keypoint matching and RANSAC, the system extracts frame-to-frame homography matrices from the reference video, warps the user&#8217;s image into a weak geometric estimate of the desired sequence, and then nudges the diffusion model&#8217;s predicted latents toward that estimate at every denoising step. This weak guidance keeps the output anchored to both the user&#8217;s scene and the reference motion without requiring any 3D reconstruction or camera pose estimation.</p>
<p>The choice of homographies over more elaborate 3D methods is deliberate. Tools like COLMAP, which recover camera trajectories from structure-from-motion, fail when a zoom effect is achieved purely by changing focal length rather than physically moving the camera, and they struggle with flat or textureless scenes. Homographies, by contrast, can be computed from almost any footage and capture the relative planar transformations, panning, tilting, and zooming, that define most cinematic camera work. The same insight underpins the team&#8217;s second contribution: a new evaluation metric called CameraScore, which measures the squared difference between homography matrices computed from consecutive frames of the reference and generated videos, providing a computationally cheap and scene-independent way to judge motion fidelity across structurally different scenes.</p>
<p>To test the system, the researchers curated a dataset of 680 reference video and user scene combinations, spanning pans, zooms, tilts, dolly shots, and 3D rotations drawn from movies, documentaries, and animation, paired with target scenes ranging from landscapes to urban settings. Quantitative comparisons against zero-shot baselines, including a naive homography approach, a DreamBooth-modified Tune-A-Video, and MotionDirector, showed the new method achieving the best trade-off between motion fidelity and scene preservation. Ablation experiments confirmed that all three components, user scene learning, the orthogonality loss, and homography guidance, are complementary and essential: removing any one of them degrades either the transferred motion or the integrity of the user&#8217;s scene.</p>
<p>Human evaluation proved even more striking. In a perceptual study with 72 participants recruited through online crowdsourcing, the method was preferred for camera motion fidelity in 90.45 percent of trials and for scene preservation in 70.31 percent of trials, with statistical analysis using generalized estimating equations confirming robust preferences across all criteria. A second, task-based interaction study with 12 participants compared three interface paradigms: pure text prompts using Google&#8217;s Flow with Veo3, preset-based motion panels using Veo2, and the reference-video-driven workflow. The reference-driven approach required an average of only 1.11 iterations and about 159 seconds per task, compared with nearly 9 minutes for text-based prompting, and it produced significantly lower cognitive load on every dimension of the NASA Task Load Index, along with dramatically higher satisfaction and preference scores.</p>
<p>The implications extend well beyond convenience. The researchers frame the work as a step toward democratizing cinematic production, noting that the expressive gap currently restricts access to advanced visual storytelling for small-scale industries, educators, and casual creators, a concern aligned with inclusive innovation goals. The system runs on a single A5000 GPU in roughly ten minutes per reference-video and image pair, comparable to existing personalization workflows, and its modular design means it could plug into stronger video backbones as they emerge, though adapting to modern architectures with unified spatiotemporal attention requires additional head-identification steps. Limitations remain: the homography guidance can be fooled by large moving foreground objects, textureless scenes provide too few keypoints, and the current interface offers no way to blend multiple reference motions or edit trajectories interactively. The team plans to address these gaps with multi-reference fusion, interactive trajectory editing, and broader studies with professional filmmakers. For now, the message is clear: the language of the camera, once the exclusive dialect of cinematographers, is becoming something anyone can speak simply by showing the machine what they mean.</p>
<p><strong>Subject of Research:</strong> Zero-shot camera motion transfer from reference videos to static images using diffusion models</p>
<p><strong>Article Title:</strong> Zero-shot personalized camera motion control for image-to-video synthesis</p>
<p><strong>Article References:</strong> Guhan, P., Kothandaraman, D., Lee, G., Huang, T.-W., Su, G.-M., &amp; Manocha, D. (2026). Zero-shot personalized camera motion control for image-to-video synthesis. <em>Discover Artificial Intelligence, 6</em>(1), Article 1377. <a href="https://doi.org/10.1007/s44163-026-02212-0" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02212-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02212-0" rel="noopener noreferrer">10.1007/s44163-026-02212-0</a></p>
<p><strong>Keywords:</strong> camera motion transfer, image-to-video synthesis, diffusion models, LoRA fine-tuning, homography, generative AI, video generation, human-computer interaction, creative tools, user study, CameraScore metric, zero-shot learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">256650</post-id>	</item>
		<item>
		<title>New Rendering Method Brings Realistic Tree Canopies to Real Time on a Memory Budget</title>
		<link>https://scienmag.com/new-rendering-method-brings-realistic-tree-canopies-to-real-time-on-a-memory-budget/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 07:05:05 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[3D Gaussian splatting for vegetation]]></category>
		<category><![CDATA[advanced rendering techniques for virtual landscapes]]></category>
		<category><![CDATA[architectural visualization with trees]]></category>
		<category><![CDATA[computer graphics]]></category>
		<category><![CDATA[dense realistic vegetation modeling]]></category>
		<category><![CDATA[diffuse lighting]]></category>
		<category><![CDATA[high-fidelity tree visualization on limited memory]]></category>
		<category><![CDATA[image-based modeling]]></category>
		<category><![CDATA[innovative tree canopy synthesis methods]]></category>
		<category><![CDATA[interactive urban planning tools]]></category>
		<category><![CDATA[landscape visualization]]></category>
		<category><![CDATA[light fields]]></category>
		<category><![CDATA[memory efficiency]]></category>
		<category><![CDATA[memory-efficient landscape visualization]]></category>
		<category><![CDATA[multimedia tools for landscape modeling]]></category>
		<category><![CDATA[neural radiance fields for foliage]]></category>
		<category><![CDATA[plant modeling]]></category>
		<category><![CDATA[real-time rendering]]></category>
		<category><![CDATA[Real-time tree canopy rendering]]></category>
		<category><![CDATA[residual lighting]]></category>
		<category><![CDATA[resource-efficient virtual environment rendering]]></category>
		<category><![CDATA[texture compression]]></category>
		<category><![CDATA[tree canopy synthesis]]></category>
		<category><![CDATA[user study]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=234030</guid>

					<description><![CDATA[Researchers at NTUST have developed a real-time method that synthesizes dense, realistic tree canopies by fitting proxy semi-ellipsoids, quantizing leaf colors, and encoding view-dependent lighting as compact proxy-based light fields.]]></description>
										<content:encoded><![CDATA[<p>Trees are the great paradox of landscape visualization. In any architectural rendering, urban planning model, or virtual environment, they are indispensable: no scene reads as believable without them. Yet in interactive design tools they typically serve as visual reference points rather than objects of scrutiny, which means pouring gigabytes of geometry and texture data into every leaf is a waste of precious resources. A research team at National Taiwan University of Science and Technology now reports a method that closes this gap, synthesizing dense, highly realistic tree canopies in real time while consuming dramatically less memory than state-of-the-art reconstruction techniques. The work, published in Multimedia Tools and Applications, demonstrates that the appearance of a lush canopy can be preserved almost perfectly even when nearly all of its internal detail is thrown away.</p>
<p>The team, led by Chia-Hsing Chiu and Yi-Ling Chen of the Department of Computer Science and Information Engineering, set out to address a blind spot in the current literature. Modern approaches to capturing vegetation, from image-based tree modeling to neural radiance fields and 3D Gaussian splatting, have become remarkably good at replicating lifelike geometric and topological detail. What they have not done, the authors argue, is constrain their appetite. Memory usage and rendering cost in these methods are essentially unconstrained, and no existing scenario reconstructs canopies under tight resource budgets. For interactive landscape design, where dozens of trees may populate a scene and the viewer is constantly moving, that omission is decisive.</p>
<p>The core idea of the new method is a deliberate act of simplification guided by human perception. Rather than reconstructing every branch and leaf, the pipeline first fits the tree&#8217;s silhouette, observed from multiple viewpoints, with semi-ellipsoids that the authors call proxies. These proxies act as lightweight stand-ins for the canopy&#8217;s bulk. Exemplar leaves are then distributed quasi-randomly across the proxy shell, producing the aggregated appearance of leaf clusters while preserving memory. The insight is that a viewer at a distance perceives the canopy as a textured surface of clustered foliage, not as a collection of individually resolvable leaves, so the shell of the proxy is the only place where detail actually matters.</p>
<p>The appearance model itself is built on a physically grounded decomposition of light. Following the rendering equation, the light leaving a surface point can be split into a diffuse component, which depends only on position and is view-independent, and a residual component that gathers everything directional: specular reflection, sub-surface scattering, transmission, and related effects. The authors derive this decomposition formally in an appendix, showing that the diffuse term can be bounded and estimated from input photographs by choosing, for each leaf, the camera whose viewing direction best aligns with the surface normal. This image-based texturing step stabilizes the coloring and minimizes perspective distortion during back projection.</p>
<p>The crucial memory saving comes from collapsing the diffuse term to a single representative radiance per leaf. Because leaves are tiny relative to the canopy, their normal variations are negligible compared with the exterior normal of the whole crown, and human perception is far more sensitive to the global color distribution of the canopy than to local color shifts within a single leaf. The team therefore assumes each leaf shares one representative normal and occlusion value, then quantizes the diffuse radiance across the entire canopy into a small set of representative index colors using K-Means clustering. Each leaf stores only its quantized color index, compressed as the product of a user-given albedo texture and the estimated light integral, rather than a high-resolution texture patch.</p>
<p>View-dependent subtlety is handled by the residual term, which is where the visual richness of real foliage lives: glints of light bouncing between leaves, soft variations as the canopy is seen from different angles. Because environment lighting varies smoothly and leaves are distributed chaotically and roughly evenly inside the canopy, the residual varies smoothly across the proxy surface. The authors exploit this by averaging the available per-leaf residual samples, weighting each input view by how closely its outgoing direction matches the current viewing angle, and then aggregating leaf-based residuals into proxy-based residuals normalized by the proxy normal. The result is that pixel-based multi-view subtlety textures are encoded as proxy-based light fields, a compact representation that captures how the canopy&#8217;s appearance changes as the observer moves.</p>
<p>The final shading model is elegantly simple: the outgoing light from any leaf is approximated as the sum of a view-independent diffuse color per leaf and a view-dependent residual lighting per proxy. In practice this means the entire canopy&#8217;s appearance is compressed into a handful of small textures, one holding quantized leaf colors and the others holding the residual light fields, while the geometry is reduced to semi-ellipsoidal shells studded with quasi-randomly placed exemplar leaves. Everything is designed to be evaluated in a single pass on a GPU, which is what makes real-time rendering feasible even for dense, visually complex canopies.</p>
<p>To validate the approach, the researchers conducted a series of experiments comparing their method against state-of-the-art baselines, including commercial and academic tools for tree reconstruction and rendering such as SpeedTree, Context Capture, and procedural painting approaches, alongside neural rendering techniques. The evaluation data has been made publicly available through the group&#8217;s project website, and data to reproduce the evaluation is available on request. The results show that the proposed method achieves superior memory usage and comparable rendering efficiency relative to the baselines, while presenting a visual appearance similar to the ground-truth imagery. In other words, the method matches its heavier competitors where it counts for viewers and beats them decisively where it counts for hardware.</p>
<p>The team also ran a user study to assess how the synthesized canopies are perceived, applying established statistical procedures for comparing ranked judgments, including Friedman-style significance testing and the method of paired comparisons, tools long used in perceptual evaluation of graphics and tone mapping. The study supports the central perceptual claim of the paper: that quantizing leaf colors and compressing view-dependent effects into proxy light fields does not produce a noticeable degradation in perceived realism for the intended viewing conditions. This perceptual grounding is what separates the work from naive level-of-detail schemes that simply swap geometry for billboards, a strategy explored in earlier vegetation rendering research dating back to slicing and blending techniques from 2000 and adaptive billboard clouds from 2014.</p>
<p>The implications extend across several domains. For interactive landscape and architectural design, the method offers a way to populate scenes with photorealistic trees without the memory blowup that has historically forced designers to accept crude stand-ins. For virtual and augmented reality, where rendering budgets are measured in milliseconds and megabytes, the proxy-based light field representation is a natural fit. And for the broader graphics community, the paper is a reminder that the path to realism does not always run through more data. By starting from the physics of light transport, observing what human vision actually registers, and aggressively compressing everything else, the NTUST team has shown that a dense canopy of thousands of leaves can live comfortably within a real-time budget. The work was supported by grants from Taiwan&#8217;s National Science and Technology Council, and the authors report no competing interests.</p>
<p><strong>Subject of Research:</strong> Memory-efficient real-time synthesis and rendering of dense tree canopies for landscape visualization</p>
<p><strong>Article Title:</strong> Real–time appearance–driven memory–efficient dense canopy synthesis</p>
<p><strong>Article References:</strong> Chiu, C.-H., Lai, Y.-C., Chang, C.-W., Du, H.-Y., Tai, W.-K., &amp; Chen, Y.-L. (2026). Real–time appearance–driven memory–efficient dense canopy synthesis. <em>Multimedia Tools and Applications, 85</em>(9), Article 748. <a href="https://doi.org/10.1007/s11042-026-21807-4" rel="noopener noreferrer">https://doi.org/10.1007/s11042-026-21807-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11042-026-21807-4" rel="noopener noreferrer">10.1007/s11042-026-21807-4</a></p>
<p><strong>Keywords:</strong> computer graphics, tree canopy synthesis, image-based modeling, light fields, real-time rendering, memory efficiency, texture compression, plant modeling, landscape visualization, diffuse lighting, residual lighting, user study</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">234030</post-id>	</item>
		<item>
		<title>Robots Learn When You Feel Unsafe: New Framework Tunes Speed and Distance in Real Time</title>
		<link>https://scienmag.com/robots-learn-when-you-feel-unsafe-new-framework-tunes-speed-and-distance-in-real-time/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 17:22:04 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[active learning]]></category>
		<category><![CDATA[adaptive control]]></category>
		<category><![CDATA[adaptive robot speed control based on human proximity]]></category>
		<category><![CDATA[balancing safety and efficiency in autonomous systems]]></category>
		<category><![CDATA[Boston Dynamics Spot]]></category>
		<category><![CDATA[control barrier functions]]></category>
		<category><![CDATA[control barrier functions in robotics]]></category>
		<category><![CDATA[dynamic safety control in robotics]]></category>
		<category><![CDATA[human-aware motion planning]]></category>
		<category><![CDATA[human-centered robot navigation]]></category>
		<category><![CDATA[human-robot interaction]]></category>
		<category><![CDATA[human-robot interaction comfort]]></category>
		<category><![CDATA[improving robot acceptance in shared workspaces]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[mobile robots]]></category>
		<category><![CDATA[model predictive control]]></category>
		<category><![CDATA[perceived safety]]></category>
		<category><![CDATA[PERSCO framework for robot speed and distance tuning]]></category>
		<category><![CDATA[real-time robot behavior adaptation]]></category>
		<category><![CDATA[real-time safety learning algorithms]]></category>
		<category><![CDATA[Robotics safety perception]]></category>
		<category><![CDATA[social robotics]]></category>
		<category><![CDATA[subjective safety versus objective safety in automation]]></category>
		<category><![CDATA[user study]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=223550</guid>

					<description><![CDATA[Researchers at Georgia Tech have developed PERSCO, a control framework that lets mobile robots learn in real time how fast and how close people are comfortable with, significantly improving perceived safety in a 54-person study.]]></description>
										<content:encoded><![CDATA[<p>A robot can be perfectly safe by every engineering metric and still terrify the people around it. A mobile platform that never collides with anyone but barrels past workers at high speed, inches from their bodies, will feel threatening no matter what the collision statistics say. Conversely, a robot that creeps along at a snail&#8217;s pace to avoid alarming anyone may be so sluggish that it fails its task entirely. This gap between objective safety and subjective comfort has long been a blind spot in robotics, and a new framework presented in the journal Autonomous Robots aims to close it by letting robots learn, in real time, exactly how fast and how close people are willing to tolerate them.</p>
<p>The framework, called PERSCO, was developed by Sanne van Waveren, Zulfiqar Zaidi, and Matthew Gombolay at the Georgia Institute of Technology. Its central insight is that perceived safety can be treated as a tunable control problem rather than a fixed design choice. The researchers build on control barrier functions, or CBFs, mathematical constructs that guarantee a robot stays within a set of safe states by constraining its control inputs at every time step. Traditional CBFs enforce physical safety with static parameters that never change during an interaction. PERSCO instead parameterizes the CBF with two variables that directly shape how the robot&#8217;s behavior feels to nearby humans: the minimum distance the robot must keep from each person, and the maximum deceleration it is allowed to use, which in turn caps how fast it may approach anyone.</p>
<p>These two parameters have intuitive physical meaning. The distance parameter defines an intimate space around each person that the robot may never enter. The deceleration parameter determines the stopping distance the robot must be able to achieve at its current speed; a robot permitted stronger braking can safely travel faster, because it can halt in a shorter distance. By adjusting the pair, the controller can make the robot behave anywhere from maximally cautious to maximally assertive. The key question is which combination a given person actually perceives as safe, and that is something no designer can hard-code in advance, because perceptions vary widely between individuals and contexts.</p>
<p>To answer it, PERSCO treats the problem as active learning. Each person carries a hidden perceived safety function that maps any parameter pair to a judgment of safe or unsafe. The robot cannot observe this function directly, so it maintains a surrogate model, a classifier trained on feedback, and probes the boundary between safe and unsafe parameter regions. Crucially, the researchers designed the feedback to be as unobtrusive as possible. Rather than asking people to fill out Likert scales mid-task, PERSCO adopts a principle of perceived safe until proven unsafe: humans only signal when they feel uncomfortable, using a simple visual cue, in this study a handheld AprilTag sign raised toward a camera. Silence is treated as implicit safe feedback, provided the robot has logged enough close encounters with that person without any complaint.</p>
<p>The learning algorithm is engineered to minimize how often people must intervene. When unsafe feedback arrives, the robot updates its classifier and selects the next candidate parameters using a novel sampling strategy that balances two criteria: entropy, which targets regions where the model is most uncertain about where the boundary lies, and diversity, which favors candidates far from previously tested ones. Importantly, the sampler only considers parameters the model predicts to be safe, so the robot never deliberately behaves in a way it believes will alarm the human. When no unsafe feedback arrives after repeated close encounters, the robot gradually relaxes its parameters, stepping toward the least restrictive pair on the safety boundary, which maximizes task efficiency while remaining at the edge of what the person tolerates.</p>
<p>All of this runs inside a model predictive control loop with a 0.1-second time step, where the perceived safety CBF is enforced over the entire planning horizon while accounting for predicted human motion. A separate, unchanging physical safety CBF with the most aggressive parameters guarantees collision avoidance at all times, so no matter how the learned parameters evolve, the robot can always brake to a standstill before reaching anyone. When parameter updates suddenly tighten the constraints and the robot temporarily finds itself outside the new safe set, a gradual recovery strategy using a slack variable steers it back smoothly; in simulation this reduced jerk by 25 percent and angular acceleration by a factor of 3.6 compared to abrupt corrections, avoiding the jarring sidesteps that quick recovery methods would produce.</p>
<p>Simulation experiments validated the technical choices. Among three candidate classifiers, a support vector classifier with a radial basis function kernel proved the clear winner, updating in about 1.3 milliseconds on average, fast enough for real-time control, while the neural network and Gaussian process alternatives exceeded the control loop&#8217;s time budget. Against a battery of classical active learning baselines and black-box optimizers, including multi-armed bandits and Bayesian optimization, PERSCO sampling achieved high accuracy in recovering ground-truth safety parameters while producing the lowest ratio of unsafe feedback events. In a simulated workplace with three moving pedestrians, the system converged to near-optimal parameters in roughly eight minutes, both with and without noise injected into the feedback.</p>
<p>The decisive test came with real humans. Fifty-four participants, organized into eighteen groups of three, performed a workplace-inspired assembly task, walking between workstations to place LED pins on breadboards while a Boston Dynamics Spot robot navigated the space autonomously, covering 7,074 meters over the course of the study. Participants experienced three conditions: individual adaptation, in which the robot learned separate parameters for each person; collective adaptation, in which one shared parameter set was updated from anyone&#8217;s feedback; and an adversarial condition, in which the robot responded to feedback by becoming more aggressive rather than more cautious. The adversarial condition served as a control to test whether adaptation itself, or only feedback-aligned adaptation, improves how safe people feel.</p>
<p>The results were striking. Both aligned conditions significantly outperformed the adversarial one on perceived safety, comfort, and anxiety, all with p-values below .001 and large effect sizes. Collective adaptation scored highest overall, and participants raised their feedback signs significantly less often under collective updates, suggesting that people benefit from feedback provided by their teammates. A mediation analysis revealed that the effect of condition on perceived safety was fully mediated by the average size of the robot&#8217;s safety boundary, meaning the psychological benefit flowed directly from the geometric changes in the robot&#8217;s enforced constraints. Notably, individual adaptation offered no task-performance advantage over collective adaptation, contrary to the researchers&#8217; hypothesis, possibly because fewer parameter changes allowed the robot to plan more consistently.</p>
<p>The authors are candid about limitations: the study took place in a controlled environment, sessions were capped at twelve minutes so parameters did not always converge, and treating silence as safe feedback assumes people are attentive enough to complain when they feel threatened. Still, the work marks a meaningful shift in how roboticists think about safety. By framing perceived safety as a quantity that can be measured, learned, and optimized alongside task performance, PERSCO argues that true safety encompasses psychological well-being, not just the absence of collisions. As robots move into warehouses, hospitals, and factories, the systems that earn human trust may be the ones that ask, in effect, how their presence feels, and adjust accordingly.</p>
<p><strong>Subject of Research:</strong> Perceived-safe control of mobile robots using active learning from human feedback</p>
<p><strong>Article Title:</strong> PERSCO: Perceived safe control of mobile robots in human groups with active learning</p>
<p><strong>Article References:</strong> van Waveren, S., Zaidi, Z., &amp; Gombolay, M. (2026). PERSCO: Perceived safe control of mobile robots in human groups with active learning. <em>Autonomous Robots, 50</em>(4), Article 43. <a href="https://doi.org/10.1007/s10514-026-10262-7" rel="noopener noreferrer">https://doi.org/10.1007/s10514-026-10262-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10514-026-10262-7" rel="noopener noreferrer">10.1007/s10514-026-10262-7</a></p>
<p><strong>Keywords:</strong> perceived safety, human-robot interaction, control barrier functions, active learning, mobile robots, model predictive control, social robotics, human-aware motion planning, adaptive control, Boston Dynamics Spot, user study, machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">223550</post-id>	</item>
		<item>
		<title>Robots Learn Faster When Humans Show Them Why, Not Just What</title>
		<link>https://scienmag.com/robots-learn-faster-when-humans-show-them-why-not-just-what/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:35:17 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[autonomous robots]]></category>
		<category><![CDATA[CALVIN benchmark]]></category>
		<category><![CDATA[causal confusion]]></category>
		<category><![CDATA[causal confusion in machine learning]]></category>
		<category><![CDATA[causal understanding in robotics]]></category>
		<category><![CDATA[demonstration-based robot training]]></category>
		<category><![CDATA[enhancing robot learning through explanations]]></category>
		<category><![CDATA[Few-shot learning]]></category>
		<category><![CDATA[human demonstration in robotics]]></category>
		<category><![CDATA[human-robot interaction]]></category>
		<category><![CDATA[human-robot teaching]]></category>
		<category><![CDATA[impact of showing why versus what]]></category>
		<category><![CDATA[improving robot learning efficiency]]></category>
		<category><![CDATA[language conditioning]]></category>
		<category><![CDATA[policy learning]]></category>
		<category><![CDATA[robot cognition and decision-making]]></category>
		<category><![CDATA[robot learning]]></category>
		<category><![CDATA[robot manipulation]]></category>
		<category><![CDATA[robot skill acquisition]]></category>
		<category><![CDATA[state representation]]></category>
		<category><![CDATA[task learning with human guidance]]></category>
		<category><![CDATA[transformer architecture]]></category>
		<category><![CDATA[user study]]></category>
		<category><![CDATA[visual imitation learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201760</guid>

					<description><![CDATA[Researchers have developed an imitation learning method called CIVIL that lets human teachers mark task-relevant objects and explain their actions, enabling robots to learn faster and generalize better than conventional demonstration-only approaches.]]></description>
										<content:encoded><![CDATA[<p>Teaching a robot a new task has always been an exercise in showing rather than explaining. A human guides a robot arm through the motions—picking up a cup, moving it to the coffee machine—and the machine records every joint angle and camera frame, then attempts to reproduce the behavior. But a fundamental gap has long lurked inside this process: the robot sees what the human does, yet never learns why the human chose those actions. Now, a team of researchers at Virginia Tech, Cornell University, and California State University, Northridge has introduced a new approach that closes this gap, and their results suggest that a small change in how humans demonstrate tasks can dramatically improve how robots learn.</p>
<p>The problem the team set out to solve is known in machine learning as causal confusion. When a robot watches a human make coffee, its camera captures far more than the cup and the coffee maker. It also sees bowls, appliances, shadows, and clutter on the counter. If, during training, the cup happens to always sit next to a bowl, the robot may wrongly conclude that the bowl matters—that reaching somewhere near the bowl is the actual goal. The learned policy may work flawlessly in the training environment, but the moment the bowl is removed or moved, the robot fails. The researchers demonstrated this failure mathematically as well as experimentally, showing that when inputs contain correlated but irrelevant features, there is no way for a robot learning purely from demonstrations to disentangle the true cause of the human&#8217;s actions from spurious coincidences.</p>
<p>Their paper, published in the journal Autonomous Robots, also establishes why learning from raw visual data is inherently expensive. Using a linear regression analysis, the authors prove that the amount of demonstration data needed to learn a policy grows exponentially with the dimensionality of the observations. Camera images are extremely high-dimensional, packed with millions of pixels, most of which have nothing to do with the task. Compressing those images into a small set of task-relevant features—say, the position and orientation of a cup—slashes the data requirement. But the catch is that the robot has no way of knowing, on its own, which features are the right ones to keep. Many different feature sets can explain the training data equally well while diverging wildly from the human&#8217;s actual reasoning, and only one of them will generalize beyond it.</p>
<p>The team&#8217;s answer is to change the teaching paradigm rather than the robot. Instead of expecting learners to infer causality from actions alone, they let human teachers communicate the reasoning behind their demonstrations directly. Their algorithm, called CIVIL for Causal and Intuitive Visual Imitation Learning, relies on two simple channels of communication that humans already use naturally: physical markers and spoken language. Before demonstrating a task, the teacher attaches small, lightweight ArUco markers—printed patterns detectable by the robot&#8217;s camera—to the objects that matter. While demonstrating, the teacher narrates what they are focusing on, saying things like &#8220;pick up the cup&#8221; or &#8220;look at the light on the coffee machine.&#8221; The robot records these cues alongside the usual stream of images, states, and actions.</p>
<p>Under the hood, CIVIL converts this augmented data into a feature representation that mirrors human reasoning. The marker poses become explicit features: the robot trains a network to encode exactly the marked positions, using an information-theoretic loss that ensures the features contain all the marker information and nothing more. The spoken instructions do their work through a language-conditioned video segmentation model, which draws bounding boxes around the objects the human mentions. Every pixel outside those boxes is masked to zero, stripping away the clutter that causes causal confusion. The robot then learns a policy—built around a transformer architecture that processes sequences of robot states and visual features—that maps this purified representation to the demonstrated actions. A second training phase distills what the robot learned into a causal network that can extract the same features from raw, unmasked images, so that once training is complete, the robot needs no markers, no language, and no external vision models at test time.</p>
<p>The team validated the approach in simulation using the CALVIN benchmark, a 3D environment with a Franka Emika Panda arm and a tabletop of blocks, drawers, sliding doors, and lights. Across three tasks—picking up a block, choosing between a drawer and a sliding door based on the state of a light bulb, and stacking blocks according to that light—the robots trained with CIVIL consistently outperformed a battery of state-of-the-art baselines, including standard behavior cloning, self-supervised feature learning, object-centric methods, and approaches built on pre-trained vision-language models. The advantage was starkest in out-of-distribution tests. When trained with 120 demonstrations, CIVIL picked up a block from the center of the table—a position never seen during training—in nearly every attempt, while the baselines succeeded less than 20 percent of the time, having latched onto misleading correlations with nearby objects.</p>
<p>Real-world experiments on a physical Franka arm echoed the simulation results. The robot performed four kitchen-table tasks, including stirring or scooping the contents of a pan, pressing a red button among a cluster of colorful cups, picking up a cup from a cluttered table, and pulling a bowl to the center of the table. In each case, the training data contained deliberate spurious correlations—a yellow cup always behind the button, a bowl always in front of the cup—that vanished at test time. CIVIL-trained robots navigated these traps successfully, achieving significantly higher success rates than object-oriented and language-conditioned baselines, especially on unseen object configurations. Notably, CIVIL required object segmentation only during offline training, avoiding the online detection failures that plagued competing methods when objects were gripped or partially occluded.</p>
<p>Perhaps the most striking findings came from a user study with ten participants, who trained the robot to pick up a cup and place it under a coffee machine. The researchers imposed a fixed five-minute teaching budget and compared CIVIL against behavior cloning. Even though attaching markers and narrating instructions consumed time—users provided about nine demonstrations with CIVIL versus eleven without—the robots trained with the enriched data far outperformed those trained on action demonstrations alone, succeeding more than 77 percent of the time versus roughly 40 percent for the baseline. Participants rated the process as intuitive and seamless, and the biggest gains appeared in the most delicate moments of the task: picking up and releasing the cup without knocking it over. The expressiveness of language also seemed to buffer against imperfect human motions, since the robot could rely on stated intent even when demonstrations were sloppy.</p>
<p>The authors also stress-tested their method against imperfect teaching. When users forgot to mark an object or placed markers on irrelevant items, performance dipped slightly but still beat the baseline. When language was vague—simply &#8220;pick up the cup&#8221; in a scene with several cups—the segmentation model sometimes masked the wrong objects, and out-of-distribution performance fell sharply. The researchers frame this as an extreme edge case and point to continuing advances in open-vocabulary segmentation as a path forward. An additional appendix evaluation against a language-conditioned pretraining approach showed CIVIL winning by more than 11 percent overall on new stacking and pouring tasks, while running faster at inference time on the same GPU.</p>
<p>The broader implication is a shift in how the field thinks about teaching machines. Rather than demanding ever more data and ever larger pre-trained models so robots can guess their way to human intent, CIVIL argues that a modest amount of structured human guidance during training—one-time marker placement and a few spoken words—buys enormous gains in learning efficiency and robustness. The robot ends up learning both what to do and why to do it, and because the guidance is only needed at training time, the deployed system behaves like any autonomous policy. The team acknowledges limitations, including reliance on humans correctly identifying all relevant objects and the current restriction to single tasks, and suggests future work on interactive reminders for teachers and scene-graph priors for multi-task settings. But the core message is likely to resonate well beyond this study: when it comes to teaching robots, a little explanation goes a very long way.</p>
<p><strong>Subject of Research:</strong> Causal and intuitive visual imitation learning for robots taught by human demonstrations with markers and language</p>
<p><strong>Article Title:</strong> Civil: causal and intuitive visual imitation learning</p>
<p><strong>Article References:</strong> Dai, Y., Ramirez Sanchez, R., Jeronimus, R., Sagheb, S., Nunez, C. M., Nemlekar, H., &amp; Losey, D. P. (2026). Civil: causal and intuitive visual imitation learning. <em>Autonomous Robots, 50</em>(4), Article 41. <a href="https://doi.org/10.1007/s10514-026-10266-3" rel="noopener noreferrer">https://doi.org/10.1007/s10514-026-10266-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10514-026-10266-3" rel="noopener noreferrer">10.1007/s10514-026-10266-3</a></p>
<p><strong>Keywords:</strong> visual imitation learning, causal confusion, robot manipulation, human-robot interaction, state representation, few-shot learning, language conditioning, policy learning, transformer architecture, autonomous robots, CALVIN benchmark, user study</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201760</post-id>	</item>
	</channel>
</rss>
