Thursday, October 8, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Robots Learn to Feel Pressure Through Cameras Alone, Thanks to Multimodal AI

October 8, 2026
in Technology and Engineering
Blake Davidson
By Blake Davidson Scienmag Editorial Profile - Data Science
Reading Time: 5 mins read
0
Robots Learn to Feel Pressure Through Cameras Alone, Thanks to Multimodal AI

Robots Learn to Feel Pressure Through Cameras Alone, Thanks to Multimodal AI

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Robotic hands have long faced a fundamental dilemma: to grip an egg without crushing it or a wrench without dropping it, a robot needs to know exactly how hard it is pressing, yet the sensors that provide that answer are often bulky, expensive, and fragile. Now, a team of researchers has unveiled a new artificial intelligence framework that lets robotic grippers estimate the pressure they exert on objects using nothing more than cameras and a remarkably rich fusion of visual, geometric, and even linguistic information. The work, published in the journal Results in Engineering, demonstrates that when a neural network is allowed to look at a scene through several complementary lenses at once, it can outperform the best purely visual methods at reading the invisible language of force.

The research, led by Yawen Liu and Wei Tang with colleagues Zhen Zhang and Xinrong Chen, addresses a persistent bottleneck in robotic manipulation. Conventional approaches to force sensing rely on physical instruments such as force and torque sensors or tactile arrays mounted directly on the gripper. Rigid-body models calculate grasping force through explicit mechanical equations, while soft-body models attempt to capture the nonlinear deformations of compliant structures. Although advances in materials science, including conductive hydrogels that endow soft robots with human-like sensory capabilities, have pushed the field forward, physical sensors bring inherent drawbacks: they add system complexity and cost, and their measurement range is limited, making it difficult to adapt to the diverse and dynamic demands of different grasping tasks.

Vision-based pressure estimation has emerged as an elegant alternative. By pointing cameras at the interaction between a gripper and an object, computer vision techniques can infer the magnitude and distribution of applied forces without any physical contact with the measured surface. This eliminates the need for structural modifications to the gripper and offers flexible, scalable perception. Yet most existing methods, including the current state-of-the-art system known as HiVPE, rely on single RGB images. That restriction means the network sees only surface appearance, missing explicit three-dimensional geometric deformation and high-level semantic context, both of which become crucial when dealing with severe occlusions or the complex hyperelastic behavior of soft materials.

The new framework breaks free of that single-modality constraint by weaving together four distinct streams of information. The first is the ordinary RGB image capturing the interaction between the gripper and a high-resolution pressure sensing array. The second is a binary segmentation mask that precisely delineates the gripper’s contact regions, generated by a pre-trained deep segmentation model that isolates the hand from the background. The third is a depth map produced from the RGB image by a monocular depth estimation network, encoding the geometric and spatial relationships of the contact. The fourth is perhaps the most surprising: a short textual prompt, either the word Straight or Roll, describing the gripper’s action, which is converted into a semantic embedding by the CLIP text encoder, the same contrastive language-image model that underpins many modern multimodal AI systems.

Architecturally, the system proceeds in three stages: feature extraction, feature fusion, and decoding. During extraction, a dual-branch hierarchical encoding scheme employs two parallel SE-ResNeXt50 networks. The first processes the raw RGB image to produce five feature maps at progressively abstract scales, spanning downsampling factors from 2 to 32 and channel dimensions from 64 to 2048, capturing everything from low-level textures to high-level semantics. The second encoder jointly processes the concatenated segmentation and depth inputs, yielding a parallel set of features rich in structural and geometric cues. In this way, one branch learns what the scene looks like while the other learns how it is shaped in space.

The fusion stage is where the framework’s distinctive machinery comes into play. Multi-head self-attention, with four attention heads of 64 dimensions each operating on 256-dimensional features, first refines the highest-resolution feature maps within each modality, strengthening intra-modal contextual dependencies through residual connections and feed-forward networks. A bidirectional cross-attention operator then lets the two streams talk to each other: the appearance features absorb structural context from the segmentation and depth branch, while the structural features incorporate appearance cues from the RGB branch. Finally, the 512-dimensional textual embedding is projected through a multilayer perceptron into a 2880-dimensional feature space, reshaped to align spatially with the visual features and concatenated to form a unified multimodal representation. A Feature Pyramid Network decoder then performs top-down aggregation, fusing high-level semantic cues with low-level spatial details before a segmentation head produces the final dense pressure prediction.

Training and evaluation were carried out on a publicly available benchmark containing roughly 650,000 synchronized RGB-pressure frames from two very different soft grippers, one tendon-actuated and one pneumatic. Ground-truth pressure annotations came from a high-resolution Sensel Morph pressure sensing array, spatially aligned with the camera images through calibration and homography transformation. The dataset covers four interaction types: contact, slide, close, and no-contact, the last serving as adversarial samples. The authors discretized continuous pressure values into nine logarithmically spaced categories, from zero background to a top bin of 64 kilopascals and above, and optimized the network with a pixel-wise weighted cross-entropy loss that penalizes underrepresented contact classes to counteract the imbalance between vast non-contact regions and sparse pressure-bearing areas. All results were averaged over six independent runs to verify statistical stability.

The results are striking for the tendon-actuated gripper, where the multimodal method surpassed the previous state of the art on every metric. It achieved a temporal accuracy of 96.57 percent, a contact area intersection-over-union of 77.20 percent, a volumetric IoU of 62.87 percent, and a mean absolute error of just 4.33 pascals, compared with 5.0 pascals for HiVPE and 5.3 for the earlier VPEC-Net. For the pneumatic gripper, the method attained the highest temporal accuracy and volumetric IoU, though it fell marginally short of HiVPE on contact IoU and mean absolute error, a shortfall the authors attribute to multimodal fusion occasionally introducing noise when soft surfaces deform so severely that depth and segmentation maps become blurred. Ablation experiments confirmed that each component earns its place: segmentation masks alone improved performance, adding attention strengthened contact-region predictions, depth maps reduced error further, and textual prompts delivered the final gains in spatial-pressure consistency.

The implications extend well beyond a leaderboard. Because the approach requires no embedded sensors, it could make force-aware grasping accessible to cheap, commodity robotic hardware, from laboratory manipulators to warehouse pickers handling fragile goods. The demonstration that a single descriptive word about the gripper’s action can measurably sharpen pressure predictions hints at a future where robots are instructed in natural language and perceive the physical consequences of their movements through vision alone. The authors are candid about the current limitations: the framework has so far been validated on static, offline benchmark data rather than deployed on physical robots in dynamic manipulation, where real-time sensor noise, unpredictable physical interactions, and inference latency constraints all come into play. They also note that soft grippers, with their complex nonlinear hyperelastic deformation, remain the hardest case, and they point toward more advanced depth estimation, adaptive fusion architectures that dynamically re-weight unreliable modalities, and explicit deformation modeling as the next frontiers. If those challenges are met, the line between seeing and touching may grow ever thinner, bringing robots one step closer to hands that truly understand what they hold.

Subject of Research: Vision-based multimodal pressure estimation for robotic gripper manipulators

Article Title: Vision-Based multimodal pressure estimation for robotic gripper manipulators

Article References: Liu, Y., Tang, W., Zhang, Z., & Chen, X. (2026). Vision-Based multimodal pressure estimation for robotic gripper manipulators. Results in Engineering, 32, Article 113295. https://doi.org/10.1016/j.rineng.2026.113295

Image Credits: AI Generated

DOI: 10.1016/j.rineng.2026.113295

Keywords: robotics, pressure estimation, multimodal AI, soft grippers, computer vision, deep learning, attention mechanism, CLIP, tactile sensing, depth estimation, feature pyramid network, robotic manipulation

Cite Scienmag News

Blake Davidson. (October 8, 2026). Robots Learn to Feel Pressure Through Cameras Alone, Thanks to Multimodal AI. Scienmag. https://scienmag.com/robots-learn-to-feel-pressure-through-cameras-alone-thanks-to-multimodal-ai/

Blake Davidson. "Robots Learn to Feel Pressure Through Cameras Alone, Thanks to Multimodal AI." Scienmag, 8 October 2026, https://scienmag.com/robots-learn-to-feel-pressure-through-cameras-alone-thanks-to-multimodal-ai/. Accessed 8 October 2026.

Blake Davidson. "Robots Learn to Feel Pressure Through Cameras Alone, Thanks to Multimodal AI." Scienmag. October 8, 2026. https://scienmag.com/robots-learn-to-feel-pressure-through-cameras-alone-thanks-to-multimodal-ai/

Tags: AI-driven robotic sensingartificial intelligence for tactile sensingattention mechanismCLIPcomputer visioncomputer vision in roboticsdeep learningDepth estimationfeature pyramid networkmultimodal AImultimodal AI in roboticsmultimodal perception in robotsmultimodal sensor fusionneural networks for robotic manipulationpressure estimationrobotic force estimationrobotic grasping force estimationrobotic manipulationrobotic manipulation without tactile sensorsroboticssensorless force detectionsoft gripperstactile sensingvisual-based pressure sensing
Share26Tweet16
Previous Post

Activated Protein C Keeps Blood Stem Cells Resting and Boosts Transplant Success

Next Post

AI Learns to Read X-rays Like a Radiologist to Spot Bone Tumors With Few False Alarms

Related Posts

When Batteries Change Jobs: Why Predictive Maintenance Must Learn to Move With Them
Technology and Engineering

When Batteries Change Jobs: Why Predictive Maintenance Must Learn to Move With Them

October 8, 2026
Shadows Betray Fakes: Wedge-Based Analysis Exposes Doctored Images
Technology and Engineering

Shadows Betray Fakes: Wedge-Based Analysis Exposes Doctored Images

October 8, 2026
New climate-economy model puts finance, innovation and debt at the heart of global warming projections
Earth Science

New climate-economy model puts finance, innovation and debt at the heart of global warming projections

October 8, 2026
New book cuts through AI hype by listening to the researchers behind the technology
Technology and Engineering

New book cuts through AI hype by listening to the researchers behind the technology

October 8, 2026
Physics Meets AI: New Machine Learning Framework Predicts How Long Electric Vehicle Drive Systems Will Last
Technology and Engineering

Physics Meets AI: New Machine Learning Framework Predicts How Long Electric Vehicle Drive Systems Will Last

October 8, 2026
Deep Beneath Tibet’s Nam Co, a 510-Meter Core Captures a Million Years of Climate History
Earth Science

Deep Beneath Tibet’s Nam Co, a 510-Meter Core Captures a Million Years of Climate History

October 8, 2026
Next Post
AI Learns to Read X-rays Like a Radiologist to Spot Bone Tumors With Few False Alarms

AI Learns to Read X-rays Like a Radiologist to Spot Bone Tumors With Few False Alarms

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • A Quiet Tide, a Turbulent Mouth: How a Brazilian Lagoon Matches the Mixing Power of the World’s Fiercest Estuaries
  • Smart Pebbles and Seismometers Reveal How Flash Floods Rattle Desert Riverbeds
  • Cancer’s Hidden Proteome: Thousands of Dark Proteins Emerge as Drug Targets
  • When Batteries Change Jobs: Why Predictive Maintenance Must Learn to Move With Them

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading