Robotic hands have long faced a fundamental dilemma: to grip an egg without crushing it or a wrench without dropping it, a robot needs to know exactly how hard it is pressing, yet the sensors that provide that answer are often bulky, expensive, and fragile. Now, a team of researchers has unveiled a new artificial intelligence framework that lets robotic grippers estimate the pressure they exert on objects using nothing more than cameras and a remarkably rich fusion of visual, geometric, and even linguistic information. The work, published in the journal Results in Engineering, demonstrates that when a neural network is allowed to look at a scene through several complementary lenses at once, it can outperform the best purely visual methods at reading the invisible language of force.
The research, led by Yawen Liu and Wei Tang with colleagues Zhen Zhang and Xinrong Chen, addresses a persistent bottleneck in robotic manipulation. Conventional approaches to force sensing rely on physical instruments such as force and torque sensors or tactile arrays mounted directly on the gripper. Rigid-body models calculate grasping force through explicit mechanical equations, while soft-body models attempt to capture the nonlinear deformations of compliant structures. Although advances in materials science, including conductive hydrogels that endow soft robots with human-like sensory capabilities, have pushed the field forward, physical sensors bring inherent drawbacks: they add system complexity and cost, and their measurement range is limited, making it difficult to adapt to the diverse and dynamic demands of different grasping tasks.
Vision-based pressure estimation has emerged as an elegant alternative. By pointing cameras at the interaction between a gripper and an object, computer vision techniques can infer the magnitude and distribution of applied forces without any physical contact with the measured surface. This eliminates the need for structural modifications to the gripper and offers flexible, scalable perception. Yet most existing methods, including the current state-of-the-art system known as HiVPE, rely on single RGB images. That restriction means the network sees only surface appearance, missing explicit three-dimensional geometric deformation and high-level semantic context, both of which become crucial when dealing with severe occlusions or the complex hyperelastic behavior of soft materials.
The new framework breaks free of that single-modality constraint by weaving together four distinct streams of information. The first is the ordinary RGB image capturing the interaction between the gripper and a high-resolution pressure sensing array. The second is a binary segmentation mask that precisely delineates the gripper’s contact regions, generated by a pre-trained deep segmentation model that isolates the hand from the background. The third is a depth map produced from the RGB image by a monocular depth estimation network, encoding the geometric and spatial relationships of the contact. The fourth is perhaps the most surprising: a short textual prompt, either the word Straight or Roll, describing the gripper’s action, which is converted into a semantic embedding by the CLIP text encoder, the same contrastive language-image model that underpins many modern multimodal AI systems.
Architecturally, the system proceeds in three stages: feature extraction, feature fusion, and decoding. During extraction, a dual-branch hierarchical encoding scheme employs two parallel SE-ResNeXt50 networks. The first processes the raw RGB image to produce five feature maps at progressively abstract scales, spanning downsampling factors from 2 to 32 and channel dimensions from 64 to 2048, capturing everything from low-level textures to high-level semantics. The second encoder jointly processes the concatenated segmentation and depth inputs, yielding a parallel set of features rich in structural and geometric cues. In this way, one branch learns what the scene looks like while the other learns how it is shaped in space.
The fusion stage is where the framework’s distinctive machinery comes into play. Multi-head self-attention, with four attention heads of 64 dimensions each operating on 256-dimensional features, first refines the highest-resolution feature maps within each modality, strengthening intra-modal contextual dependencies through residual connections and feed-forward networks. A bidirectional cross-attention operator then lets the two streams talk to each other: the appearance features absorb structural context from the segmentation and depth branch, while the structural features incorporate appearance cues from the RGB branch. Finally, the 512-dimensional textual embedding is projected through a multilayer perceptron into a 2880-dimensional feature space, reshaped to align spatially with the visual features and concatenated to form a unified multimodal representation. A Feature Pyramid Network decoder then performs top-down aggregation, fusing high-level semantic cues with low-level spatial details before a segmentation head produces the final dense pressure prediction.
Training and evaluation were carried out on a publicly available benchmark containing roughly 650,000 synchronized RGB-pressure frames from two very different soft grippers, one tendon-actuated and one pneumatic. Ground-truth pressure annotations came from a high-resolution Sensel Morph pressure sensing array, spatially aligned with the camera images through calibration and homography transformation. The dataset covers four interaction types: contact, slide, close, and no-contact, the last serving as adversarial samples. The authors discretized continuous pressure values into nine logarithmically spaced categories, from zero background to a top bin of 64 kilopascals and above, and optimized the network with a pixel-wise weighted cross-entropy loss that penalizes underrepresented contact classes to counteract the imbalance between vast non-contact regions and sparse pressure-bearing areas. All results were averaged over six independent runs to verify statistical stability.
The results are striking for the tendon-actuated gripper, where the multimodal method surpassed the previous state of the art on every metric. It achieved a temporal accuracy of 96.57 percent, a contact area intersection-over-union of 77.20 percent, a volumetric IoU of 62.87 percent, and a mean absolute error of just 4.33 pascals, compared with 5.0 pascals for HiVPE and 5.3 for the earlier VPEC-Net. For the pneumatic gripper, the method attained the highest temporal accuracy and volumetric IoU, though it fell marginally short of HiVPE on contact IoU and mean absolute error, a shortfall the authors attribute to multimodal fusion occasionally introducing noise when soft surfaces deform so severely that depth and segmentation maps become blurred. Ablation experiments confirmed that each component earns its place: segmentation masks alone improved performance, adding attention strengthened contact-region predictions, depth maps reduced error further, and textual prompts delivered the final gains in spatial-pressure consistency.
The implications extend well beyond a leaderboard. Because the approach requires no embedded sensors, it could make force-aware grasping accessible to cheap, commodity robotic hardware, from laboratory manipulators to warehouse pickers handling fragile goods. The demonstration that a single descriptive word about the gripper’s action can measurably sharpen pressure predictions hints at a future where robots are instructed in natural language and perceive the physical consequences of their movements through vision alone. The authors are candid about the current limitations: the framework has so far been validated on static, offline benchmark data rather than deployed on physical robots in dynamic manipulation, where real-time sensor noise, unpredictable physical interactions, and inference latency constraints all come into play. They also note that soft grippers, with their complex nonlinear hyperelastic deformation, remain the hardest case, and they point toward more advanced depth estimation, adaptive fusion architectures that dynamically re-weight unreliable modalities, and explicit deformation modeling as the next frontiers. If those challenges are met, the line between seeing and touching may grow ever thinner, bringing robots one step closer to hands that truly understand what they hold.
Subject of Research: Vision-based multimodal pressure estimation for robotic gripper manipulators
Article Title: Vision-Based multimodal pressure estimation for robotic gripper manipulators
Article References: Liu, Y., Tang, W., Zhang, Z., & Chen, X. (2026). Vision-Based multimodal pressure estimation for robotic gripper manipulators. Results in Engineering, 32, Article 113295. https://doi.org/10.1016/j.rineng.2026.113295
Image Credits: AI Generated
DOI: 10.1016/j.rineng.2026.113295
Keywords: robotics, pressure estimation, multimodal AI, soft grippers, computer vision, deep learning, attention mechanism, CLIP, tactile sensing, depth estimation, feature pyramid network, robotic manipulation
Cite Scienmag News
Blake Davidson. (October 8, 2026). Robots Learn to Feel Pressure Through Cameras Alone, Thanks to Multimodal AI. Scienmag. https://scienmag.com/robots-learn-to-feel-pressure-through-cameras-alone-thanks-to-multimodal-ai/
Blake Davidson. "Robots Learn to Feel Pressure Through Cameras Alone, Thanks to Multimodal AI." Scienmag, 8 October 2026, https://scienmag.com/robots-learn-to-feel-pressure-through-cameras-alone-thanks-to-multimodal-ai/. Accessed 8 October 2026.
Blake Davidson. "Robots Learn to Feel Pressure Through Cameras Alone, Thanks to Multimodal AI." Scienmag. October 8, 2026. https://scienmag.com/robots-learn-to-feel-pressure-through-cameras-alone-thanks-to-multimodal-ai/

