<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Depth Anything V2 &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/depth-anything-v2/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sat, 26 Sep 2026 00:59:19 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Depth Anything V2 &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>FuseDepth Combines Frozen AI Visual Priors to Unlock True Metric Depth From a Single Photo</title>
		<link>https://scienmag.com/fusedepth-combines-frozen-ai-visual-priors-to-unlock-true-metric-depth-from-a-single-photo/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 00:59:19 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-based scene understanding]]></category>
		<category><![CDATA[combining relative and absolute depth models]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[computer vision depth sensing]]></category>
		<category><![CDATA[Depth Anything V2]]></category>
		<category><![CDATA[depth discontinuities]]></category>
		<category><![CDATA[domain generalization in depth estimation]]></category>
		<category><![CDATA[foundation models]]></category>
		<category><![CDATA[fusion of foundation models]]></category>
		<category><![CDATA[image matting]]></category>
		<category><![CDATA[improved depth boundary accuracy]]></category>
		<category><![CDATA[LoRa]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[metric depth]]></category>
		<category><![CDATA[metric depth from single images]]></category>
		<category><![CDATA[monocular depth estimation]]></category>
		<category><![CDATA[off-the-shelf AI models]]></category>
		<category><![CDATA[open-access depth estimation frameworks]]></category>
		<category><![CDATA[prior fusion]]></category>
		<category><![CDATA[scene scale calibration]]></category>
		<category><![CDATA[surface normals]]></category>
		<category><![CDATA[zero-shot depth prediction]]></category>
		<category><![CDATA[zero-shot learning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215787</guid>

					<description><![CDATA[A new framework called FuseDepth fuses frozen detection, matting and surface-normal models with an adapted Depth Anything backbone to achieve sharper, truly metric depth estimation on unseen scenes without any test-time calibration.]]></description>
										<content:encoded><![CDATA[<p>Estimating how far away every pixel in a photograph truly lies, in meters rather than merely in relative order, has long been one of computer vision&#8217;s most stubborn challenges. A new open-access study published in Machine Learning with Applications introduces FuseDepth, a framework that pushes zero-shot monocular depth estimation closer to practical deployment by fusing several frozen, off-the-shelf AI models into a single pipeline. Rather than training yet another giant network, the approach borrows the strengths of existing foundation models and composes them in a structured way, achieving sharper boundaries and more reliable absolute scale on scenes the system has never seen.</p>
<p>The central problem the paper identifies is a coupled failure mode. Relative-depth models such as Depth Anything v2 and Marigold produce visually convincing depth maps that generalize well across domains, but their outputs are not metrically calibrated, so converting them to true distances typically requires per-image scale and shift fitting against ground-truth data, which is unavailable in real deployments. Conversely, metric-aware systems such as ZoeDepth, Metric3Dv2, DepthPro, ZeroDepth and UniDepth anchor absolute scale better, yet can blur object boundaries and produce locally inconsistent geometry when the scene distribution shifts. FuseDepth attacks both weaknesses simultaneously instead of trading one for the other.</p>
<p>The architecture works in two stages. In the first, a frozen Depth Anything v2 backbone is adapted with lightweight low-rank adapters, known as LoRA modules, and paired with a boundary-regularized local metric-scale field. Instead of applying a single global affine calibration, a small convolutional head predicts spatially varying scale and shift values, biased toward low-frequency variation and piecewise smoothness. Crucially, the regularization is relaxed near estimated occlusion boundaries, allowing metric scale to change legitimately where objects end, while penalizing arbitrary pixelwise calibration everywhere else. No per-image or per-dataset scale fitting occurs at inference.</p>
<p>The second stage injects explicit high-fidelity visual priors. A pretrained YOLOv8 detector proposes object regions, and each proposal is refined with BiRefNet, an image-matting model that recovers far sharper alpha boundaries than conventional segmentation masks. The gradients of these alpha mattes are aggregated into a scene-wide contour map, and overlapping or adjacent surface hypotheses are linked into a deterministic surface-interaction graph. From this graph, the method derives a soft jump set marking locations where depth may legitimately discontinuously change and where smoothness assumptions should be suspended.</p>
<p>A complementary geometric prior comes from Omnidata, a frozen monocular surface-normal estimator. Surface normals describe the local orientation of planes and curved surfaces, information that is orthogonal to boundary evidence. FuseDepth enforces a projective compatibility between predicted depth gradients and the normal field, computed under the camera&#8217;s intrinsic matrix, but this constraint is deliberately suppressed wherever the graph-derived jump set indicates an occlusion boundary. The result is a piecewise regularization strategy: smoothness and normal consistency hold within surfaces, while boundary-preserving correction takes over at discontinuities.</p>
<p>A compact residual network of roughly 0.2 million parameters then refines the coarse metric depth. Rather than simply concatenating inputs, the refiner emits three residual proposals per pixel, one each for contour, normal and coarse-depth corrections, together with softmax-normalized arbitration weights that mix them locally. These learned gates allow the model to favor the contour branch near detected boundaries and lean on normals or the coarse depth in smooth or uncertain regions. The trainable components are learned only through the final depth losses, without any explicit reliability labels.</p>
<p>Evaluation follows a deliberately strict protocol. FuseDepth is trained on roughly 50,000 images from NYUv2, ScanNet and KITTI, and tested zero-shot on three unseen benchmarks spanning panoramic HDR scenes with LiDAR depth, long-range driving footage and diverse indoor and outdoor RGB-D imagery. Relative-depth baselines receive only one affine mapping fitted once on the training mixture and then frozen. Across all targets, FuseDepth posts the lowest errors, reaching an absolute relative error of 0.105 on SYNS, 0.171 on DDAD and 0.182 and 0.276 on the indoor and outdoor DIODE splits, while also recording the lowest per-image log-scale bias, a direct measure of absolute-scale stability.</p>
<p>Mechanism-isolation experiments support the design choices. Removing the contour prior, the normal prior, or the graph structure each degrades accuracy, and replacing the learned gate with plain concatenation or fixed equal weighting is consistently worse. Deliberately corrupted external priors, such as alpha masks taken from the wrong image or normals from a different scene, cause only graceful degradation because the gate shifts its mass away from the corrupted branch. Statistical analysis shows the improvements over the strongest baseline, UniDepth, are significant at the 95 percent level on SYNS and DIODE, while the DDAD advantage is small but reproduced across three independent training seeds and all geographic subsets.</p>
<p>The authors are candid about limitations. The pipeline depends on the quality and cost of its external priors, runs at 15.4 frames per second with 7.4 gigabytes of peak GPU memory, and cannot fully certify that the frozen upstream models never saw the evaluation datasets during their own pretraining. Low-quality contour priors naturally occur on roughly 6 to 12 percent of target images. Even so, the study demonstrates that thoughtfully composing existing foundation models can beat monolithic systems on strict metric transfer, suggesting a future where visual AI advances by orchestrating specialized experts rather than simply scaling up single networks.</p>
<p><strong>Subject of Research:</strong> Zero-shot metric monocular depth estimation via fusion of frozen foundation visual priors</p>
<p><strong>Article Title:</strong> FuseDepth: Zero-shot metric depth with semantic–geometric fusion of foundation visual priors</p>
<p><strong>Article References:</strong> Javidnia, H. (2026). FuseDepth: Zero-shot metric depth with semantic–geometric fusion of foundation visual priors. <em>Machine Learning with Applications, 26</em>, Article 101010. <a href="https://doi.org/10.1016/j.mlwa.2026.101010" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.101010</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.101010" rel="noopener noreferrer">10.1016/j.mlwa.2026.101010</a></p>
<p><strong>Keywords:</strong> monocular depth estimation, zero-shot learning, metric depth, computer vision, foundation models, Depth Anything v2, image matting, surface normals, LoRA, machine learning, depth discontinuities, prior fusion</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215787</post-id>	</item>
		<item>
		<title>AI Learns Where to Trust Depth for Sharper Chili Pepper Segmentation</title>
		<link>https://scienmag.com/ai-learns-where-to-trust-depth-for-sharper-chili-pepper-segmentation/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 20:18:11 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[AI-based plant organ segmentation]]></category>
		<category><![CDATA[chili pepper]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[convolutional and transformer segmentation models]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[Depth Anything V2]]></category>
		<category><![CDATA[depth routing]]></category>
		<category><![CDATA[depth-guided selective attention in agriculture]]></category>
		<category><![CDATA[depth-routed segmentation accuracy]]></category>
		<category><![CDATA[handling overlapping plant organs in computer vision]]></category>
		<category><![CDATA[improving crop treatment precision]]></category>
		<category><![CDATA[machine learning for agriculture]]></category>
		<category><![CDATA[monocular depth]]></category>
		<category><![CDATA[multi-sensor depth and RGB integration]]></category>
		<category><![CDATA[open-access plant imaging research]]></category>
		<category><![CDATA[organ segmentation]]></category>
		<category><![CDATA[overcoming occlusion in field images]]></category>
		<category><![CDATA[plant methods]]></category>
		<category><![CDATA[plant organ recognition in messy field conditions]]></category>
		<category><![CDATA[precision agriculture]]></category>
		<category><![CDATA[precision spraying for chili peppers]]></category>
		<category><![CDATA[selective attention]]></category>
		<category><![CDATA[site-specific spraying]]></category>
		<category><![CDATA[spray-aware perception]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202188</guid>

					<description><![CDATA[A new depth-routing AI framework sharply improves organ-level segmentation of chili pepper plants for precision spraying.]]></description>
										<content:encoded><![CDATA[<p>Researchers in Xinjiang, China, have unveiled a new artificial intelligence framework that teaches a segmentation network exactly when and where to trust estimated depth information, dramatically improving how computers distinguish leaves, peppers, and flowers in messy field photographs. The method, called Depth-Routed Selective Attention, or DRSA, is described in an open-access study published in the journal Plant Methods, and it could become a key perception module for precision spraying systems that aim to hit only the plant organs that need treatment.</p>
<p>The problem the team set out to solve is deceptively simple to state but notoriously difficult in practice. Site-specific spraying in chili pepper production requires a machine to separate individual organs—leaves, fruits, and flowers—from handheld images captured in open fields. Under ideal studio lighting, modern convolutional and transformer-based segmentation models handle such tasks well. But real pepper canopies are unforgiving: organs overlap and occlude one another, dust coats leaf surfaces, and the waxy, glossy skin of chili fruits produces specular highlights that scramble the color and texture cues on which RGB-only models depend. When appearance fails, the network&#8217;s predictions smear across organ boundaries, and any downstream spraying decision inherits that error.</p>
<p>Depth information offers an obvious escape route. A second camera or a laser scanner can supply geometric structure that survives bad lighting, but RGB-D hardware adds cost, calibration burden, and fragility for handheld field use. The researchers instead turned to monocular depth estimation, using the publicly available pretrained Depth Anything V2 model to infer a depth map from each ordinary phone photograph. This estimated depth acts as an accessible structural prior—no special sensors required. Yet the team recognized a subtlety that most depth-fusion approaches ignore: the reliability of estimated monocular depth is not uniform across an image. It tends to be trustworthy in some regions, particularly near strong geometric boundaries, and questionable elsewhere. Fusing depth indiscriminately can therefore inject noise precisely where the network can least afford it.</p>
<p>DRSA&#8217;s central innovation is a single per-pixel routing field that jointly governs where two depth-derived mechanisms contribute. The first mechanism is depth-boundary cross-attention, which lets the network consult geometric cues near organ contours, where they matter most for separating touching leaves and fruits. The second is residual depth fusion, which blends depth features into the RGB representation in the regions the routing field selects. Through one shared decision, DRSA ensures that geometric cues act near organ boundaries while RGB remains the default carrier of information everywhere else. In other words, the network does not have to choose globally between trusting color or trusting depth; it makes that choice locally, pixel by pixel, for every image it sees.</p>
<p>Crucially, the routing field is calibrated online from the network&#8217;s own depth-on and depth-suppressed predictions, without requiring any manually annotated trust maps. This design sidesteps what would otherwise be a laborious labeling burden: nobody has to sit down and mark which parts of each depth estimate are reliable. Instead, the model compares its own behavior with and without depth, learns where depth helps, and routes accordingly. The approach reflects a broader principle gaining traction in agricultural AI—estimated cues from foundation models are useful, but only if the system knows their limits and applies them selectively.</p>
<p>To train and evaluate the framework, the team built PepperField-EstDepth, a self-constructed dataset of 3,940 handheld field images of chili pepper canopies, each paired with estimated monocular depth. The images were collected with commodity phone cameras in open field plots in southern Xinjiang, with a field-acquisition team assisting with collection and annotation. On this benchmark, DRSA achieved a mean intersection over union of 90.20 percent and a boundary mIoU of 84.48 percent, outperforming both RGB-only baselines and attention-based RGB-D fusion baselines. Relative to the RGB segmentation reference, the gains amounted to 1.98 and 2.67 percentage points respectively—modest-sounding margins that translate into substantially cleaner organ boundaries in exactly the ambiguous, occluded regions where spraying errors originate.</p>
<p>The authors also stress-tested generalization using group cross-validation, a protocol that holds out entire groups of images to simulate deployment on unseen field conditions. Under this stricter regime, DRSA reached an mIoU of 0.8919 plus or minus 0.0031 and a boundary mIoU of 0.8294 plus or minus 0.0046, indicating that the performance is stable rather than an artifact of particular images. Because the study used only handheld phone photographs and a publicly available pretrained depth checkpoint, with no novel physical materials produced, the pipeline is deliberately reproducible by other laboratories working on similar crops.</p>
<p>For the intended spraying application, the numbers matter most at the organ level. DRSA attained a target recall of 0.9814 and a target precision of 0.9756, meaning that nearly all organs requiring spray are detected and very few non-target organs are wrongly activated. The organ-level off-target activation rate was just 2.44 percent—a figure that speaks directly to reducing chemical waste and collateral deposition on flowers or leaves that should remain untreated. Timing measurements show a segmentation-only latency of 43.0 milliseconds when depth is pre-generated, rising to 219.6 milliseconds for the full RGB-to-mask visual pipeline when online Depth Anything V2-L depth generation is included. Those latencies position DRSA as a pre-spray perception module rather than a real-time closed-loop controller, a distinction the authors make explicitly.</p>
<p>The work was supported by the Joint Foundation of Tarim University and Nanjing Agricultural University, the Bingtuan Science and Technology Program, the Tianshan Talents Cultivation Program of Xinjiang Uygur Autonomous Region, and the Presidential Foundation of Tarim University. The research team, based at Tarim University&#8217;s College of Information Engineering and the Key Laboratory of Tarim Oasis Agriculture under the Ministry of Education, with corresponding author Tiecheng Bai, sees DRSA as part of a larger shift toward spray-aware perception in precision agriculture. As foundation models for depth, segmentation, and language continue to mature, the selective-use philosophy embodied in DRSA—borrow a powerful prior, but route it only where it pays—offers a template that could extend well beyond chili peppers to other row crops, orchard systems, and any vision task where sensor estimates are helpful but imperfect.</p>
<p><strong>Subject of Research:</strong> Depth-guided selective attention for chili pepper organ segmentation in precision agriculture</p>
<p><strong>Article Title:</strong> DRSA: Depth-Routed Selective Attention for chili pepper organ segmentation with selective use of estimated monocular depth</p>
<p><strong>Article References:</strong> Zhou, W., Wang, Z., Chi, J., Chen, H., Yan, P., &amp; Bai, T. (2026). DRSA: Depth-Routed Selective Attention for chili pepper organ segmentation with selective use of estimated monocular depth. <em>Plant Methods</em>. <a href="https://doi.org/10.1186/s13007-026-01581-y" rel="noopener noreferrer">https://doi.org/10.1186/s13007-026-01581-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s13007-026-01581-y" rel="noopener noreferrer">10.1186/s13007-026-01581-y</a></p>
<p><strong>Keywords:</strong> precision agriculture, chili pepper, organ segmentation, monocular depth, Depth Anything V2, selective attention, depth routing, site-specific spraying, computer vision, deep learning, Plant Methods, spray-aware perception</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202188</post-id>	</item>
	</channel>
</rss>
