<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>selective attention &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/selective-attention/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 20 Sep 2026 20:18:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>selective attention &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns Where to Trust Depth for Sharper Chili Pepper Segmentation</title>
		<link>https://scienmag.com/ai-learns-where-to-trust-depth-for-sharper-chili-pepper-segmentation/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 20:18:11 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[AI-based plant organ segmentation]]></category>
		<category><![CDATA[chili pepper]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[convolutional and transformer segmentation models]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[Depth Anything V2]]></category>
		<category><![CDATA[depth routing]]></category>
		<category><![CDATA[depth-guided selective attention in agriculture]]></category>
		<category><![CDATA[depth-routed segmentation accuracy]]></category>
		<category><![CDATA[handling overlapping plant organs in computer vision]]></category>
		<category><![CDATA[improving crop treatment precision]]></category>
		<category><![CDATA[machine learning for agriculture]]></category>
		<category><![CDATA[monocular depth]]></category>
		<category><![CDATA[multi-sensor depth and RGB integration]]></category>
		<category><![CDATA[open-access plant imaging research]]></category>
		<category><![CDATA[organ segmentation]]></category>
		<category><![CDATA[overcoming occlusion in field images]]></category>
		<category><![CDATA[plant methods]]></category>
		<category><![CDATA[plant organ recognition in messy field conditions]]></category>
		<category><![CDATA[precision agriculture]]></category>
		<category><![CDATA[precision spraying for chili peppers]]></category>
		<category><![CDATA[selective attention]]></category>
		<category><![CDATA[site-specific spraying]]></category>
		<category><![CDATA[spray-aware perception]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202188</guid>

					<description><![CDATA[A new depth-routing AI framework sharply improves organ-level segmentation of chili pepper plants for precision spraying.]]></description>
										<content:encoded><![CDATA[<p>Researchers in Xinjiang, China, have unveiled a new artificial intelligence framework that teaches a segmentation network exactly when and where to trust estimated depth information, dramatically improving how computers distinguish leaves, peppers, and flowers in messy field photographs. The method, called Depth-Routed Selective Attention, or DRSA, is described in an open-access study published in the journal Plant Methods, and it could become a key perception module for precision spraying systems that aim to hit only the plant organs that need treatment.</p>
<p>The problem the team set out to solve is deceptively simple to state but notoriously difficult in practice. Site-specific spraying in chili pepper production requires a machine to separate individual organs—leaves, fruits, and flowers—from handheld images captured in open fields. Under ideal studio lighting, modern convolutional and transformer-based segmentation models handle such tasks well. But real pepper canopies are unforgiving: organs overlap and occlude one another, dust coats leaf surfaces, and the waxy, glossy skin of chili fruits produces specular highlights that scramble the color and texture cues on which RGB-only models depend. When appearance fails, the network&#8217;s predictions smear across organ boundaries, and any downstream spraying decision inherits that error.</p>
<p>Depth information offers an obvious escape route. A second camera or a laser scanner can supply geometric structure that survives bad lighting, but RGB-D hardware adds cost, calibration burden, and fragility for handheld field use. The researchers instead turned to monocular depth estimation, using the publicly available pretrained Depth Anything V2 model to infer a depth map from each ordinary phone photograph. This estimated depth acts as an accessible structural prior—no special sensors required. Yet the team recognized a subtlety that most depth-fusion approaches ignore: the reliability of estimated monocular depth is not uniform across an image. It tends to be trustworthy in some regions, particularly near strong geometric boundaries, and questionable elsewhere. Fusing depth indiscriminately can therefore inject noise precisely where the network can least afford it.</p>
<p>DRSA&#8217;s central innovation is a single per-pixel routing field that jointly governs where two depth-derived mechanisms contribute. The first mechanism is depth-boundary cross-attention, which lets the network consult geometric cues near organ contours, where they matter most for separating touching leaves and fruits. The second is residual depth fusion, which blends depth features into the RGB representation in the regions the routing field selects. Through one shared decision, DRSA ensures that geometric cues act near organ boundaries while RGB remains the default carrier of information everywhere else. In other words, the network does not have to choose globally between trusting color or trusting depth; it makes that choice locally, pixel by pixel, for every image it sees.</p>
<p>Crucially, the routing field is calibrated online from the network&#8217;s own depth-on and depth-suppressed predictions, without requiring any manually annotated trust maps. This design sidesteps what would otherwise be a laborious labeling burden: nobody has to sit down and mark which parts of each depth estimate are reliable. Instead, the model compares its own behavior with and without depth, learns where depth helps, and routes accordingly. The approach reflects a broader principle gaining traction in agricultural AI—estimated cues from foundation models are useful, but only if the system knows their limits and applies them selectively.</p>
<p>To train and evaluate the framework, the team built PepperField-EstDepth, a self-constructed dataset of 3,940 handheld field images of chili pepper canopies, each paired with estimated monocular depth. The images were collected with commodity phone cameras in open field plots in southern Xinjiang, with a field-acquisition team assisting with collection and annotation. On this benchmark, DRSA achieved a mean intersection over union of 90.20 percent and a boundary mIoU of 84.48 percent, outperforming both RGB-only baselines and attention-based RGB-D fusion baselines. Relative to the RGB segmentation reference, the gains amounted to 1.98 and 2.67 percentage points respectively—modest-sounding margins that translate into substantially cleaner organ boundaries in exactly the ambiguous, occluded regions where spraying errors originate.</p>
<p>The authors also stress-tested generalization using group cross-validation, a protocol that holds out entire groups of images to simulate deployment on unseen field conditions. Under this stricter regime, DRSA reached an mIoU of 0.8919 plus or minus 0.0031 and a boundary mIoU of 0.8294 plus or minus 0.0046, indicating that the performance is stable rather than an artifact of particular images. Because the study used only handheld phone photographs and a publicly available pretrained depth checkpoint, with no novel physical materials produced, the pipeline is deliberately reproducible by other laboratories working on similar crops.</p>
<p>For the intended spraying application, the numbers matter most at the organ level. DRSA attained a target recall of 0.9814 and a target precision of 0.9756, meaning that nearly all organs requiring spray are detected and very few non-target organs are wrongly activated. The organ-level off-target activation rate was just 2.44 percent—a figure that speaks directly to reducing chemical waste and collateral deposition on flowers or leaves that should remain untreated. Timing measurements show a segmentation-only latency of 43.0 milliseconds when depth is pre-generated, rising to 219.6 milliseconds for the full RGB-to-mask visual pipeline when online Depth Anything V2-L depth generation is included. Those latencies position DRSA as a pre-spray perception module rather than a real-time closed-loop controller, a distinction the authors make explicitly.</p>
<p>The work was supported by the Joint Foundation of Tarim University and Nanjing Agricultural University, the Bingtuan Science and Technology Program, the Tianshan Talents Cultivation Program of Xinjiang Uygur Autonomous Region, and the Presidential Foundation of Tarim University. The research team, based at Tarim University&#8217;s College of Information Engineering and the Key Laboratory of Tarim Oasis Agriculture under the Ministry of Education, with corresponding author Tiecheng Bai, sees DRSA as part of a larger shift toward spray-aware perception in precision agriculture. As foundation models for depth, segmentation, and language continue to mature, the selective-use philosophy embodied in DRSA—borrow a powerful prior, but route it only where it pays—offers a template that could extend well beyond chili peppers to other row crops, orchard systems, and any vision task where sensor estimates are helpful but imperfect.</p>
<p><strong>Subject of Research:</strong> Depth-guided selective attention for chili pepper organ segmentation in precision agriculture</p>
<p><strong>Article Title:</strong> DRSA: Depth-Routed Selective Attention for chili pepper organ segmentation with selective use of estimated monocular depth</p>
<p><strong>Article References:</strong> Zhou, W., Wang, Z., Chi, J., Chen, H., Yan, P., &amp; Bai, T. (2026). DRSA: Depth-Routed Selective Attention for chili pepper organ segmentation with selective use of estimated monocular depth. <em>Plant Methods</em>. <a href="https://doi.org/10.1186/s13007-026-01581-y" rel="noopener noreferrer">https://doi.org/10.1186/s13007-026-01581-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s13007-026-01581-y" rel="noopener noreferrer">10.1186/s13007-026-01581-y</a></p>
<p><strong>Keywords:</strong> precision agriculture, chili pepper, organ segmentation, monocular depth, Depth Anything V2, selective attention, depth routing, site-specific spraying, computer vision, deep learning, Plant Methods, spray-aware perception</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202188</post-id>	</item>
		<item>
		<title>Where Sound Meets Sight: Spatial Coincidence Decides When Attention Fails</title>
		<link>https://scienmag.com/where-sound-meets-sight-spatial-coincidence-decides-when-attention-fails/</link>
		
		<dc:creator><![CDATA[Glenn Wilkins]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 20:15:02 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[attention and sensory processing]]></category>
		<category><![CDATA[attentional load]]></category>
		<category><![CDATA[audiovisual integration]]></category>
		<category><![CDATA[auditory stimuli]]></category>
		<category><![CDATA[auditory-visual stimulus fusion]]></category>
		<category><![CDATA[cognitive psychology]]></category>
		<category><![CDATA[cross-modal attention and perception]]></category>
		<category><![CDATA[cross-modal interaction]]></category>
		<category><![CDATA[effects of attentional load on perception]]></category>
		<category><![CDATA[influence of spatial alignment on sensory detection]]></category>
		<category><![CDATA[limits of multisensory attention]]></category>
		<category><![CDATA[multisensory integration]]></category>
		<category><![CDATA[multisensory perception]]></category>
		<category><![CDATA[multisensory perception under cognitive load]]></category>
		<category><![CDATA[neural mechanisms of multisensory binding]]></category>
		<category><![CDATA[perceptual load]]></category>
		<category><![CDATA[role of superior colliculus in sensory integration]]></category>
		<category><![CDATA[RSVP]]></category>
		<category><![CDATA[selective attention]]></category>
		<category><![CDATA[spatial coincidence]]></category>
		<category><![CDATA[spatial coincidence in perception]]></category>
		<category><![CDATA[spatial localization of sounds and sights]]></category>
		<category><![CDATA[spatial rule]]></category>
		<category><![CDATA[visual target detection]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=202100</guid>

					<description><![CDATA[New research shows that sounds boost visual detection under heavy attentional load only when they share the same spatial location as the visual target.]]></description>
										<content:encoded><![CDATA[<p>Everyday perception feels effortless, yet beneath the surface the brain is constantly deciding which fragments of sight and sound belong together. A new study published in Attention, Perception, and Psychophysics by Qingqing Li, Huazhi Li, Hecheng Jiang, Yulong Liu, Mengni Zhou, Jinglong Wu, Jiajia Yang, and Qiong Wu has now mapped, with unusual precision, the conditions under which the brain can still fuse what it hears with what it sees when attention is stretched to its limits. The central finding is striking: when a sound arrives at exactly the same place as a visual target, it can boost the detection of that target even when the observer&#8217;s attention is almost completely consumed by another demanding task. When the sound comes from somewhere else, that boost evaporates under high load, as if the brain simply no longer has the resources to bind signals scattered across space.</p>
<p>The question the researchers tackled is one of the oldest debates in multisensory science. Decades of work, beginning with the classic neurophysiological studies of the superior colliculus by Stein and Meredith, established that neurons in the midbrain respond most powerfully when visual and auditory inputs converge on the same spatial location. This gave rise to the so-called spatial rule of multisensory integration: signals from different senses enhance one another most when they plausibly originate from the same object or event. Yet behavioral studies of simple, meaningless stimuli, such as a flash paired with a brief tone, have sometimes suggested that cross-modal interactions persist even when observers are instructed to ignore one modality entirely. That persistence has been interpreted as evidence that audiovisual integration is automatic, running to completion regardless of attentional control, much like the Stroop effect or preattentive feature binding described in Anne Treisman&#8217;s feature-integration theory.</p>
<p>But automaticity has its skeptics. Work by Nilli Lavie on perceptual load has shown that when the primary task is easy, spare attentional capacity spills over onto irrelevant stimuli, producing what looks like automatic processing. Under high load, that spillover disappears, and distractors are effectively filtered out. Critics such as Tsal and Benoni have argued that many apparent load effects are actually dilution effects, driven by the number of items competing for processing rather than by a genuine exhaustion of perceptual resources. Against this backdrop, the question of whether audiovisual integration truly requires attentional resources, or merely appears to, remained unresolved, particularly for the simple, arbitrary sound-flash pairings that dominate the experimental literature.</p>
<p>There was a second, equally important gap. Previous experiments had rarely asked whether the spatial relationship between the sound and the visual target changes how attentional load affects integration. Most studies either presented stimuli from a single location or did not systematically manipulate spatial coincidence. Yet if the spatial rule holds, then a spatially congruent sound and a spatially incongruent sound might tap into fundamentally different neural mechanisms, one that is robust and resource-independent, the other fragile and dependent on spare capacity. The new study was designed to separate these possibilities cleanly.</p>
<p>To manipulate attentional load, the researchers adopted a rapid serial visual presentation paradigm, one of the most reliable tools in cognitive psychology for controlling how much attention a distractor task consumes. Participants watched a fast-moving stream of characters at fixation while searching for targets within the stream. In the no-load condition, the stream demanded minimal attention; in the low-load condition, it demanded a moderate amount; and in the high-load condition, the task was tuned to consume nearly all available attentional resources. This graded approach allowed the team to trace how integration behaves as resources are progressively drained, rather than simply comparing easy and hard tasks.</p>
<p>On top of this load manipulation, the researchers controlled spatial coincidence. Visual targets and task-irrelevant auditory stimuli were presented either at the same spatial position or at different positions. Participants were instructed to ignore the sounds completely, so any influence of the tones on visual target identification would reflect an involuntary cross-modal interaction. The design therefore crossed three levels of attentional load with two levels of spatial congruency, producing a matrix of conditions in which the contributions of resources and space could be disentangled statistically.</p>
<p>The results were clear and, in places, surprising. Spatially congruent auditory stimuli improved the identification of visual targets across all three load conditions, including the high-load condition in which the RSVP stream was consuming the bulk of participants&#8217; attention. In other words, even when observers were pushed close to their attentional limits, a sound arriving from the same location as the visual target still made that target easier to detect. This resilience suggests that spatially coincident audiovisual integration operates through a mechanism that is largely automatic, one that does not compete meaningfully for the limited resources taxed by the RSVP task. It is consistent with the idea that spatially aligned signals are bound early and efficiently, perhaps at subcortical or early cortical levels where the spatial rule was originally discovered.</p>
<p>The story changed dramatically for spatially incongruent sounds. When the auditory stimulus appeared at a different location from the visual target, it failed to enhance visual identification under high-load conditions. Under no load and low load, some cross-modal influence could still be observed, but as the RSVP task drained resources, the benefit of the mismatched sound vanished. This dissociation is the paper&#8217;s key contribution: it demonstrates that not all audiovisual integration is created equal. Spatially congruent integration survives the harshest attentional conditions, whereas spatially incongruent integration depends on the availability of spare attentional capacity. The findings therefore reconcile two seemingly contradictory literatures, showing that studies reporting automatic integration may have relied on conditions, or on spatial arrangements, in which the congruent, resource-independent mechanism was doing the work.</p>
<p>The theoretical implications reach into several active debates. For proponents of load theory, the results support the view that high load filters out stimuli that lack a privileged link to the attended event, while leaving intact interactions that are structurally embedded in the spatial layout of the scene. For multisensory researchers, the study adds a crucial qualification to the spatial rule: spatial coincidence is not merely a facilitator of integration but a determinant of whether integration can occur without attention. The work also echoes earlier findings by Ho, Santangelo, and Spence on multisensory warning signals, which showed that spatial correspondence matters enormously for the effectiveness of cross-modal alerts, and by McDonald and colleagues, who identified neural substrates of perceptual enhancement by cross-modal spatial attention. The new data extend this line by showing that the spatial rule becomes decisive precisely when attention runs out.</p>
<p>Beyond the laboratory, the findings carry practical weight. Warning signals in aircraft cockpits, operating theaters, and vehicles often pair a sound with a visual indicator, and designers generally assume the pairing will help even when operators are overloaded. This study suggests that assumption holds only when the sound and the visual signal share a location. A warning tone emitted from a speaker far from the relevant display may fail to boost detection in a stressed, overloaded operator, whereas a spatially aligned cue could still cut through. Similarly, the results inform the design of assistive technologies and virtual reality environments, where multisensory cues are increasingly used to guide attention, and they may help explain why multisensory enhancement can break down in conditions of fatigue or divided attention.</p>
<p>The study, conducted with approval from the Academic Committee of the Department of Psychology at Soochow University in line with the Declaration of Helsinki, was supported by the Japan Society for the Promotion of Science, the Pre-approved Project of the Wenzhou Key Research Base for Philosophy and Social Sciences, and the Social Science project of Suzhou University of Science and Technology. Data and code are available from the corresponding authors upon reasonable request. As multisensory research moves toward real-world applications, this work delivers a deceptively simple message with deep consequences: the brain&#8217;s ability to merge sight and sound is not a single switch but a layered system, and only the layer built on spatial coincidence keeps working when everything else is asked to give.</p>
<p><strong>Subject of Research:</strong> How attentional load and spatial coincidence modulate audiovisual integration of simple stimuli</p>
<p><strong>Article Title:</strong> The effect of attentional loads on audiovisual integration: When spatial coincidence matters</p>
<p><strong>Article References:</strong> The effect of attentional loads on audiovisual integration: When spatial coincidence matters. (n.d.). <a href="https://doi.org/10.3758/s13414-026-03257-0" rel="noopener noreferrer">https://doi.org/10.3758/s13414-026-03257-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13414-026-03257-0" rel="noopener noreferrer">10.3758/s13414-026-03257-0</a></p>
<p><strong>Keywords:</strong> attentional load, audiovisual integration, spatial coincidence, multisensory perception, RSVP, cross-modal interaction, selective attention, perceptual load, spatial rule, visual target detection, auditory stimuli, cognitive psychology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">202100</post-id>	</item>
	</channel>
</rss>
