<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>digital surface model &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/digital-surface-model/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 12:49:06 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>digital surface model &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Learns When to Switch On Global Attention for Sharper Satellite Maps</title>
		<link>https://scienmag.com/ai-learns-when-to-switch-on-global-attention-for-sharper-satellite-maps/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 12:49:06 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[adaptive attention in deep learning models]]></category>
		<category><![CDATA[automatic extraction of roads and buildings from satellite images]]></category>
		<category><![CDATA[axial attention]]></category>
		<category><![CDATA[digital surface model]]></category>
		<category><![CDATA[dynamic computation]]></category>
		<category><![CDATA[efficiency improvements in satellite map generation]]></category>
		<category><![CDATA[global attention mechanisms in remote sensing]]></category>
		<category><![CDATA[high-resolution aerial photograph processing]]></category>
		<category><![CDATA[innovation in remote sensing neural networks]]></category>
		<category><![CDATA[ISPRS Potsdam]]></category>
		<category><![CDATA[ISPRS Vaihingen]]></category>
		<category><![CDATA[multimodal fusion]]></category>
		<category><![CDATA[neural network fusion of optical and elevation data]]></category>
		<category><![CDATA[pixel-perfect land cover classification]]></category>
		<category><![CDATA[policy gradient]]></category>
		<category><![CDATA[reinforcement learning]]></category>
		<category><![CDATA[remote sensing]]></category>
		<category><![CDATA[remote sensing change detection techniques]]></category>
		<category><![CDATA[RL-Axial framework for image segmentation]]></category>
		<category><![CDATA[satellite imagery analysis]]></category>
		<category><![CDATA[semantic segmentation]]></category>
		<category><![CDATA[semantic segmentation in urban mapping]]></category>
		<category><![CDATA[Transformer]]></category>
		<category><![CDATA[urban mapping]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=227803</guid>

					<description><![CDATA[Researchers have developed RL-Axial, a segmentation framework that uses reinforcement learning to decide for each image patch whether to apply row, column, both, or no axial attention when fusing RGB and elevation data for satellite mapping.]]></description>
										<content:encoded><![CDATA[<p>Every aerial photograph tells two stories at once. The first is written in light: rooftops, roads, trees, and shadows rendered in the red, green, and blue bands that cameras capture from hundreds of meters above a city. The second is written in geometry: a digital surface model, or DSM, that records the precise height of every pixel, distinguishing a flat parking lot from a warehouse roof even when both look identical from orbit. For years, researchers have tried to teach neural networks to read both stories simultaneously, fusing optical imagery with elevation data to produce pixel-perfect maps of urban environments. The task, known as high-resolution remote sensing semantic segmentation, underpins land cover mapping, change detection, and the automatic extraction of roads and buildings. Now a team at the Changchun Institute of Optics, Fine Mechanics and Physics of the Chinese Academy of Sciences has introduced a framework that takes an unusual approach to the problem: instead of fixing how the network thinks, it lets the network decide, patch by patch, how much thinking each piece of an image actually deserves.</p>
<p>The framework, called RL-Axial and described in the journal Cognitive Computation, attacks a stubborn inefficiency at the heart of modern segmentation models. Convolutional encoder-decoder architectures such as U-Net excel at capturing local texture through skip connections and multi-scale aggregation, but their limited receptive fields leave them weak at modeling global context. Transformer-style models fix that weakness with self-attention, which lets every pixel consult every other pixel, yet the computational cost of that conversation grows quadratically with image size, a serious problem for the gigapixel tiles typical of aerial surveys. Worse, transformers apply the same attention budget everywhere: easy regions of sky or uniform pavement receive the same expensive global computation as tangled street intersections full of small, ambiguous objects. The result is redundant computation on simple areas and, paradoxically, under-modeling of the hard cases that matter most. RL-Axial&#8217;s central idea is to allocate global modeling capacity on demand rather than imposing a single, fixed computation path on every sample.</p>
<p>The architecture rests on two pillars. The first is a Cross-Modal Gated Fusion head, or CMGF, which handles the delicate business of merging RGB and DSM features. Naive fusion strategies, whether early stacking of co-registered inputs or simple feature concatenation, are fragile in practice: resolution mismatches, imperfect geometric correction, and different noise statistics between sensors introduce subtle misalignments that static fusion can amplify into boundary blurring and semantic drift. CMGF instead constructs an explicit cross-modal state by concatenating the RGB features, the DSM features, and their difference, a term that makes the module directly sensitive to inconsistencies between the two modalities. A small predictor built from a 3&#215;3 convolution, group normalization, ReLU, and a 1&#215;1 convolution then produces fusion logits, and a softmax over the two modalities yields per-pixel weight maps. The fused feature is a weighted sum, so every single pixel can lean toward appearance or geometry as the evidence dictates, and unreliable modality responses are suppressed locally rather than globally.</p>
<p>The second pillar is the element that gives the framework its name: a controllable axial-attention operator with four distinct execution modes. Full two-dimensional self-attention scales as the square of the number of pixels, but axial attention decomposes the problem into row-wise and column-wise one-dimensional attention, reducing complexity to a far more manageable order that suits dense prediction. RL-Axial goes further by making the operator switchable: for any given image patch, attention can be turned off entirely, applied along rows only, applied along columns only, or applied along both axes. The off mode is not a mere convenience; it provides a stable fallback during training and inference, ensuring the network never depends on attention to function. Each axial branch follows a projection-residual-normalization pattern, with its output added back to the input and normalized position-wise, a design that helps avoid the distribution shifts attention branches can introduce.</p>
<p>Deciding which mode to use is framed as a reinforcement learning problem. A lightweight policy network, a two-layer multilayer perceptron, receives a globally average-pooled state vector from each patch&#8217;s fused representation and maps it to a categorical distribution over the four modes. Crucially, states are never averaged across the mini-batch, so different patches in the same batch can select different modes independently. Training the selector uses what the authors call a group-normalized policy gradient, an objective that borrows the relative-reward normalization idea popularized by group-relative methods but deliberately avoids the probability-ratio clipping and reference-policy regularization of PPO and GRPO. For each patch, a reference prediction is generated with attention off, then K candidate actions are sampled and evaluated. Each candidate&#8217;s reward is the difference in soft-IoU between its prediction and the reference, and these within-sample rewards are normalized into advantages. An entropy bonus, linearly annealed over training, encourages broad exploration early and more deterministic choices later.</p>
<p>Training proceeds in two carefully separated stages. In the first, axial attention is disabled and the dual modality encoders, the CMGF module, the decoder, and the segmentation head are trained jointly with standard cross-entropy loss using stochastic gradient descent. In the second stage, the encoders and fusion head are frozen, and the policy, axial operator, decoder, and segmentation head continue training with Adam. Gradient paths are explicitly partitioned: the policy loss, built from detached rewards, updates only the selector parameters, while the segmentation loss updates only the axial operator and decoder, never differentiating through the discrete action choice. This separation keeps the non-differentiable scheduling decision cleanly isolated from supervised learning. At inference time, per-sample decisions do not force batch size one; samples selecting the same mode are grouped, the corresponding axial branch is executed for each group, and the original order is restored before decoding, preserving patch-wise adaptivity within ordinary batched tensor operations.</p>
<p>Evaluated on the ISPRS Vaihingen and Potsdam benchmarks, the standard proving grounds for urban remote sensing segmentation, RL-Axial delivered strong single-run results. On Vaihingen, aerial imagery at 9-centimeter ground sampling distance paired with a DSM, the model reached 84.58 percent mean intersection over union, along with 92.73 percent overall accuracy and 91.50 percent mean F1. On the larger Potsdam dataset, with 6000-by-6000-pixel orthophotos at 5-centimeter resolution, it achieved 86.73 mIoU, 91.73 overall accuracy, and 92.74 mean F1. Those mIoU figures sit numerically 0.35 and 0.53 points above the literature-reported values for FTransUNet, a strong multimodal transformer baseline. The authors are notably careful here: because baseline numbers come from their original publications with source-specific evaluation protocols, the comparison is descriptive rather than a controlled head-to-head test, and no claims of statistical superiority are made. Qualitatively, the predictions show cleaner region layouts, sharper boundaries at class transitions, fewer boundary breaks, and more coherent object shapes, particularly in the dense, complex urban layouts of Potsdam.</p>
<p>The ablation studies reveal exactly where the gains come from, and the honesty of the decomposition is refreshing. An optical-only backbone reaches 82.14 mIoU; naive RGB-DSM fusion adds 1.23 points, confirming the value of elevation even without learned gating. Replacing naive fusion with CMGF contributes a further 0.54 points, the best fixed axial mode adds 0.40, and adaptive scheduling contributes the final 0.27. In a controlled comparison where every fixed attention branch and every alternative selector shared the same backbone, initialization, and training budget, the learned selector beat the best fixed native mode, a fixed guided focal-axial adaptation, random scheduling, a supervised action classifier, and a Gumbel-Softmax selector, while an oracle that exhaustively evaluates all four actions per patch sets an upper bound of 85.07 mIoU, leaving a 0.49-point gap. The reward design matters too: optimizing the change in mIoU directly outperformed rewards based on F1 or overall accuracy, evidence that aligning the learning signal with the target metric pays off. Group size ablations showed performance rising from K equals 2 to a peak at K equals 8, the value adopted.</p>
<p>Perhaps the most revealing result is what the learned policy actually does. Rather than collapsing onto a single favorite mode, it uses all four: row-plus-column attention for 30 percent of test patches, the off path for 29 percent, row-only for 23 percent, and column-only for 18 percent. Profiling each branch on a single RTX 3090 GPU, the authors estimate that this distribution yields roughly 40.125 gigaflops and 28.8 milliseconds per patch, a branch-only reduction of 7.55 percent in FLOPs and 10.51 percent in latency compared with always running full row-plus-column attention, though they stress these figures exclude selector overhead and are not end-to-end speedup claims. The broader significance lies in the paradigm: axial attention itself is not new to remote sensing, and the authors explicitly disclaim that novelty. What is new is treating global context as a schedulable resource, learned through direct feedback from the segmentation metric itself. As earth observation constellations multiply and the appetite for fine-grained urban mapping grows, frameworks that spend computation where it matters, and skip it where it does not, may define the next generation of geospatial AI.</p>
<p><strong>Subject of Research:</strong> Reinforcement-learning-based adaptive axial attention scheduling for multimodal RGB and digital surface model segmentation in high-resolution remote sensing</p>
<p><strong>Article Title:</strong> RL-Axial: Learning When to Compute Global Context for Multimodal Segmentation</p>
<p><strong>Article References:</strong> Xing, D., Yang, H., Zhang, J., Liu, P., &amp; Wang, Y. (2026). RL-Axial: Learning When to Compute Global Context for Multimodal Segmentation. <em>Cognitive Computation, 18</em>(1), Article 113. <a href="https://doi.org/10.1007/s12559-026-10661-z" rel="noopener noreferrer">https://doi.org/10.1007/s12559-026-10661-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s12559-026-10661-z" rel="noopener noreferrer">10.1007/s12559-026-10661-z</a></p>
<p><strong>Keywords:</strong> remote sensing, semantic segmentation, multimodal fusion, reinforcement learning, axial attention, digital surface model, ISPRS Vaihingen, ISPRS Potsdam, policy gradient, dynamic computation, urban mapping, transformer</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">227803</post-id>	</item>
	</channel>
</rss>
