Drones have become one of the most disruptive technologies of the decade, buzzing over cities, airports, battlefields, and stadiums. But the same small, agile machines that deliver packages and capture cinematic footage can also smuggle contraband, spy on restricted sites, or collide with manned aircraft. Detecting them reliably has become an urgent problem for public safety agencies, air traffic authorities, and militaries around the world. Now, a team of researchers in China has unveiled a new artificial intelligence system that appears to solve one of the hardest versions of this problem: spotting a tiny drone from the air, against a background full of visual clutter.
The system, called FUS-DETR, was developed by Siyuan Duan, Geng Zhang, Xin Li, Shang Wang, Zhimin Mao, Bingliang Hu, and Xuebin Liu at the Xi’an Institute of Optics and Precision Mechanics of the Chinese Academy of Sciences, together with the University of Chinese Academy of Sciences. Writing in the journal Applied Intelligence, the team reports that their transformer-based detector achieves state-of-the-art results on two demanding air-to-air drone detection benchmarks, reaching 96.6 percent mean average precision at a 50 percent overlap threshold on the Det-Fly dataset and 98.1 percent on the UAV-Eagle dataset.
Those numbers matter because the air-to-air detection scenario is uniquely brutal for computer vision. When a camera drone is used to hunt another drone, the target often occupies only a handful of pixels in the frame. At long range, a quadcopter can shrink to a few dozen pixels across, losing nearly all of the distinctive shape cues that ground-based detectors rely on. Worse, the hunting drone’s own camera is moving, so the background — buildings, trees, roads, clouds, birds — streams past and creates a constantly shifting field of clutter that can masquerade as a target or swallow a real one entirely.
Traditional approaches to this problem have leaned heavily on convolutional neural networks, particularly the YOLO family of real-time detectors, which have dominated practical object detection for years. But convolutional networks process images through local receptive fields, which limits their ability to reason about the global context of a scene. A transformer-based detector, by contrast, uses self-attention mechanisms that let the model relate any part of the image to any other part, making it easier to distinguish a small intruder from its surroundings based on scene-wide patterns rather than purely local appearance.
FUS-DETR builds on the DETR, or Detection Transformer, architecture — specifically the real-time RT-DETR lineage that has recently begun to challenge YOLO models on speed and accuracy. The researchers introduced three key innovations. The first is a Feature-Focused Diffusion Pyramid Network, or FFDP, which tackles a chronic weakness of feature pyramid networks: as information flows between different scales of the image representation, fine details tend to be lost. The FFDP combines feature focus modules with a feature diffusion mechanism that enriches multi-scale feature maps with contextual information, deliberately spreading information across scales so that the faint signature of a small drone is not diluted or discarded along the way.
The second innovation is a Multilevel Attention Fusion Module, or MAF. Detection networks typically blend features from several layers, each capturing different levels of detail — shallow layers hold fine-grained edges and textures, while deeper layers encode broader semantic context. The MAF uses a hierarchical attention mechanism to adaptively fuse local and global features, learning on the fly which combination best captures the fine details of a small target while suppressing the background noise that would otherwise trigger false detections. In effect, the module lets the network decide, for every part of every image, how much to trust close-up texture versus wide-angle context.
The third component consists of three independent information fusion blocks arranged in concatenation and parallel structures. These blocks aggregate multi-scale information efficiently, giving the detection head a richer, more complete representation of each candidate target. Together, the three modules form a pipeline that is explicitly engineered around the two defining difficulties of drone detection from the air: extreme target smallness and heavy background clutter.
The experimental results are notable not just for their accuracy but for the rigor behind them. The team benchmarked FUS-DETR against a broad field of competitors, including classic detectors such as Faster R-CNN, SSD, and RetinaNet, the full YOLO series from YOLOv3 through YOLOv10, and the latest transformer-based detectors including RT-DETRv2 and DEIM. In an unusual display of transparency, the researchers disclosed in supplementary materials that they had initially reproduced some baseline results with an insufficient 50-epoch training schedule; they retrained RT-DETRv2 and DEIM under the full 100-epoch protocol and updated the comparison tables accordingly. All baselines were run under unified, fully documented training settings with the same 640 by 640 input resolution, ensuring that the comparison reflects the architectures rather than differences in training recipes.
The researchers also published detailed module-level complexity analyses, breaking down the parameter counts and computational costs of the FocusBlock and MAF components, including the multi-kernel depth-wise convolutions with kernel sizes of 5, 7, 9, and 11 that drive the feature diffusion process. This kind of accounting matters for a practical reason the authors highlight directly: achieving an optimal balance between top-tier detection performance and a model size suitable for aerial deployment remains a critical challenge. A detector that runs only on a data-center GPU is useless for a drone that must spot intruders in real time using onboard compute. The team’s attention to lightweight, re-parameterizable convolutional refinement operators suggests the design was shaped by deployment constraints from the start.
The datasets underpinning the results are publicly available — Det-Fly on GitHub and UAV-Eagle in an open repository — which means other groups can verify the numbers and build on them. The code is available from the corresponding author upon request. The work was supported by the Open Research Fund of the Shaanxi Key Laboratory of Optical Remote Sensing and Intelligent Information Processing, and the authors declare no competing financial interests.
If the reported performance holds up in independent testing, the implications extend well beyond counter-drone applications. The core techniques — context-preserving feature diffusion across pyramid scales, adaptive local-global attention fusion, and efficient multi-scale aggregation — address the general problem of small object detection in cluttered scenes, which also plagues medical imaging, satellite reconnaissance, wildlife monitoring, and autonomous driving. As drones multiply in the airspace, the ability of one flying machine to reliably see another may become as fundamental to aviation safety as radar was to the last century. FUS-DETR offers a glimpse of what that machine vision might look like: fast, compact, and sharp-eyed enough to find a speck of a drone in a sky full of noise.
Subject of Research: Transformer-based deep learning for detecting small unmanned aerial vehicles in cluttered airborne images
Article Title: FUS-DETR: a robust algorithm for drone detection from a cluttered background in airborne images
Article References: Duan, S., Zhang, G., Li, X., Wang, S., Mao, Z., Hu, B., & Liu, X. (2026). FUS-DETR: a robust algorithm for drone detection from a cluttered background in airborne images. Applied Intelligence, 56(15), Article 424. https://doi.org/10.1007/s10489-026-07448-y
Image Credits: AI Generated
DOI: 10.1007/s10489-026-07448-y
Keywords: UAV detection, FUS-DETR, transformer, object detection, small object detection, air-to-air detection, computer vision, feature pyramid network, attention mechanism, deep learning, drone safety, Applied Intelligence
Cite Scienmag News
Blake Davidson. (October 5, 2026). New AI Model Spots Tiny Drones Hiding in Cluttered Skies. Scienmag. https://scienmag.com/new-ai-model-spots-tiny-drones-hiding-in-cluttered-skies/
Blake Davidson. "New AI Model Spots Tiny Drones Hiding in Cluttered Skies." Scienmag, 5 October 2026, https://scienmag.com/new-ai-model-spots-tiny-drones-hiding-in-cluttered-skies/. Accessed 5 October 2026.
Blake Davidson. "New AI Model Spots Tiny Drones Hiding in Cluttered Skies." Scienmag. October 5, 2026. https://scienmag.com/new-ai-model-spots-tiny-drones-hiding-in-cluttered-skies/

