Across the windswept sandlands of northern China, where scattered trees rise from a sea of grass, ecologists have long struggled with a deceptively simple question: how many trees are there, and how big are they? A new study published in Smart Agricultural Technology offers a striking answer. Researchers have developed DeepTree, a lightweight deep-learning framework that can pick out and measure every individual tree crown in drone imagery of temperate savanna woodlands, using a fraction of the computing power that conventional approaches demand. The work promises to transform how scientists monitor some of the world’s most ecologically fragile transition zones between forest and grassland.
Temperate savanna woodlands form a broad ecotone between closed forests and open grasslands, and the ones studied here span two vast sandy regions of Inner Mongolia: the Hunsandak Sandland, covering roughly 52,000 square kilometers, and the Horqin Sandland, extending over about 35,100 square kilometers. These landscapes support exceptionally rich biodiversity and act as reservoirs of genetic resources, yet the woody vegetation that anchors them has received far less attention than the grasses and herbs around it. Because trees and shrubs in these systems control carbon storage, community stability, and the spatial distribution of biodiversity, knowing exactly where each individual tree stands, and how large its crown is, is fundamental to nearly every ecological question that can be asked about the region.
The challenge is that savanna trees are maddeningly heterogeneous. Natural savanna woodlands stretch across large environmental gradients, and the woody vegetation responds with pronounced spatial variation, producing wildly diverse phenotypes at the level of single trees. Classical image-processing techniques such as local maximum filtering, marker-controlled watershed segmentation, template matching, region growing, and edge detection can work under simple conditions, but they falter in natural forests with complex canopy structures. They depend on manual parameter tuning, extract features only crudely, and generalize poorly. Deep convolutional neural networks offered a way forward, but the standard workhorse for instance segmentation, Mask R-CNN, carries a heavy price: its ResNet-50 backbone and Feature Pyramid Network demand substantial computational resources, and the pyramid module can lose information when propagating features to high resolutions, undermining the detection of small or sparsely distributed crowns.
To build DeepTree, the research team first assembled an unusually rich dataset. Between July 20 and 30, 2020, a DJI Phantom 4 RTK quadcopter flew nine sample plots across six banners of Inner Mongolia under clear, cloudless skies, capturing thousands of photographs at roughly 100 meters altitude with 75 percent forward and 70 percent side overlap. The imagery was processed into nine digital orthophotos with pixel resolutions between 2.63 and 3.06 centimeters, then resampled to a uniform 10 centimeters and sliced into 1,000-by-1,000-pixel tiles, each covering one hectare, yielding 663 tiles in total. From eight of the plots, the team manually annotated 8,724 individual tree crowns using Labelme, exporting the labels in COCO format. The ninth plot was deliberately held back as an independent cross-site test, ensuring that no spatially overlapping regions leaked between training and evaluation data.
The architectural surgery at the heart of DeepTree is elegant in its simplicity. The researchers replaced ResNet-50 with MobileNet V3, a backbone that uses just 13 layers and depthwise separable convolutions instead of 50 conventional layers, dramatically cutting computational complexity. They swapped the Feature Pyramid Network for NAS-FPN, a feature-fusion module discovered through neural architecture search, which enhances multi-scale representation and improves detection of small targets while using only 32 channels. The framework then branches into two variants: MNMS R-CNN (S), which pairs this lightweight backbone and neck with a standard region-of-interest head for maximum speed, and MNMS R-CNN (L), which adds an enhanced SCNet-based head with Squeeze-and-Excitation, Global Context, and Feature Relay components to prioritize detection completeness and segmentation accuracy.
The performance gains are remarkable. MNMS R-CNN (S) slashed computational cost from 248.0 GFLOPs to just 5.6 GFLOPs and shrank the parameter count from 42.9 million to 3.8 million, while boosting inference speed from 10.2 to 19.2 frames per second, an 88.2 percent improvement, all on 1.5 gigabytes of GPU memory instead of 6.2. The larger variant, MNMS R-CNN (L), required only 9.0 GFLOPs and 8.1 million parameters yet achieved the best overall accuracy in the study: bounding-box average precision rose from 34.9 to 46.5 percent, a relative improvement of about 33 percent, while recall jumped from 62.4 to 81.2 percent and the F1 score climbed to 81.1 percent. Comparisons against YOLO11n-seg and the high-capacity HTC framework confirmed that neither compact size nor raw capacity alone delivers the right balance; MNMS R-CNN (L) outperformed both on the key detection metrics while remaining far leaner than HTC’s 76.93 million parameters and 401.1 GFLOPs.
Perhaps the most ecologically telling results emerged when the models were tested across different canopy densities. In high-density scenes, MNMS R-CNN (L) lifted recall from Mask R-CNN’s 45.5 percent to 84.9 percent, cutting missed crowns from 72 to just 20 in a representative case. Even in sparse, low-density stands, where isolated small crowns blend into the grassy background, it nearly doubled recall, from 28.2 to 46.5 percent. Stratified analysis revealed a persistent pattern: omission errors concentrated among smaller crowns, particularly in low-density scenes where weak spectral and textural contrast against grass undermined detection. The team also found that performance degraded as imagery was downsampled, with F1 falling from 81.1 percent at 10 centimeters to 74.1 percent at 0.5 meters and 55.4 percent at 1.0 meter, underscoring how much fine spatial detail matters for resolving individual crowns.
Beyond simply finding trees, DeepTree extracts the structural traits that ecologists actually need. For crown width, the coefficient of determination reached 0.98 with a root mean square error of just 0.30 meters; for crown area, R squared was again 0.98 with an error of 2.66 square meters. The model recovered crown-width estimates spanning 1.41 to 16.05 meters and crown areas from 1.33 to 203.71 square meters, closely matching reference ranges that Mask R-CNN systematically truncated. An independent field validation using 166 measured elm trees, matched to drone detections by geographic coordinates, yielded R squared values of 0.82 for crown width and 0.78 for tree height, with the modest gap attributable to the one-year interval between imagery acquisition and field measurement, GPS positioning uncertainty, and differences in how crown width is defined on the ground versus in a segmentation mask.
Applied across all nine plots, the model mapped tree densities ranging from about 30 to 102.7 trees per hectare and canopy areas from roughly 709 to 3,117 square meters per hectare, revealing stark spatial heterogeneity in woody cover. In the held-out ninth plot, never seen during training, the model still achieved 75.8 percent precision and 76.0 percent recall, demonstrating genuine cross-site transferability, the property that separates a local demonstration from a scalable monitoring tool. The authors acknowledge limitations: RGB imagery alone struggles when crowns, shrubs, and grass share similar colors, and future work should fuse in LiDAR and spectral data, expand training libraries to more species and seasons, and test the framework across broader environmental gradients. Even so, DeepTree points toward a future where drone-based biodiversity assessment, carbon accounting, and long-term monitoring of the world’s fragile forest-grassland ecotones can run on modest hardware, one tree at a time.
Subject of Research: Lightweight deep learning instance segmentation for individual tree crown delineation and structural trait extraction from UAV imagery in temperate savanna woodlands
Article Title: DeepTree: Deep learning-based lightweight instance segmentation for individual tree crown delineation and structural trait extraction
Article References: Duan, T., Yang, B., Li, X., Hou, D., Hu, P., Li, X., Cong, W., & Wang, F. (2026). DeepTree: Deep learning-based lightweight instance segmentation for individual tree crown delineation and structural trait extraction. Smart Agricultural Technology, 15, Article 102552. https://doi.org/10.1016/j.atech.2026.102552
Image Credits: AI Generated
DOI: 10.1016/j.atech.2026.102552
Keywords: deep learning, UAV remote sensing, instance segmentation, tree crown delineation, Mask R-CNN, MobileNet V3, NAS-FPN, savanna woodlands, Inner Mongolia, ecological monitoring, phenotype extraction, lightweight models
Cite Scienmag News
Margaret Porter. (September 30, 2026). Lightweight AI Maps Every Tree in China’s Fragile Savanna Woodlands. Scienmag. https://scienmag.com/lightweight-ai-maps-every-tree-in-chinas-fragile-savanna-woodlands/
Margaret Porter. "Lightweight AI Maps Every Tree in China’s Fragile Savanna Woodlands." Scienmag, 30 September 2026, https://scienmag.com/lightweight-ai-maps-every-tree-in-chinas-fragile-savanna-woodlands/. Accessed 1 October 2026.
Margaret Porter. "Lightweight AI Maps Every Tree in China’s Fragile Savanna Woodlands." Scienmag. September 30, 2026. https://scienmag.com/lightweight-ai-maps-every-tree-in-chinas-fragile-savanna-woodlands/

