Thursday, September 24, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Earth Science

Lightweight AI Network Counts Crowds in Real Time on Tiny Embedded Chips

September 24, 2026
in Earth Science
Blake Davidson
By Blake Davidson Scienmag Editorial Profile - Data Science
Reading Time: 5 mins read
0
Lightweight AI Network Counts Crowds in Real Time on Tiny Embedded Chips

Lightweight AI Network Counts Crowds in Real Time on Tiny Embedded Chips

Lightweight AI Network Counts Crowds in Real Time on Tiny Embedded Chips

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Crowded train stations, packed stadiums, and bustling city squares generate exactly the kinds of scenes where knowing how many people are present can mean the difference between smooth crowd management and a dangerous bottleneck. Estimating headcounts from images, a task known as crowd counting, has long been a staple of computer vision research, but most state-of-the-art models are far too heavy to run on the low-power hardware that actually watches those scenes. A research team led by Zhiyuan Zhao, Yubin Wen, and Junyu Gao, with collaborators at the Institute of Artificial Intelligence (TeleAI) of China Telecom and Northwestern Polytechnical University, has now unveiled a strikingly compact neural network that pushes real-time crowd counting onto ordinary embedded devices. Published in the open-access journal Vicinagearth, the work reports inference speeds of 381.7 frames per second on an NVIDIA GTX 1080Ti graphics card and 71.9 frames per second on the modest NVIDIA Jetson TX1, all while keeping accuracy competitive with far larger models.

The motivation behind the study is straightforward: intelligence gathered after the fact is of limited use to security operators and urban planners who need answers as events unfold. Earlier crowd counting approaches generally fall into three camps. Detection-based methods scan video frames for individual people, but they falter badly in dense crowds where bodies and heads occlude one another. Regression-based approaches instead learn global image cues such as texture and gradient statistics to estimate numbers, yet they often miss fine local detail that precision counting demands. The dominant modern paradigm, density map estimation, trains a convolutional network to predict a continuous map in which each head contributes a small blob, and summing the map yields the total count. Pioneering models such as the Multi-column Convolutional Neural Network showed the promise of this idea, but their deep, parameter-heavy architectures demand server-class GPUs and large memory footprints that embedded surveillance hardware simply cannot provide.

Some researchers have chased efficiency with deliberately slimmed-down networks. PCC-Net-light reduced parameters for single-image counting, MobileCount introduced an efficient encoder-decoder framework, and structured knowledge transfer distilled large models into smaller ones. Generic lightweight backbones such as SqueezeNet, MobileNet, and ShuffleNet demonstrated that careful architectural choices, including depthwise separable convolutions, group convolutions, and channel shuffling, could slash computational cost without catastrophic accuracy loss. The new work builds on this lineage but targets a stricter bar the authors call super real-time performance. Previous lightweight counters, they note, cut parameters yet still failed to deliver the fastest possible inference, leaving a gap for applications such as intelligent surveillance, public safety management, urban planning, and intelligent transportation where every millisecond counts.

The proposed architecture follows a stem-encoder-decoder blueprint. The stem network performs early down-sampling that compresses spatially redundant pixel data to one quarter of the original resolution, then applies unusually large convolution kernels of sizes 9, 7, and 5. Large kernels enlarge the network’s receptive field, allowing it to capture detailed head features that small kernels miss, a choice the authors validated in ablation experiments. Because the model trains from scratch without pretrained weights, expanding the receptive field at the front end proves especially valuable. The stem also incorporates ShuffleNetV2-style shuffle blocks, which split channels into two branches, process them with convolutions, and re-mix information through concatenation and channel shuffling to keep the representation expressive at minimal cost.

The encoder, where much of the speed gain originates, organizes features into multi-scale branches at one quarter, one eighth, and one sixteenth of the input resolution, with channel counts of 36, 64, and 96 respectively. Each of its two stages stacks components containing two Conditional Channel Weighting blocks and one Multi-branch Local Fusion block. Conditional Channel Weighting, first introduced in Lite-HRNet and adapted here for crowd counting for the first time, replaces ordinary convolutions with element-wise weighting operations governed by cross-resolution and spatial weight functions, adaptively selecting which feature channels matter at each resolution. The Multi-branch Local Fusion block, a new design from this team, merges multi-scale features exclusively through down-sampling and summation, keeping feature scales small during fusion and thereby holding computational consumption down.

To compensate for the inevitable incompleteness of purely local fusion, the decoder borrows Feature Pyramid Networks, a proven mechanism from object detection. The lowest-resolution encoder output is up-sampled, fused with lateral feature maps produced by one-by-one convolutions, and refined with three-by-three convolutions that smooth the aliasing artifacts of up-sampling; this iterative process repeats until a final density map emerges. Two one-by-one convolution layers then regress the combined features down to a single-channel prediction map. The entire decoder adds only about 0.085 megabytes of parameters. Training uses a straightforward mean squared error loss between predicted and ground-truth density maps, with the Adam optimizer, a cosine-annealed learning rate schedule starting at one times ten to the minus four, 300 epochs, and standard augmentation including random cropping and horizontal flipping.

The numbers are remarkable for a model this size: the whole network weighs just 0.15 megabytes and requires roughly 1.32 gigafloating-point operations per image. Across three standard benchmarks, the small-scale ShanghaiTech dataset with its SHHA and SHHB subsets, the diverse and challenging UCF-QNRF collection of 1,535 dense crowd images, and NWPU-Crowd, currently the largest benchmark with 5,109 images and more than 2.1 million annotated heads spanning densities from zero to 20,033 people per image, the network delivers competitive accuracy, including a mean absolute error near 65 on the SHHA test split while running at roughly 380 frames per second. Runtime tests across four hardware platforms, the GTX 1080Ti, RTX 3090, Jetson TX1, and Jetson Xavier, show per-image inference never exceeding 15 milliseconds at a resolution of 576 by 768 pixels, comfortably inside real-time territory even on low-power modules.

The study also confronts the elephant in the modern research room: large language models and vision-language systems. Although models such as BLIP-2 and LLaVA excel at few-shot and zero-shot vision tasks, the authors quantify why they are hopeless fits for embedded counting today. BLIP-2, with 11 billion parameters, and LLaVA, with 7 billion, demand 15 to 22 gigabytes of GPU memory and take 1.0 to 1.3 seconds per 224 by 224 image even on an NVIDIA A100. On a Jetson TX1, these models cannot complete a single forward pass in under 10 to 12 seconds and effectively exhaust available memory, and even a compact MiniGPT-4-tiny variant exceeds 500 milliseconds per frame, far beyond the 33-millisecond budget that 30 frames-per-second operation requires. The team’s own tests found that no evaluated language-model-based approach exceeded 0.1 frames per second on the TX1, versus more than 70 frames per second for their lightweight convolutional network on the same chip.

Ablation studies reinforce each design decision. Large kernels of 9, 7, and 5 outperformed both stacks of small three-by-three kernels and dilated kernels, with the authors speculating that dilated convolutions ignore very small head information in dense scenes and introduce gridding artifacts. Neither Conditional Channel Weighting nor Multi-branch Local Fusion alone matches the accuracy-and-efficiency combination of the pair working together, and the stem, encoder, and decoder each prove indispensable. The team further introduces an Accuracy-Efficiency Score that jointly weighs mean square error, parameter count, and frame rate, and their model tops this metric across benchmarks, trading roughly 2 to 3 mean absolute error points for a five-to-tenfold speedup compared with heavier competitors whose hundreds of millions of parameters confine them below 50 frames per second.

The researchers acknowledge limits and point the way forward. The model occasionally underestimates counts in extremely dense regions where head boundaries overlap severely, and future work may pair the lightweight architecture with density-aware loss functions or replace the Feature Pyramid Network decoder with something even leaner to reach more marginalized devices. For now, the paper demonstrates a crucial principle for practical artificial intelligence: raw accuracy on a leaderboard tells only part of the story, and a deliberately balanced design can put genuinely useful, super-real-time crowd analytics within reach of the inexpensive embedded hardware that actually keeps watch over the world’s crowds.

Subject of Research: A lightweight deep learning architecture for real-time crowd counting on embedded systems

Article Title: Real-time crowd counting for embedded systems with lightweight architecture

Article References: Zhao, Z., Wen, Y., Yang, S., Ning, L., Liu, Y., & Gao, J. (2025). Real-time crowd counting for embedded systems with lightweight architecture. Vicinagearth, 2(1), Article 13. https://doi.org/10.1007/s44336-025-00025-w

Image Credits: AI Generated

DOI: 10.1007/s44336-025-00025-w

Keywords: crowd counting, embedded systems, lightweight neural network, computer vision, density map estimation, real-time inference, NVIDIA Jetson, convolutional neural network, intelligent surveillance, public safety, Feature Pyramid Networks, efficiency

Cite Scienmag News

Blake Davidson. (September 24, 2026). Lightweight AI Network Counts Crowds in Real Time on Tiny Embedded Chips. Scienmag. https://scienmag.com/lightweight-ai-network-counts-crowds-in-real-time-on-tiny-embedded-chips/

Blake Davidson. "Lightweight AI Network Counts Crowds in Real Time on Tiny Embedded Chips." Scienmag, 24 September 2026, https://scienmag.com/lightweight-ai-network-counts-crowds-in-real-time-on-tiny-embedded-chips/. Accessed 24 September 2026.

Blake Davidson. "Lightweight AI Network Counts Crowds in Real Time on Tiny Embedded Chips." Scienmag. September 24, 2026. https://scienmag.com/lightweight-ai-network-counts-crowds-in-real-time-on-tiny-embedded-chips/

Tags: accuracy of small-scale neural networks in crowd countingAI-powered crowd monitoring at train stations and stadiumsapplications of computer vision in urban safetycollaboration between AI research institutionscomputer visionconvolutional neural networkcrowd countingcrowd counting on embedded devicesdensity map estimationdeployment of AI on edge devicesefficiencyembedded systemsFeature Pyramid Networkshardware-efficient AI models for surveillanceinference speed of compact neural networksintelligent surveillancelightweight neural networklightweight neural networks for real-time crowd estimationlow-power AI models for crowd managementNVIDIA Jetsonopen-access AI research publicationspublic safetyreal-time image analysis for crowded scenesreal-time inference
Share26Tweet16
Previous Post

Virtual Emergency Department Keeps Four in Five Patients Away From Hospital EDs, Landmark Data Show

Next Post

China Updates National Playbook for Robot-Assisted Colorectal Cancer Surgery

Related Posts

One Flood, Two Different Groundwater Stories Beneath Iran’s Caspian Plain
Earth Science

One Flood, Two Different Groundwater Stories Beneath Iran’s Caspian Plain

September 24, 2026
Ten Rules Could Unlock the Power of NASA’s Earth Observation Data
Earth Science

Ten Rules Could Unlock the Power of NASA’s Earth Observation Data

September 24, 2026
Fertilizer choice, not dose, drives nitrous oxide losses in drip-irrigated desert wheat
Earth Science

Fertilizer choice, not dose, drives nitrous oxide losses in drip-irrigated desert wheat

September 24, 2026
Underground Life Takes Center Stage as Soil Book Claims Top Ecology Prize
Earth Science

Underground Life Takes Center Stage as Soil Book Claims Top Ecology Prize

September 24, 2026
Heavy Metal Pollution Builds Up in India’s Scenic National Waterway, Study Warns
Earth Science

Heavy Metal Pollution Builds Up in India’s Scenic National Waterway, Study Warns

September 23, 2026
Warming World Severs the Ancient Monsoon Link That Governs Mediterranean Summers
Earth Science

Warming World Severs the Ancient Monsoon Link That Governs Mediterranean Summers

September 23, 2026
Next Post
China Updates National Playbook for Robot-Assisted Colorectal Cancer Surgery

China Updates National Playbook for Robot-Assisted Colorectal Cancer Surgery

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Tiny creases in soft materials act as rewritable gates that steer and sort droplets
  • Crystal-Plastic Hybrids Could Rewrite How Food Moves From Farm to Table
  • China Updates National Playbook for Robot-Assisted Colorectal Cancer Surgery
  • Lightweight AI Network Counts Crowds in Real Time on Tiny Embedded Chips

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading