Wednesday, September 23, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Earth Science

One AI Model, Any Device: Flexibly Slicable Network Cleans Up Speech from Earbuds to the Cloud

September 23, 2026
in Earth Science
Violet Maxwell
By Violet Maxwell Scienmag Editorial Profile - Natural Hazards
Reading Time: 4 mins read
0
One AI Model, Any Device: Flexibly Slicable Network Cleans Up Speech from Earbuds to the Cloud

One AI Model, Any Device: Flexibly Slicable Network Cleans Up Speech from Earbuds to the Cloud

One AI Model, Any Device: Flexibly Slicable Network Cleans Up Speech from Earbuds to the Cloud

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

A single neural network that can shrink itself to one percent of its full size and still rescue speech buried in noise, reverberation, clipping, and lost data packets has been unveiled by researchers at Northwestern Polytechnical University and the Institute of Artificial Intelligence (TeleAI) at China Telecom. The system, called SEFlow, published in the open-access journal Vicinagearth, is designed for a future the authors describe as AI Flow: a vision of distributed intelligence in which devices, edge servers, and cloud computers collaborate seamlessly, each drawing on the same underlying model at whatever scale its hardware can afford.

The problem the team set out to solve is a familiar one for anyone deploying artificial intelligence in the real world. Modern deep networks achieve remarkable quality by growing ever larger, but a model that runs comfortably on a cloud server is hopeless on a pair of wireless earbuds. Traditionally, engineers have trained separate models for each computational tier, or resorted to pruning and knowledge distillation, both of which require additional fine-tuning after the fact. SEFlow takes a different route: a single network is trained once in such a way that it can be dynamically sliced into subnetworks of wildly different sizes and deployed directly, with no retraining, onto anything from a smartphone to a data center.

The technical heart of the approach is a set of flexible modules the authors call FlexAttention, FlexLinear, and FlexRMSNorm. In a conventional transformer, the width of each layer is fixed: a linear layer has a set number of inputs and outputs, and multi-head attention has a fixed number of heads. FlexLinear instead slices its weight matrix to any chosen output size and, crucially, activates only a chosen subset of its input neurons, following the same logic as early-exit methods that activate only early layers. FlexAttention adjusts the number of active attention heads, and FlexRMSNorm adapts the size of its normalization parameters. Because each output is a weighted sum of many neurons, slicing works mathematically without changing what the layer computes—only how much of it runs.

Depth is handled by early exit. The network is built from a stack of residual blocks whose outputs all share the same shape, so a decoder can read the result after any number of blocks. Because of how residual connections accumulate, decoding after block number B-bar is equivalent to decoding a fusion of the features produced by the first B-bar blocks; early layers capture the main content of the signal while later layers refine details, so cutting depth sacrifices fine polish but preserves the essentials. Adjusting depth changes computation roughly linearly, while adjusting width changes it quadratically, giving the deployer a two-dimensional dial for matching any computational budget.

Training such a shape-shifting network requires care. At every training step, the team duplicates each batch: one copy trains the full network, the other trains a randomly sampled subnetwork with a random depth and width, ensuring every neuron is updated regularly. To keep multi-GPU training efficient—since GPUs assigned smaller subnetworks would otherwise finish early and sit idle—a synchronized pseudo-random generator gives every GPU the same subnetwork index each step. According to the paper, this synchronization alone cut average training computation by about 38 percent and training time by about 35 percent.

SEFlow is not merely flexible; it is also unified. Instead of training separate models for denoising, dereverberation, declipping, and packet loss concealment, the team applied dynamic data augmentation: clean speech was corrupted with random combinations of noise at signal-to-noise ratios between minus 5 and 20 decibets, simulated room reverberation, waveform clipping, and Markov-chain-simulated packet loss. One model learned to handle all of these degradations, alone or together. A lightweight auxiliary voice activity detection decoder, and a loss combining complex-spectrogram and magnitude reconstruction with the VAD objective, further boosted quality across metrics such as PESQ, STOI, and downstream speaker and speech recognition scores.

The backbone builds on BS-RoFormer, the winning system of the NeurIPS 2024 speech enhancement challenge, splitting the frequency axis into 41 sub-bands distributed approximately uniformly on the Mel scale. A two-stage band-splitting scheme makes the model sampling-rate agnostic: the same architecture handles audio at 8, 16, 22.05, 24, 32, 44.1, and 48 kilohertz simply by using fewer sub-bands at lower rates. A Distributed Grouped Sampler keeps each training batch at a single sampling rate, avoiding wasteful upsampling and GPU idle time, reducing computation by a further 17 percent in their experiments.

The results are striking in their scalability. The full network has about 27 million parameters and costs roughly 24.7 GMACs per second on 16-kilohertz audio; the smallest subnetwork, with a single residual block and a single attention head, has 1.69 million parameters and costs about 0.19 GMACs per second—two orders of magnitude less—yet still measurably improves speech quality over the unprocessed input. At full size, SEFlow performs comparably to state-of-the-art task-specific models on denoising benchmarks from the INTERSPEECH 2020 DNS Challenge and on packet loss concealment benchmarks from the 2022 PLC Challenge, with demo-quality declipping shown on the project homepage. Interesting quirks emerged: a six-block, one-head configuration beat a one-block, four-head early-exit variant across all metrics despite using only 37 percent of its compute, suggesting width can matter more than depth in some regimes.

The authors are candid about trade-offs. Flexibly trained models slightly underperform fixed full-scale networks at the same nominal size, packet loss concealment demands at least two blocks for usable performance, and automatic speech recognition accuracy downstream remains limited, hinting that the backbone or loss may need task-specific tuning. There are also overheads in training multiple subnetworks simultaneously and in deciding which subnetwork to invoke at inference. Still, the researchers argue the approach carries real environmental promise: fixed models burn peak energy even on easy inputs, whereas adaptive slicing reduces multiply-accumulate operations and memory accesses—the dominant energy costs on edge chips—potentially trimming the carbon footprint of always-on speech processing.

What makes SEFlow resonate beyond acoustics is its implication for how AI might be delivered everywhere at once. Rather than a patchwork of bespoke models scattered across earbuds, phones, cars, and servers, a single family-model system could flow intelligence across the device-edge-cloud continuum, expanding and contracting to fit each platform. If the same recipe—flexible width, early exit, unified multi-task training—transfers to language, vision, and audio generation models, the paper’s vision of ubiquitous, resource-aware intelligence moves a step closer to reality. Demonstrations of SEFlow’s outputs, from heavily noisy cafe chatter to clipped and packet-mangled calls, are publicly available, and the code is available on request, inviting the community to stress-test this elegantly elastic architecture.

Subject of Research: A flexibly scalable, unified neural architecture for multi-task speech enhancement

Article Title: Towards a flexible and unified architecture for speech enhancement

Article References: Feng, L., Zhang, C., & Zhang, X.-L. (2025). Towards a flexible and unified architecture for speech enhancement. Vicinagearth, 2(1), Article 14. https://doi.org/10.1007/s44336-025-00022-z

Image Credits: AI Generated

DOI: 10.1007/s44336-025-00022-z

Keywords: speech enhancement, SEFlow, flexible neural networks, FlexAttention, early exit, AI Flow, edge computing, slimmable networks, packet loss concealment, denoising, BS-RoFormer, resource-constrained inference

Cite Scienmag News

Violet Maxwell. (September 23, 2026). One AI Model, Any Device: Flexibly Slicable Network Cleans Up Speech from Earbuds to the Cloud. Scienmag. https://scienmag.com/one-ai-model-any-device-flexibly-slicable-network-cleans-up-speech-from-earbuds-to-the-cloud/

Violet Maxwell. "One AI Model, Any Device: Flexibly Slicable Network Cleans Up Speech from Earbuds to the Cloud." Scienmag, 23 September 2026, https://scienmag.com/one-ai-model-any-device-flexibly-slicable-network-cleans-up-speech-from-earbuds-to-the-cloud/. Accessed 23 September 2026.

Violet Maxwell. "One AI Model, Any Device: Flexibly Slicable Network Cleans Up Speech from Earbuds to the Cloud." Scienmag. September 23, 2026. https://scienmag.com/one-ai-model-any-device-flexibly-slicable-network-cleans-up-speech-from-earbuds-to-the-cloud/

Tags: adaptive speech enhancement technologyAI FlowAI model deployment across multiple devicesBS-RoFormerdenoisingdistributed AI for edge and cloud devicesdynamic network slicingearly exitedge computingend-to-end speech restoration solutionsFlexAttentionflexible neural networksflexible neural networks for wireless earbudsneural network model size reductionnoise and reverberation suppression in audioopen-access AI research on speech enhancementpacket loss concealmentreal-world speech signal processingresource-constrained inferencescalable neural network architectureSEFlowsingle training AI models for diverse hardwareslimmable networksspeech enhancement
Share26Tweet16
Previous Post

New Knee Surgery Scorecard Captures What Matters to Chinese Patients

Next Post

Welded Conductive Skin Turns Stretchy Rubber Into Ultra-Sensitive Wearable Sensor

Related Posts

Hybrid FRP and Stainless Steel Rebars Give Concrete Joints Corrosion Resistance and Seismic Ductility
Earth Science

Hybrid FRP and Stainless Steel Rebars Give Concrete Joints Corrosion Resistance and Seismic Ductility

September 23, 2026
Seaweed Farming Could Reshape Vietnam’s Mekong Delta, but the Evidence Is Thinner Than the Hype
Earth Science

Seaweed Farming Could Reshape Vietnam’s Mekong Delta, but the Evidence Is Thinner Than the Hype

September 23, 2026
Brazil’s Atlantic Forest Hides a Carbon Surprise Beneath Its Grasslands
Earth Science

Brazil’s Atlantic Forest Hides a Carbon Surprise Beneath Its Grasslands

September 23, 2026
Why Agile Small Firms Still Fail to Go Green: New Evidence From Indonesia
Earth Science

Why Agile Small Firms Still Fail to Go Green: New Evidence From Indonesia

September 23, 2026
Plant Diversity and Body Width Govern Hidden Soil Nematode Worlds
Earth Science

Plant Diversity and Body Width Govern Hidden Soil Nematode Worlds

September 23, 2026
Ancient Water Hides Deep Beneath Finland’s Buried Valleys, Tracers Reveal
Earth Science

Ancient Water Hides Deep Beneath Finland’s Buried Valleys, Tracers Reveal

September 23, 2026
Next Post
Welded Conductive Skin Turns Stretchy Rubber Into Ultra-Sensitive Wearable Sensor

Welded Conductive Skin Turns Stretchy Rubber Into Ultra-Sensitive Wearable Sensor

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Optica Foundation Honors Six Rising Stars in Optics and Photonics with 2026 Prizes and Fellowships
  • CPAP Machines Can Mistake Blocked Noses for a Dangerous Heart-Linked Breathing Pattern
  • Welded Conductive Skin Turns Stretchy Rubber Into Ultra-Sensitive Wearable Sensor
  • One AI Model, Any Device: Flexibly Slicable Network Cleans Up Speech from Earbuds to the Cloud

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading