<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>AI Flow &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/ai-flow/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 24 Sep 2026 23:16:36 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>AI Flow &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Graph Networks Turn Scattered Smart Speakers Into Powerful Microphone Arrays</title>
		<link>https://scienmag.com/graph-networks-turn-scattered-smart-speakers-into-powerful-microphone-arrays/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 23:16:36 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[ad-hoc microphone arrays]]></category>
		<category><![CDATA[AI Flow]]></category>
		<category><![CDATA[ambient noise suppression in smart devices]]></category>
		<category><![CDATA[channel selection]]></category>
		<category><![CDATA[collaborative microphone array technology]]></category>
		<category><![CDATA[deep learning for sound source localization]]></category>
		<category><![CDATA[edge computing]]></category>
		<category><![CDATA[equal error rate]]></category>
		<category><![CDATA[far-field speech processing]]></category>
		<category><![CDATA[far-field voice recognition challenges]]></category>
		<category><![CDATA[graph attention networks]]></category>
		<category><![CDATA[graph neural networks for speaker verification]]></category>
		<category><![CDATA[graph-based speech signal processing]]></category>
		<category><![CDATA[informed machine learning]]></category>
		<category><![CDATA[multi-channel audio]]></category>
		<category><![CDATA[multi-device speech processing]]></category>
		<category><![CDATA[multi-microphone collaboration algorithms]]></category>
		<category><![CDATA[noise reduction in smart home devices]]></category>
		<category><![CDATA[reverberation mitigation in voice recognition]]></category>
		<category><![CDATA[smart home]]></category>
		<category><![CDATA[Smart speaker microphone array enhancement]]></category>
		<category><![CDATA[spatial-temporal graph neural network]]></category>
		<category><![CDATA[speaker identification in noisy environments]]></category>
		<category><![CDATA[speaker verification]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=213107</guid>

					<description><![CDATA[Researchers have developed a spatial-temporal graph attention network that lets scattered smart devices collaborate as an ad-hoc microphone array, cutting far-field speaker verification error rates by up to 17.70 percent relative on real-world data.]]></description>
										<content:encoded><![CDATA[<p>Speaker verification, the technology that decides whether a voice belongs to a claimed identity, has quietly become one of the most important building blocks of modern life. It unlocks smartphones, guards bank accounts, and wakes up smart home hubs. Yet the technology has a persistent weakness: it works brilliantly when a microphone is close to the speaker&#8217;s mouth, and it falters badly when the voice must travel across a noisy, reverberant room. A new study published in the open-access journal Vicinagearth by Yijiang Chen, Chengdong Liang, Xiao-Lei Zhang and colleagues at Northwestern Polytechnical University and collaborating institutions tackles this far-field problem head-on, and its solution is as elegant as it is ambitious: treat every smart device in a room as a node in a graph, and let the devices learn to collaborate.</p>
<p>The core difficulty is physics. When a person speaks from across a room, the sound that reaches a distant microphone is attenuated, smeared by echoes bouncing off walls and furniture, and buried under background noise from fans, traffic, or televisions. Early speaker verification systems, dating back to the 1960s, relied on statistical models such as Gaussian mixture models with universal background models and later i-vectors. The deep learning era brought neural embeddings like x-vectors that dramatically improved accuracy, but the fundamental problem remained: a single distant microphone simply does not capture enough of the speaker&#8217;s characteristic vocal signature. Previous remedies included deep-learning speech enhancement front-ends that attempt to strip noise before verification, and domain adaptation techniques that treat noisy speech as a shifted version of clean speech. These help, but they leave a crucial resource untapped: spatial information.</p>
<p>Multi-channel approaches try to recover that spatial information by combining signals from several microphones. Fixed microphone arrays, the kind built into dedicated hardware, can apply beamforming algorithms that steer acoustic sensitivity toward the speaker. Researchers have combined neural beamforming with verification back-ends, fed multi-channel signals directly into convolutional networks, and encoded direction-of-arrival estimates into spatially aware speaker vectors. But fixed arrays have small apertures and fixed geometries. When the speaker is far away or moves around, even these sophisticated systems struggle. The truly interesting opportunity, the authors argue, lies in the devices people already own: smart speakers, phones, tablets, and other internet-connected gadgets scattered around a home or office, each with its own microphone.</p>
<p>This is the concept of the ad-hoc microphone array. Instead of a rigid cluster of microphones, an ad-hoc array is a loose federation of independently placed devices, each acting as an intelligent edge endpoint that can process audio locally. The architecture aligns with AI Flow, a decentralized artificial intelligence paradigm in which intelligent agents distributed across edge devices cooperate through coordinated computation and communication rather than shipping everything to a central server. The catch is that in an ad-hoc array, nobody controls where the devices sit. Some will be close to the speaker and capture clean audio; others will be far away, tucked behind furniture, or near a noise source, and their signals may actively harm the system. Earlier work on ad-hoc speaker verification used attention mechanisms to reweight all channels, but the new study makes a sharper observation: channels that are too noisy should not merely be down-weighted, they should be discarded.</p>
<p>The team&#8217;s framework, called a spatial-temporal graph attention network, or ST-GAT, reformulates the entire multi-channel problem as learning on a graph. In graph neural networks, data points become nodes and relationships between them become edges, described by an adjacency matrix. What makes this paper unusual is its treatment of time. Existing graph-based approaches to multi-channel audio modeled only the relationships between microphones, ignoring how information flows between successive time frames. The authors instead make each frame of each channel a node in a single static graph, so that both the spatial relationships between microphones and the temporal relationships between frames are captured in one structure. To their knowledge, this is the first time spatial-temporal data has been formulated as a static graph learning problem, a choice that is simpler, requires fewer parameters, and trains with standard graph neural network machinery.</p>
<p>There is a practical obstacle: the full adjacency matrix would be enormous. A ten-second recording from forty microphones, analyzed in ten-millisecond frames, produces a graph with forty thousand nodes and an adjacency matrix of forty thousand by forty thousand. The authors sidestep this by decomposing the aggregation into two successive modules. A temporal module builds a graph over the frames of each individual channel, and a spatial module builds a graph over the channels at each frame. They implemented the aggregation in two ways: a self-attention mechanism that uses the adjacency matrix as a mask over attention scores, and a true graph attention network in which each node attends to its neighbors using its own representation as the query. Both use multi-head attention, and the blocks are stacked to deepen the model.</p>
<p>The second innovation is a graph-based channel selection block that exploits prior knowledge during training. When the training data includes labeled positions of microphones, speakers, and noise sources, the authors construct an auxiliary adjacency matrix that connects only the most useful channels, for example the microphones closest to the speaker, or those with the best signal-to-noise ratio, while masking out nodes near a noise source or behind the speaker&#8217;s head. Crucially, this selection happens only during training. At test time, when the layout of devices and speakers is unknown, the network applies a fully connected graph and autonomously decides which channels to trust, because the attention parameters have learned the selection rule through backpropagation. The authors frame this as a form of informed machine learning, where handcrafted rules inject prior knowledge into the graph topology.</p>
<p>Training proceeds in two stages. First, a standard single-channel verification system, built from a residual convolutional network front-end, self-attentive pooling, and a classification layer, is trained on abundant single-channel speech. The frame-level feature extractor is then frozen, and the graph-based channel fusion modules are trained on spatial-temporal data from ad-hoc arrays. The evaluation was thorough: two simulated datasets, LibriSIMU-noise and LibriSIMU-reverb, generated with random room dimensions, reverberation times up to 1.2 seconds, and signal-to-noise ratios down to minus five decibels, plus two real-world corpora, Libri-adhoc40, a forty-node replayed array recorded in a highly reverberant office, and Hi-mia, a far-field text-dependent smart home dataset. Six representative baselines were compared, including oracle selection of the closest microphone, classical beamforming, energy-envelope channel selection, and attention-based multi-channel aggregation methods.</p>
<p>The results are striking. On the simulated datasets, the best proposed variant achieved a relative reduction in equal error rate of 15.39 percent compared with the strongest reference method, and on the real-world data the reduction reached 17.70 percent. Ablation studies showed that the auxiliary adjacency matrix brought especially large gains when combined with the ST-GAT backbone, cutting the error rate by roughly 15.65 percent relative on noisy simulated data and about 18.94 percent relative on the real Libri-adhoc40 corpus in the eight-channel scenario. The system remained robust across signal-to-noise ratios from minus five to twenty decibels and reverberation times up to 1.2 seconds, and it transferred to a different verification architecture, ECAPA-TDNN, with a further relative improvement of 12.3 percent over the best baseline. Analysis of the learned attention weights confirmed the mechanism: when trained with the distance-based or SNR-based auxiliary matrix, the channels receiving the highest attention weights were consistently the microphones physically closest to the speaker.</p>
<p>The implications reach well beyond the laboratory. As homes and offices fill with voice-capable edge devices, the ability to fuse their microphones into a virtual array, without centralizing raw audio and without knowing where the devices sit, points toward privacy-preserving, robust voice interfaces that work as well from across the room as they do at arm&#8217;s length. The authors are candid about open questions, notably that the adjacency matrix design improves the graph attention mechanism but not the plain self-attention variant, and they envision future extensions using signal-to-interference ratios for multi-speaker scenes and estimated SNR when distances are unknown. For now, the study demonstrates a compelling principle: when microphones learn to collaborate as a graph, the sum of many imperfect ears becomes a remarkably sharp listener.</p>
<p><strong>Subject of Research:</strong> Far-field speaker verification using spatial-temporal graph attention networks with ad-hoc microphone arrays</p>
<p><strong>Article Title:</strong> Edge-collaborative multi-channel speaker verification via spatial-temporal graph with ad-hoc microphone arrays</p>
<p><strong>Article References:</strong> Chen, Y., Liang, C., Chen, S., Feng, L., Zhu, B., Zhang, C., &amp; Zhang, X.-L. (2025). Edge-collaborative multi-channel speaker verification via spatial-temporal graph with ad-hoc microphone arrays. <em>Vicinagearth, 2</em>(1), Article 12. <a href="https://doi.org/10.1007/s44336-025-00023-y" rel="noopener noreferrer">https://doi.org/10.1007/s44336-025-00023-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-025-00023-y" rel="noopener noreferrer">10.1007/s44336-025-00023-y</a></p>
<p><strong>Keywords:</strong> speaker verification, ad-hoc microphone arrays, graph attention networks, far-field speech processing, edge computing, AI Flow, channel selection, spatial-temporal graph neural network, multi-channel audio, equal error rate, smart home, informed machine learning</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">213107</post-id>	</item>
		<item>
		<title>One AI Model, Any Device: Flexibly Slicable Network Cleans Up Speech from Earbuds to the Cloud</title>
		<link>https://scienmag.com/one-ai-model-any-device-flexibly-slicable-network-cleans-up-speech-from-earbuds-to-the-cloud/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 22:49:35 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[adaptive speech enhancement technology]]></category>
		<category><![CDATA[AI Flow]]></category>
		<category><![CDATA[AI model deployment across multiple devices]]></category>
		<category><![CDATA[BS-RoFormer]]></category>
		<category><![CDATA[denoising]]></category>
		<category><![CDATA[distributed AI for edge and cloud devices]]></category>
		<category><![CDATA[dynamic network slicing]]></category>
		<category><![CDATA[early exit]]></category>
		<category><![CDATA[edge computing]]></category>
		<category><![CDATA[end-to-end speech restoration solutions]]></category>
		<category><![CDATA[FlexAttention]]></category>
		<category><![CDATA[flexible neural networks]]></category>
		<category><![CDATA[flexible neural networks for wireless earbuds]]></category>
		<category><![CDATA[neural network model size reduction]]></category>
		<category><![CDATA[noise and reverberation suppression in audio]]></category>
		<category><![CDATA[open-access AI research on speech enhancement]]></category>
		<category><![CDATA[packet loss concealment]]></category>
		<category><![CDATA[real-world speech signal processing]]></category>
		<category><![CDATA[resource-constrained inference]]></category>
		<category><![CDATA[scalable neural network architecture]]></category>
		<category><![CDATA[SEFlow]]></category>
		<category><![CDATA[single training AI models for diverse hardware]]></category>
		<category><![CDATA[slimmable networks]]></category>
		<category><![CDATA[speech enhancement]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=210966</guid>

					<description><![CDATA[Researchers have built SEFlow, a single neural network that can be dynamically sliced into subnetworks of vastly different sizes to perform denoising, dereverberation, declipping, and packet loss concealment across devices from earbuds to cloud servers.]]></description>
										<content:encoded><![CDATA[<p>A single neural network that can shrink itself to one percent of its full size and still rescue speech buried in noise, reverberation, clipping, and lost data packets has been unveiled by researchers at Northwestern Polytechnical University and the Institute of Artificial Intelligence (TeleAI) at China Telecom. The system, called SEFlow, published in the open-access journal Vicinagearth, is designed for a future the authors describe as AI Flow: a vision of distributed intelligence in which devices, edge servers, and cloud computers collaborate seamlessly, each drawing on the same underlying model at whatever scale its hardware can afford.</p>
<p>The problem the team set out to solve is a familiar one for anyone deploying artificial intelligence in the real world. Modern deep networks achieve remarkable quality by growing ever larger, but a model that runs comfortably on a cloud server is hopeless on a pair of wireless earbuds. Traditionally, engineers have trained separate models for each computational tier, or resorted to pruning and knowledge distillation, both of which require additional fine-tuning after the fact. SEFlow takes a different route: a single network is trained once in such a way that it can be dynamically sliced into subnetworks of wildly different sizes and deployed directly, with no retraining, onto anything from a smartphone to a data center.</p>
<p>The technical heart of the approach is a set of flexible modules the authors call FlexAttention, FlexLinear, and FlexRMSNorm. In a conventional transformer, the width of each layer is fixed: a linear layer has a set number of inputs and outputs, and multi-head attention has a fixed number of heads. FlexLinear instead slices its weight matrix to any chosen output size and, crucially, activates only a chosen subset of its input neurons, following the same logic as early-exit methods that activate only early layers. FlexAttention adjusts the number of active attention heads, and FlexRMSNorm adapts the size of its normalization parameters. Because each output is a weighted sum of many neurons, slicing works mathematically without changing what the layer computes—only how much of it runs.</p>
<p>Depth is handled by early exit. The network is built from a stack of residual blocks whose outputs all share the same shape, so a decoder can read the result after any number of blocks. Because of how residual connections accumulate, decoding after block number B-bar is equivalent to decoding a fusion of the features produced by the first B-bar blocks; early layers capture the main content of the signal while later layers refine details, so cutting depth sacrifices fine polish but preserves the essentials. Adjusting depth changes computation roughly linearly, while adjusting width changes it quadratically, giving the deployer a two-dimensional dial for matching any computational budget.</p>
<p>Training such a shape-shifting network requires care. At every training step, the team duplicates each batch: one copy trains the full network, the other trains a randomly sampled subnetwork with a random depth and width, ensuring every neuron is updated regularly. To keep multi-GPU training efficient—since GPUs assigned smaller subnetworks would otherwise finish early and sit idle—a synchronized pseudo-random generator gives every GPU the same subnetwork index each step. According to the paper, this synchronization alone cut average training computation by about 38 percent and training time by about 35 percent.</p>
<p>SEFlow is not merely flexible; it is also unified. Instead of training separate models for denoising, dereverberation, declipping, and packet loss concealment, the team applied dynamic data augmentation: clean speech was corrupted with random combinations of noise at signal-to-noise ratios between minus 5 and 20 decibets, simulated room reverberation, waveform clipping, and Markov-chain-simulated packet loss. One model learned to handle all of these degradations, alone or together. A lightweight auxiliary voice activity detection decoder, and a loss combining complex-spectrogram and magnitude reconstruction with the VAD objective, further boosted quality across metrics such as PESQ, STOI, and downstream speaker and speech recognition scores.</p>
<p>The backbone builds on BS-RoFormer, the winning system of the NeurIPS 2024 speech enhancement challenge, splitting the frequency axis into 41 sub-bands distributed approximately uniformly on the Mel scale. A two-stage band-splitting scheme makes the model sampling-rate agnostic: the same architecture handles audio at 8, 16, 22.05, 24, 32, 44.1, and 48 kilohertz simply by using fewer sub-bands at lower rates. A Distributed Grouped Sampler keeps each training batch at a single sampling rate, avoiding wasteful upsampling and GPU idle time, reducing computation by a further 17 percent in their experiments.</p>
<p>The results are striking in their scalability. The full network has about 27 million parameters and costs roughly 24.7 GMACs per second on 16-kilohertz audio; the smallest subnetwork, with a single residual block and a single attention head, has 1.69 million parameters and costs about 0.19 GMACs per second—two orders of magnitude less—yet still measurably improves speech quality over the unprocessed input. At full size, SEFlow performs comparably to state-of-the-art task-specific models on denoising benchmarks from the INTERSPEECH 2020 DNS Challenge and on packet loss concealment benchmarks from the 2022 PLC Challenge, with demo-quality declipping shown on the project homepage. Interesting quirks emerged: a six-block, one-head configuration beat a one-block, four-head early-exit variant across all metrics despite using only 37 percent of its compute, suggesting width can matter more than depth in some regimes.</p>
<p>The authors are candid about trade-offs. Flexibly trained models slightly underperform fixed full-scale networks at the same nominal size, packet loss concealment demands at least two blocks for usable performance, and automatic speech recognition accuracy downstream remains limited, hinting that the backbone or loss may need task-specific tuning. There are also overheads in training multiple subnetworks simultaneously and in deciding which subnetwork to invoke at inference. Still, the researchers argue the approach carries real environmental promise: fixed models burn peak energy even on easy inputs, whereas adaptive slicing reduces multiply-accumulate operations and memory accesses—the dominant energy costs on edge chips—potentially trimming the carbon footprint of always-on speech processing.</p>
<p>What makes SEFlow resonate beyond acoustics is its implication for how AI might be delivered everywhere at once. Rather than a patchwork of bespoke models scattered across earbuds, phones, cars, and servers, a single family-model system could flow intelligence across the device-edge-cloud continuum, expanding and contracting to fit each platform. If the same recipe—flexible width, early exit, unified multi-task training—transfers to language, vision, and audio generation models, the paper&#8217;s vision of ubiquitous, resource-aware intelligence moves a step closer to reality. Demonstrations of SEFlow&#8217;s outputs, from heavily noisy cafe chatter to clipped and packet-mangled calls, are publicly available, and the code is available on request, inviting the community to stress-test this elegantly elastic architecture.</p>
<p><strong>Subject of Research:</strong> A flexibly scalable, unified neural architecture for multi-task speech enhancement</p>
<p><strong>Article Title:</strong> Towards a flexible and unified architecture for speech enhancement</p>
<p><strong>Article References:</strong> Feng, L., Zhang, C., &amp; Zhang, X.-L. (2025). Towards a flexible and unified architecture for speech enhancement. <em>Vicinagearth, 2</em>(1), Article 14. <a href="https://doi.org/10.1007/s44336-025-00022-z" rel="noopener noreferrer">https://doi.org/10.1007/s44336-025-00022-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-025-00022-z" rel="noopener noreferrer">10.1007/s44336-025-00022-z</a></p>
<p><strong>Keywords:</strong> speech enhancement, SEFlow, flexible neural networks, FlexAttention, early exit, AI Flow, edge computing, slimmable networks, packet loss concealment, denoising, BS-RoFormer, resource-constrained inference</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">210966</post-id>	</item>
		<item>
		<title>AI Flow Framework Aims to Bring Powerful Artificial Intelligence to Every Device</title>
		<link>https://scienmag.com/ai-flow-framework-aims-to-bring-powerful-artificial-intelligence-to-every-device/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sun, 13 Sep 2026 03:01:08 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[6G networks]]></category>
		<category><![CDATA[advancements in communication technology for AI]]></category>
		<category><![CDATA[AI Flow]]></category>
		<category><![CDATA[AI integration in small devices]]></category>
		<category><![CDATA[AI model compression techniques]]></category>
		<category><![CDATA[AI-powered edge computing]]></category>
		<category><![CDATA[bridging AI model size with device memory constraints]]></category>
		<category><![CDATA[challenges of deploying large AI models on mobile devices]]></category>
		<category><![CDATA[device-edge-cloud collaboration]]></category>
		<category><![CDATA[diffusion models]]></category>
		<category><![CDATA[edge AI]]></category>
		<category><![CDATA[familial models]]></category>
		<category><![CDATA[future of AI in wearable and IoT devices]]></category>
		<category><![CDATA[intelligence emergence]]></category>
		<category><![CDATA[large language model scalability]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[limitations of current AI hardware]]></category>
		<category><![CDATA[making AI accessible on smartphones and sensors]]></category>
		<category><![CDATA[multidisciplinary AI framework development]]></category>
		<category><![CDATA[speculative decoding]]></category>
		<category><![CDATA[task-oriented feature compression]]></category>
		<category><![CDATA[ubiquitous AI services]]></category>
		<category><![CDATA[ubiquitous intelligence]]></category>
		<category><![CDATA[vision-language models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201104</guid>

					<description><![CDATA[Researchers have introduced AI Flow, a framework combining device-edge-cloud collaboration, familial models, and networked intelligence emergence to make powerful AI accessible on resource-constrained devices.]]></description>
										<content:encoded><![CDATA[<p>A sweeping new framework called AI Flow promises to dissolve the barrier between today&#8217;s massive artificial intelligence models and the small devices people carry every day. In a comprehensive review published in the journal Vicinagearth, researchers at the Institute of Artificial Intelligence (TeleAI) at China Telecom, led by Xuelong Li, lay out a multidisciplinary blueprint that fuses advances in information technology and communication technology to deliver what they call ubiquitous intelligence: AI services that are fast, accessible, and available anywhere, from smartphones and sensors to drones and smart glasses. The work traces its intellectual lineage to Claude Shannon&#8217;s information theory and Alan Turing&#8217;s vision of machine intelligence, arguing that the long convergence of computing and communication has now reached a decisive moment with large AI models.</p>
<p>The core problem the researchers identify is a dual bottleneck. Modern large language models have grown from the roughly 12 to 60 million parameters of ResNet in 2016 to hundreds of billions or even trillions of parameters in systems like Llama-4, released in 2025. That hundredfold expansion in less than a decade means inference can demand tens to hundreds of gigabytes of memory, far beyond the 4 to 32 gigabytes typical of consumer devices. Compression techniques such as quantization and pruning help, but they trade away model capability. At the same time, communication networks strain under the load: split-inference schemes that ship high-dimensional activation features from devices to servers can generate tens to hundreds of megabytes per inference step, while multi-agent systems that synchronize reasoning traces amplify overhead further. Congestion, jitter, and wireless instability compound the challenge.</p>
<p>AI Flow responds with three interlocking pillars. The first is a device-edge-cloud architecture that treats the network itself as a computational hierarchy. End devices handle lightweight tasks and real-time interaction; edge servers at base stations and roadside units provide nearby, low-latency processing; and cloud clusters supply the scalable horsepower for training and compute-intensive inference. By orchestrating workloads across these tiers, the framework balances resource scalability against latency, offloading latency-critical inference to the edge while reserving the cloud for heavy operations.</p>
<p>Within that hierarchy, the team introduces two collaboration techniques designed to cut communication costs. The first, task-oriented feature compression, targets vision-language model inference. Rather than transmitting raw images, the device merges visual features produced by a CLIP-style encoder using density peaks clustering based on K nearest neighbors, then encodes the merged features with a hyperprior-based entropy model whose parameters are modeled on a Laplacian distribution. A router network selects the best entropy model for each feature. In experiments on the LLaVA-OneVision-7B model using an NVIDIA Jetson AGX Orin device and an RTX 4090 edge server, the method reduced transmitted data by 25 to 45 percent compared with WebP and 35 to 60 percent compared with JPEG at equal accuracy on the RealWorldQA benchmark, and cut inference latency to roughly a third of server-only inference on the MME benchmark.</p>
<p>The second technique, hierarchical collaborative decoding, accelerates large language model generation through speculative decoding spread across network tiers. A lightweight model on the device drafts tokens locally, while a larger edge model validates and corrects them using a soft acceptance strategy, shifting the big model&#8217;s role from full generation to error correction. The researchers extend this into a parallel pipeline in which the device keeps generating without blocking while the edge server refines tokens at intervals. On the MATH-500 benchmark, a two-tier configuration pairing a 1.5-billion-parameter device model with a 7-billion-parameter edge model achieved about 40 tokens per second, a 1.25-fold speedup over edge-only decoding at the same accuracy, with a three-tier setup adding a 14-billion-parameter cloud model for further gains.</p>
<p>The second pillar of AI Flow is the concept of familial models: families of different-sized models whose hidden features are aligned so that intermediate results from a small model can be directly reused by a larger one without any middleware. Two enabling techniques make this possible. Early exit allows inference to terminate at intermediate layers while preserving acceptable accuracy, with lightweight branch modules refining features before prediction. Weight decomposition splits the linear layers of transformer blocks into pairs of low-rank matrices whose combined parameter count is smaller than the original, with the hidden dimension tuned to hit nearly any target size. Initialization via singular value decomposition on whitened data keeps distortion low, and the team shows that compression loss is quantitatively determined by the squared singular values of discarded components, enabling per-layer compression decisions.</p>
<p>The researchers demonstrate two implementation strategies. Hierarchical principal component decomposition trains a series of low-rank components that progressively fit the residuals of earlier ones, producing TeleChat-based models from 2.38 billion to 6.30 billion parameters that, despite limited training tokens, perform comparably to established models such as LLaMA2-7B and ChatGLM2-6B on benchmarks including MMLU, CMMLU, C-Eval, GSM8K, MATH, and BBH. The second strategy, early exiting with scalable branches, inserts decomposed transformer blocks between exit points and a shared language model head. Applied to LLaVA-1.5-7B, it retained 98.2 percent of the backbone&#8217;s average performance on six visual question answering benchmarks using only 3.17 billion parameters, while a baseline without the branch design needed at least 4.63 billion parameters to reach 90 percent.</p>
<p>The third pillar is perhaps the most provocative: connectivity- and interaction-based intelligence emergence. Here, the network becomes a medium through which heterogeneous models, including large language models, vision-language models, and diffusion models, collaborate to achieve capabilities exceeding any single model. A device-server collaboration scheme lets specialized on-device models generate preliminary responses in parallel, which a central server model aggregates into a unified answer that is then returned to devices for revision. Evaluations on MT-Bench, AlpacaEval 2.0, and Arena-Hard showed consistent gains, with weaker models benefiting most, and performance on Arena-Hard rose nearly linearly with the number of participating agents, suggesting practical scalability.</p>
<p>Diffusion models receive their own collaboration paradigms. A serial scheme for multi-person motion generation chains an interleaved interaction synthesis module with a relative coordination refinement module, achieving state-of-the-art results on the InterHuman benchmark, including a 25.3 percent improvement in Top-1 R-Precision and a 50.6 percent reduction in Fréchet inception distance compared with prior methods. A parallel scheme for monocular depth estimation splits processing into near-field and far-field decoder branches fused around a sliding anchor, topping benchmarks on both indoor NYU-V2 and outdoor KITTI data. A networked scheme, OmniVDiff, unifies RGB, depth, segmentation, and edge modalities within a single video diffusion transformer, outperforming baselines on depth-conditioned video generation.</p>
<p>The authors ground the framework in application scenarios that include embodied AI, where drones and ground robots share aligned intermediate features to avoid redundant computation; wearable devices, where smart glasses offload heavy recognition tasks to edge and cloud tiers while keeping latency-sensitive processing local; and smart cities, where the low-altitude economy of delivery drones and aerial mobility systems demands ultra-low-latency coordination across thousands of heterogeneous devices. Future directions include federated learning adapted to large models, distributed edge inference resilient to device churn, and adaptive network orchestration for volatile wireless conditions. The team also articulates guiding principles, including a Law of Information Capacity that defines efficiency as the ratio of text compression gain to inference cost, and a Law of Multi-model Collaboration showing that ensembles of diverse models follow power-law scaling with a better loss floor than single-series collaboration. Together, the researchers argue, these ideas chart a path toward AI that is not confined to data centers but flows through the networks that already surround us.</p>
<p><strong>Subject of Research:</strong> A multidisciplinary framework integrating AI and communication technologies for ubiquitous, low-latency intelligence across device-edge-cloud networks</p>
<p><strong>Article Title:</strong> AI Flow: perspectives, scenarios, and approaches</p>
<p><strong>Article References:</strong> An, H., Hu, W., Huang, S., Huang, S., Li, R., Liang, Y., Shao, J., Song, Y., Wang, Z., Yuan, C., Zhang, C., Zhang, H., Zhuang, W., &amp; Li, X. (2026). AI Flow: perspectives, scenarios, and approaches. <em>Vicinagearth, 3</em>(1), Article 1. <a href="https://doi.org/10.1007/s44336-025-00031-y" rel="noopener noreferrer">https://doi.org/10.1007/s44336-025-00031-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-025-00031-y" rel="noopener noreferrer">10.1007/s44336-025-00031-y</a></p>
<p><strong>Keywords:</strong> AI Flow, edge AI, device-edge-cloud collaboration, familial models, large language models, speculative decoding, intelligence emergence, task-oriented feature compression, vision-language models, diffusion models, ubiquitous intelligence, 6G networks</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201104</post-id>	</item>
		<item>
		<title>Cloud-Edge AI System Translates Speech While Protecting Speaker Identity</title>
		<link>https://scienmag.com/cloud-edge-ai-system-translates-speech-while-protecting-speaker-identity/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 20:25:31 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[adaptive computational architecture for multilingual speech translation]]></category>
		<category><![CDATA[AI Flow]]></category>
		<category><![CDATA[AI-driven voice anonymization techniques]]></category>
		<category><![CDATA[Cloud-Edge AI speech translation]]></category>
		<category><![CDATA[cloud-edge collaboration]]></category>
		<category><![CDATA[cloud-edge collaboration in speech translation systems]]></category>
		<category><![CDATA[collaborative cloud-edge AI for speech-to-speech translation]]></category>
		<category><![CDATA[early exit]]></category>
		<category><![CDATA[edge computing]]></category>
		<category><![CDATA[heterogeneous device support for speech translation]]></category>
		<category><![CDATA[machine translation]]></category>
		<category><![CDATA[multilingual speech translation without raw voice transfer]]></category>
		<category><![CDATA[neural network model scalability for edge devices]]></category>
		<category><![CDATA[neural networks]]></category>
		<category><![CDATA[privacy]]></category>
		<category><![CDATA[privacy and efficiency in real-time speech translation]]></category>
		<category><![CDATA[privacy-aware voice data processing in AI translation]]></category>
		<category><![CDATA[resource-constrained inference]]></category>
		<category><![CDATA[secure voice data transmission in AI translation pipelines]]></category>
		<category><![CDATA[speaker identity]]></category>
		<category><![CDATA[speaker privacy preservation in neural translation systems]]></category>
		<category><![CDATA[speech synthesis]]></category>
		<category><![CDATA[speech translation]]></category>
		<category><![CDATA[voice preservation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=198320</guid>

					<description><![CDATA[A new cloud-edge collaborative framework adapts speech-to-speech translation to any device's compute budget while preserving speaker identity through a privacy-preserving retrieval system.]]></description>
										<content:encoded><![CDATA[<p>Speech-to-speech translation has long promised a world in which language barriers simply dissolve: a traveler speaks Japanese into a phone and their companion hears fluent English, a business negotiator conducts a multilingual conference call without an interpreter, and a filmmaker dubs content across dozens of languages. Modern neural systems have grown remarkably good at translating the words themselves. Yet two stubborn problems have limited real-world deployment. Most research teams release only a single model size, forcing every device—from a smartwatch to a data center—to run the same network regardless of its computing budget, and the standard trick for preserving a speaker&#8217;s distinctive voice requires feeding sensitive acoustic data directly into the translation pipeline, raising uncomfortable privacy questions. A new study published in the journal Vicinagearth tackles both problems at once with a cloud-edge collaborative architecture that adapts its own computational effort to the hardware it runs on while never transmitting the speaker&#8217;s raw voice across the network.</p>
<p>The research, led by Boyu Zhu and Xiao-Lei Zhang of Northwestern Polytechnical University together with Rujin Chen and Chi Zhang of the Institute of Artificial Intelligence at China Telecom, builds on the AI Flow framework, which reconceives the network&#8217;s job as transmitting intelligence flows rather than raw information flows. In this design, edge devices, edge servers, and cloud servers share an inference pipeline, each contributing the compute it has available. Applied to speech translation, the system splits into three cooperating modules: a speech-to-text translation module distributed between the sender&#8217;s device and the cloud, a voice preservation module running entirely on the sender&#8217;s device, and a text-to-speech synthesis module on the receiver&#8217;s device. The result is a pipeline that translates source speech into target-language text, retrieves a compact identity reference, and regenerates natural target speech that carries the original speaker&#8217;s acoustic character.</p>
<p>The central technical innovation is a family of early-exit heads attached to the translation backbone, a strategy borrowed from efficient inference research on models such as HuBERT and Whisper. Rather than forcing every input through all decoder layers, an early-exit system evaluates its own uncertainty after each layer and stops when confidence is high enough. In the traditional version of this strategy, the team computes the Shannon entropy of the predicted token distribution after each decoder layer; if entropy falls below a fixed threshold, computation terminates immediately, saving the cost of the remaining layers. Static evaluations at layers 16, 18, 20, and 24 showed that usable translations emerge well before the full model completes its pass, meaning the same trained network can serve phones, laptops, and servers with very different resources simply by deciding where to stop.</p>
<p>But a simple entropy threshold is blunt: easy sentences exit too late and hard ones exit too early, degrading translation quality. The team&#8217;s solution is a small language model that learns to predict, for each input, which exit layer will produce the best result. Training this predictor required difficulty labels, which the authors generated with a teacher-guided classifier built on the GPT-4o API. The large language model scored each training sample on a one-to-ten difficulty scale, focusing on the rarity and frequency of long-tail vocabulary, structural divergence between source and target languages, and ambiguity or context dependence. Because difficulty judgments can shift systematically across language pairs, the researchers calibrated the scores statistically—standardizing within each pair and applying quantile mapping to a common reference distribution—so that a score of seven in French-to-English means the same thing as a seven in Polish-to-German.</p>
<p>Once the difficulty labels existed, the team trained a Flan-T5-based small model that takes the output of an automatic speech recognition pass and predicts the optimal early-exit layer for the translation model. During collaborative inference, the large translation model sits in the cloud while a pruned small model runs on the sender&#8217;s edge device, both fed by a lightweight convolutional preprocessing module that converts raw waveforms into compact feature tensors instead of streaming audio. The small model applies an entropy test: if it exits confidently, it raises a flag and the receiver synthesizes speech immediately from the local translation. If not, the sender waits up to a bounded number of seconds for the cloud result, falling back to the local output if the network stalls—guaranteeing bounded latency under real network conditions.</p>
<p>The second major contribution addresses voice preservation without privacy leakage. Conventional expressive S2ST systems pass speaker-related acoustic embeddings into the translation model itself, exposing biometric voice information to the cloud. The new system instead performs retrieval-based voice preservation. On the sender&#8217;s device, a lightweight pipeline extracts three kinds of features from the input speech: a speaker embedding generated by a RawNet3-based verification model, an emotion category from the emotion2vec+ representation, and a speaking rate computed as syllables per unit of voiced duration using energy-based peak detection. These features drive a hierarchical search through a large multilingual acoustic database built from the M3PDB dataset, which contains multiple speakers per language with samples spanning varied emotions and speaking rates.</p>
<p>The hierarchical matching narrows candidates in stages: first the speaker embedding locates the most acoustically similar speaker, then emotion category must match exactly, and finally the speech-rate similarity selects the reference sample whose cadence most closely mirrors the input. Only the identifier of that winning reference—never the voice itself—is transmitted to the receiver. The receiving device looks up the corresponding speech tokens, extracted with the tokenizer from IndexTTS2, and conditions its synthesis on both the translated text and this token, reconstructing target speech that echoes the source speaker&#8217;s timbre, emotion, and pacing without any raw acoustic data ever leaving the sender.</p>
<p>Experimental validation spanned two fronts. On the CoVoST2 French translation benchmark, evaluated with BLEU scores on Whisper-Large-V2, the SLM-assisted early-exit strategy outperformed entropy-, confidence-, and cosine-similarity-based early-exit baselines, indicating that predicting the optimal exit depth preserves translation quality better than uncertainty thresholds alone. For voice preservation, the team constructed demanding test sets from LibriTTS samples corrupted with noise at signal-to-noise ratios of negative five, five, and fifteen decibels, drawing noise from AudioSet, FreeSound, WHAM!, and FSD50K sources, with reverberation added probabilistically using simulated room impulse responses. Across these conditions, the retrieval-based method achieved lower word error rates and higher UTMOS naturalness scores than a baseline that transmitted the source speech directly, with the advantage widening as noise increased—a striking demonstration that a compact retrieved reference can be more robust than the original recording.</p>
<p>Cross-lingual generalization tests on the VoxPopuli corpus covered sixteen European languages, including English, German, French, Spanish, Polish, Italian, Romanian, Hungarian, Czech, Dutch, Finnish, Croatian, Slovak, Slovene, Estonian, and Lithuanian. Word error rates dropped for fourteen of the sixteen languages relative to the direct-transmission baseline, while speaking-rate consistency held nearly constant at the looser tolerance threshold and emotion consistency showed no substantial overall difference. The authors note that although objective speaker-similarity scores dipped below the baseline in noisy conditions, human listeners are typically far less sensitive to such differences than automated metrics suggest, so the perceptual gap is likely smaller than the numbers imply.</p>
<p>The broader significance of this work lies in what it reframes rather than what it merely accelerates. By treating model size as a deployment decision rather than a design constraint, the early-exit architecture lets one trained network flexibly serve the full spectrum of hardware, from battery-constrained earbuds to cloud GPUs. By relocating speaker identity from the model&#8217;s input to a retrieval lookup, it decouples expressive fidelity from biometric exposure, addressing a concern that will only grow as voice interfaces proliferate. And by showing that a few transmitted identifiers can outperform full audio transmission under noise, the study makes a compelling case that the future of real-time speech translation may rest less on bigger models than on smarter collaboration between the devices we carry and the servers that await our hardest questions. The framework, the authors conclude, offers a practical path toward deploying speech translation where it matters most: out in the noisy, bandwidth-limited, privacy-sensitive real world.</p>
<p><strong>Subject of Research:</strong> A cloud-edge collaborative speech-to-speech translation system with early-exit inference and retrieval-based voice preservation.</p>
<p><strong>Article Title:</strong> Speech to speech translation system based on cloud-edge collaboration</p>
<p><strong>Article References:</strong> Zhu, B., Chen, R., Zhang, C., &amp; Zhang, X.-L. (2026). Speech to speech translation system based on cloud-edge collaboration. <em>Vicinagearth, 3</em>(1), Article 4. <a href="https://doi.org/10.1007/s44336-026-00033-4" rel="noopener noreferrer">https://doi.org/10.1007/s44336-026-00033-4</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-026-00033-4" rel="noopener noreferrer">10.1007/s44336-026-00033-4</a></p>
<p><strong>Keywords:</strong> speech translation, cloud-edge collaboration, early exit, voice preservation, privacy, speech synthesis, neural networks, edge computing, machine translation, speaker identity, AI Flow, resource-constrained inference</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">198320</post-id>	</item>
		<item>
		<title>AI-Powered Video Compression Nears the 0.01% Frontier</title>
		<link>https://scienmag.com/ai-powered-video-compression-nears-the-0-01-frontier/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Fri, 11 Sep 2026 00:23:47 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[AI Flow]]></category>
		<category><![CDATA[AI-based content reconstruction]]></category>
		<category><![CDATA[AI-powered video data reduction]]></category>
		<category><![CDATA[bitrate]]></category>
		<category><![CDATA[data efficiency in communication systems]]></category>
		<category><![CDATA[diffusion models]]></category>
		<category><![CDATA[edge computing]]></category>
		<category><![CDATA[future of AI-driven multimedia compression]]></category>
		<category><![CDATA[generative models in video transmission]]></category>
		<category><![CDATA[generative video compression]]></category>
		<category><![CDATA[innovative video codecs]]></category>
		<category><![CDATA[LPIPS]]></category>
		<category><![CDATA[neural encoder]]></category>
		<category><![CDATA[pixel-level coding alternatives]]></category>
		<category><![CDATA[revolutionary video compression techniques]]></category>
		<category><![CDATA[Shannon-Weaver communication theory]]></category>
		<category><![CDATA[Shannon-Weaver model]]></category>
		<category><![CDATA[surveillance]]></category>
		<category><![CDATA[surveillance video compression]]></category>
		<category><![CDATA[task-oriented communication]]></category>
		<category><![CDATA[ultra-low bitrate video transmission]]></category>
		<category><![CDATA[video compression]]></category>
		<category><![CDATA[video transmission]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=192093</guid>

					<description><![CDATA[Researchers have introduced Generative Video Compression, a framework that transmits video at rates as low as 0.02% and below 0.01% in some scenarios by encoding compact tokens and using generative AI models at the receiver to reconstruct high-quality footage.]]></description>
										<content:encoded><![CDATA[<p>A team of researchers at the Institute of Artificial Intelligence (TeleAI), China Telecom, has unveiled a radically new way to squeeze video down to a tiny fraction of its original size, and the results are turning heads across the worlds of communications and artificial intelligence. The technique, called Generative Video Compression, or GVC, pushes bitrates to levels that traditional codecs cannot approach: in some cases the transmitted data amounts to just 0.02% of the original video, and in surveillance scenarios the team reports compression rates beyond the long-sought 0.01% threshold. Rather than refining the familiar machinery of pixel-level coding, GVC hands the hard work of reconstruction to a generative video model waiting at the receiving end, converting the act of communication from copying pixels into describing content and letting artificial intelligence fill in the rest.</p>
<p>The conceptual foundation of the new framework departs sharply from the dominant view of video compression that has prevailed for decades. Classical communication theory, rooted in the Shannon-Weaver model articulated by Claude Shannon in 1948, distinguishes three levels of communication. Level A concerns the technical problem of transmitting data accurately, Level B addresses whether transmitted symbols convey the intended meaning, and Level C concerns the effectiveness problem: whether the received information produces the desired outcome. Video standards such as HEVC have concentrated almost exclusively on Level A, maximizing signal fidelity under constrained bandwidth by minimizing distortion between the original and the reconstructed signal. GVC instead places Level C at the heart of its design, asking not whether every pixel survives the journey but whether the reconstructed video meets perceptual expectations or supports the task at hand.</p>
<p>The driving principle behind GVC is elegantly simple to state: trade computation for compression rate. Instead of transmitting detailed visual data, the framework encodes video into extremely compact representations and delegates content reconstruction to the receiver, where powerful generative priors synthesize high-quality video from minimal transmitted information. The researchers offer a vivid metaphor to explain the shift. Traditional compression is like photographing a painting and sending the photograph; GVC is like describing the painting&#8217;s composition and style, then relying on an AI painter at the far end to recreate it. Modern generative video models are so expressive that they can synthesize convincing footage from sparse latent representations, or in the limit even from pure noise guided by learned priors. That capability transforms the encoder&#8217;s job from preserving every pixel to selecting and transmitting only the most task-relevant information.</p>
<p>The system is built from two primary components working in tandem. On the sending side, a neural encoder, a pre-trained neural network, ingests an input video sequence, which might be surveillance footage, a video call stream, or a live broadcast, and compresses it into a set of compact representations called compressed tokens. These tokens blend discrete and continuous elements: compressed keyframes, high-level descriptors of video segments, and low-level continuous features that together capture the essential semantics and motion dynamics of the scene while drastically reducing dimensionality. The tokens are further encoded into a bitstream using techniques such as residual coding to squeeze out remaining redundancy. On the receiving side, a pre-trained diffusion-based generative video model performs what is essentially a conditional video generation task. Some tokens serve as direct inputs to the denoising process while others act as conditioning signals, and the model synthesizes frames that are visually faithful to the original input.</p>
<p>What gets transmitted depends, critically, on the purpose of the reconstruction. If the goal is human perception, the encoder sends features that help the generative decoder produce perceptually similar content. If the goal is machine understanding, for instance segmentation or recognition by a downstream algorithm, the encoder focuses on semantically meaningful representations instead. This task-oriented orientation is where GVC aligns itself with the AI Flow framework, proposed by TeleAI at the end of 2024, which envisions communication networks distributing intelligence for ubiquitous AI-powered services. The theoretical underpinning draws on the concept of Information Capacity, a measure of how efficiently generative models compress data, as well as earlier work on task-oriented feature compression for multimodal understanding via device-edge co-inference. The GVC concept itself was first introduced publicly by TeleAI at the World Artificial Intelligence Conference in mid-2025, where a prototype for maritime communications demonstrated ultra-low bitrate video transmission over bandwidth-limited satellite links.</p>
<p>Extreme compression, however, introduces a new bottleneck: the computational cost of high-quality generative reconstruction. Diffusion-based decoders are computationally intensive, and hardware, power, and latency constraints impose an upper bound on how much computation can realistically be traded for compression, especially in real-time applications like video conferencing or edge-device streaming. The team&#8217;s answer is a second, complementary principle: trading compression rate for practicality. By sacrificing a small fraction of the compression ratio, the system can send richer latent representations that reduce reliance on massive generative models, unlocking the use of smaller and faster decoders. The researchers further apply model compression techniques to shrink key components such as 3D variational autoencoders, and they employ distillation and sampling acceleration methods for the diffusion-based decoder to lower inference time. The result is a flexible balance across the compression-computation-quality triangle that adapts to whatever resources the deployment environment offers.</p>
<p>The empirical results are striking. Benchmarked on the standard MCL-JCV dataset using a 14-billion-parameter video generative model, GVC maintained competitively high perceptual quality, measured with the Learned Perceptual Image Patch Similarity metric, at an average bitrate of just 0.008 bits per pixel, equivalent to roughly a 0.02% compression rate. Conventional video coding schemes exhibit a substantial performance gap at this bitrate; on certain challenging sequences, traditional methods require approximately six times more bandwidth to match the perceptual quality achieved by GVC. In one benchmark example, the framework achieved visually compelling reconstruction at 0.005 bits per pixel. In real-world surveillance scenarios, the compression can go further still: the team reports bitrates below 0.002 bits per pixel, crossing the 0.01% compression threshold, while retaining sufficient visual quality for the task.</p>
<p>Importantly, the extreme compression does not come at the expense of semantic integrity. To test downstream utility, the researchers applied the reconstructed videos to video object segmentation on the DAVIS2017 benchmark, evaluating performance with the Jaccard index, contour accuracy, their average, and contour recall. The compressed-and-regenerated videos achieved highly competitive segmentation results, indicating that even at astonishingly low bitrates the framework preserves the semantic information machines need to understand a scene. This validates the core promise of effectiveness-level communication: the reconstructed video is not merely visually plausible but genuinely useful for the tasks the transmission was intended to serve.</p>
<p>Deployment readiness was demonstrated on real hardware. After miniaturization, distillation, and quantization of the generative decoder, the system can reconstruct a group of 29 frames in a single pass with inference latency of around two seconds on consumer-grade GPUs, a response time comparable to what users routinely experience with large language models. Although the miniaturized model incurs some loss in visual quality and bandwidth efficiency relative to its full-scale counterpart, it still maintains competitively high perceptual quality, with a demonstrated LPIPS score of 0.273 on a sample sequence. That combination of speed, quality, and modest hardware requirements makes GVC a plausible candidate for the environments that need it most: emergency rescue operations, remote surveillance, narrowband mobile networks, in-vehicle and wearable devices, and maritime satellite links where bandwidth is scarce and expensive.</p>
<p>The authors frame GVC not merely as another codec but as a task-oriented communication paradigm tailored for the era of generative intelligence. By transmitting only what is necessary for perception and decision making, and letting generative priors at the receiver do the heavy lifting of reconstruction, the framework opens the door to communication systems that are more efficient, adaptive, and intelligent than the fidelity-obsessed pipelines of the past. Whether video transmission at one hundredth of one percent of its original size becomes a routine capability will depend on further advances in generative model efficiency and edge computing, but this work offers a credible, empirically validated path toward that frontier, and a glimpse of a future in which the networks we build carry descriptions rather than copies, and understanding rather than pixels.</p>
<p>The work appears as a brief communication in Vicinagearth, an open-access journal, published on 23 March 2026 as volume 3, article number 7, with a correction issued on 3 June 2026. Its placement at the intersection of coding and information theory, computer vision, and multimedia systems reflects the increasingly hybrid nature of compression research, where ideas from generative modeling are being grafted onto classical transmission problems.</p>
<p>Historically, the Shannon-Weaver model dates to the 1940s, and the authors note that video communication technology has spent decades optimizing its Level A, the technical problem of accurate signal delivery. The Information Capacity metric, proposed as a way to evaluate how effectively generative models compress data, laid methodological groundwork for this line of research, and in early 2025 the same group extended the approach to task-oriented communications for multimodal understanding via device-edge co-inference. The maritime prototype unveiled at the World Artificial Intelligence Conference demonstrated ultra-low bitrate transmission over bandwidth-limited satellite connections, a setting where every saved bit carries direct operational value.</p>
<p>Within the AI Flow framework, the researchers position GVC as opening new possibilities for video communication in bandwidth- and resource-constrained environments such as emergency rescue, remote surveillance, and mobile edge computing, describing it as a viable path toward an effective, efficient, scalable, and practical video communication paradigm.</p>
<p><strong>Subject of Research:</strong> Extreme low-bitrate video compression using generative AI models to reconstruct video from minimal transmitted information</p>
<p><strong>Article Title:</strong> Generative video compression: towards 0.01% compression rate for video transmission</p>
<p><strong>Article References:</strong> Generative video compression: towards 0.01% compression rate for video transmission. (n.d.). <a href="https://doi.org/10.1007/s44336-026-00035-2" rel="noopener noreferrer">https://doi.org/10.1007/s44336-026-00035-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-026-00035-2" rel="noopener noreferrer">10.1007/s44336-026-00035-2</a></p>
<p><strong>Keywords:</strong> generative video compression, video compression, task-oriented communication, AI Flow, diffusion models, bitrate, Shannon-Weaver model, edge computing, video transmission, LPIPS, surveillance, neural encoder</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">192093</post-id>	</item>
	</channel>
</rss>
