<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Earth Mover&#8217;s Distance &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/earth-movers-distance/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Mon, 05 Oct 2026 05:45:32 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Earth Mover&#8217;s Distance &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Earth Mover&#8217;s Distance Steers Smarter Knowledge Sharing in Federated Learning</title>
		<link>https://scienmag.com/earth-movers-distance-steers-smarter-knowledge-sharing-in-federated-learning/</link>
		
		<dc:creator><![CDATA[Veronica Carney]]></dc:creator>
		<pubDate>Mon, 05 Oct 2026 05:45:32 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[CIFAR-100]]></category>
		<category><![CDATA[client drift]]></category>
		<category><![CDATA[client drift mitigation in federated AI]]></category>
		<category><![CDATA[data heterogeneity]]></category>
		<category><![CDATA[decoupled knowledge distillation]]></category>
		<category><![CDATA[decoupled knowledge distillation for federated models]]></category>
		<category><![CDATA[distributed machine learning]]></category>
		<category><![CDATA[distributional distance metrics in federated learning]]></category>
		<category><![CDATA[Earth Mover's Distance]]></category>
		<category><![CDATA[Earth Mover's Distance in distributed machine learning]]></category>
		<category><![CDATA[enhancing AI collaboration across devices]]></category>
		<category><![CDATA[federated learning]]></category>
		<category><![CDATA[federated learning data heterogeneity]]></category>
		<category><![CDATA[federated learning with non-i.i.d. data]]></category>
		<category><![CDATA[federated meta-learning techniques]]></category>
		<category><![CDATA[improved knowledge sharing in federated systems]]></category>
		<category><![CDATA[knowledge distillation]]></category>
		<category><![CDATA[meta-learning]]></category>
		<category><![CDATA[model aggregation]]></category>
		<category><![CDATA[model aggregation strategies in heterogeneous environments]]></category>
		<category><![CDATA[multi-device AI model training challenges]]></category>
		<category><![CDATA[non-IID data]]></category>
		<category><![CDATA[privacy-preserving AI]]></category>
		<category><![CDATA[scalable federated learning frameworks]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=236988</guid>

					<description><![CDATA[Researchers have developed FedML-DKD, a federated learning framework that uses Earth Mover's Distance to adaptively balance knowledge distillation and meta-learning, boosting accuracy by up to 10.87 percentage points under highly heterogeneous data.]]></description>
										<content:encoded><![CDATA[<p>Federated learning has long promised a world in which smartphones, hospitals, and factories can jointly train powerful artificial intelligence models without ever shipping their raw data to a central server. Yet the reality on the ground has been messier than the vision. Each participating device holds a sliver of data that reflects only its own users and environment, so the statistical landscapes seen by individual clients can differ wildly from one another. A new study published in Cluster Computing by Debao Wang and Shaopeng Guan of Shandong Technology and Business University tackles this stubborn problem, known as data heterogeneity, with a framework called FedML-DKD that combines two ideas from the machine learning toolbox: decoupled knowledge distillation and federated meta-learning, tied together by a distributional yardstick known as the Earth Mover&#8217;s Distance.</p>
<p>The core difficulty the researchers set out to address is client drift. In a standard federated training round, each client updates a shared model on its own local data before sending the updated parameters back to a server for aggregation. When those local datasets are skewed, for example when one phone contains mostly images of cats while another holds mostly images of cars, each client pulls the model in a different direction. The averaged global model ends up oscillating between incompatible preferences, converging slowly and often settling at a worse solution than a model trained on pooled data would reach. This non-independent-and-identically-distributed data problem has spawned an entire subfield of research, and knowledge distillation has emerged as one of its more promising remedies.</p>
<p>Knowledge distillation, in its classic form, asks a smaller or newer model to mimic the output probabilities of a teacher model, transferring not just the correct answers but the teacher&#8217;s nuanced confidence across all classes. In federated settings, the global model on the server typically acts as the teacher for local students. But Wang and Guan point out a subtle weakness in the conventional, coupled formulation of distillation when data are highly heterogeneous. The distillation loss blends two distinct signals: the knowledge about the target class, meaning the probability assigned to the correct label, and the knowledge about non-target classes, the relative ordering of all the wrong answers. Under skewed local distributions, a client can develop biased confidence in its dominant target classes, and that bias contaminates the coupled loss, misguiding the learning process rather than correcting it.</p>
<p>Their remedy builds on decoupled knowledge distillation, a technique introduced at the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, which separates the distillation objective into a target-class component, abbreviated TCKD, and a non-target-class component, abbreviated NCKD, each weighted independently. Decoupling alone, however, uses fixed weights that ignore how heterogeneous a given client&#8217;s data actually are. The key innovation of FedML-DKD is a heterogeneity-aware weighting mechanism that adjusts the balance between TCKD and NCKD dynamically for each client. The measure used to gauge heterogeneity is the Earth Mover&#8217;s Distance, a metric with roots in computer vision and transportation theory that quantifies the minimum work needed to transform one probability distribution into another, like measuring how much soil must be moved to reshape one pile of earth into the shape of another.</p>
<p>In practice, the framework computes the distributional divergence between each client&#8217;s local label distribution and a reference distribution, then uses that EMD value to set the relative contributions of the two distillation components. Clients whose data deviate sharply from the global distribution receive a rebalanced distillation signal that prevents them from overfitting to their biased local samples, while clients with more representative data can lean more heavily on the standard knowledge transfer pathway. This adaptive decoupling improves the reliability of the knowledge flowing between server and clients, because the teacher&#8217;s guidance is no longer distorted by the statistical quirks of any single participant. The result, according to the authors, is a training process that remains stable even as the degree of heterogeneity across clients increases.</p>
<p>The second pillar of FedML-DKD concerns what happens on the server after clients upload their updates. Conventional federated learning aggregates those parameters, typically through weighted averaging, to produce the next global model. Wang and Guan instead replace this aggregation step with a meta-learning objective. Meta-learning, often described as learning to learn, seeks an initialization from which a model can rapidly adapt to new tasks with only a few gradient steps. In the federated context, the uploaded client parameters are treated as a shared initialization that must generalize across the diverse local tasks represented by the participating devices. Rather than simply averaging away the differences between clients, the meta-learning objective explicitly optimizes for a starting point that can be fine-tuned quickly and effectively on each client&#8217;s particular distribution.</p>
<p>This design choice addresses a persistent tension in federated systems between personalization and generalization. A single global model must serve clients with very different data profiles, and pure averaging can wash out the specialized competence each client develops locally. By framing the global model as a meta-learned initialization, the framework preserves the ability of each client to adapt rapidly while ensuring that the shared foundation remains broadly capable. The approach connects to a growing literature on federated meta-learning, including model-agnostic meta-learning formulations with theoretical guarantees and group-based personalization methods, but the integration with adaptive decoupled distillation guided by EMD is what distinguishes the new framework from its predecessors.</p>
<p>The empirical evidence presented in the paper is substantial. The authors evaluated FedML-DKD on three widely used image classification benchmarks: CIFAR-10, CIFAR-100, and EMNIST-Letters. These datasets were partitioned across simulated clients under different levels of data heterogeneity, replicating the skewed, non-IID conditions that plague real deployments. Across these experiments, FedML-DKD consistently outperformed six representative baseline methods drawn from the federated learning and distillation literature. The reported gains are striking: accuracy improvements of up to 10.87 percentage points on the more challenging classification tasks, alongside stable convergence behavior, meaning the training did not exhibit the erratic oscillations that often accompany heterogeneous federated optimization.</p>
<p>The significance of a ten-percentage-point improvement in this domain should not be understated. Benchmarks like CIFAR-100, with one hundred fine-grained classes, are notoriously sensitive to distributional skew, and many published federated methods report gains of only one or two points over baselines. A framework that maintains both accuracy and convergence stability under severe heterogeneity could matter for applications where data cannot be centralized for legal or practical reasons, such as healthcare systems that must protect patient privacy, intrusion detection in resource-constrained networks, and mobile-edge computing scenarios where devices learn collaboratively over limited bandwidth. The paper&#8217;s reference list itself maps this landscape, citing recent work on privacy-preserving federated learning for smart healthcare, green federated learning, and data-free distillation techniques that avoid sharing even synthetic data.</p>
<p>Like any study, the work has boundaries worth noting. The experiments were conducted on standard computer vision benchmarks rather than production-scale deployments, and the authors state that no new datasets were generated or analyzed during the study. The computational cost of computing Earth Mover&#8217;s Distance between distributions at each round, while modest compared with model training itself, is a factor that future deployments will need to characterize. Nevertheless, the conceptual contribution is clear and elegant: by measuring how far each client&#8217;s data distribution sits from the collective norm and using that measurement to recalibrate what kind of knowledge gets distilled, FedML-DKD turns heterogeneity from a silent saboteur into an explicit signal that the training algorithm can respond to. As federated learning moves from research prototypes toward the infrastructure of privacy-conscious AI, techniques that make collaborative learning robust to the messy reality of uneven data will only grow in importance, and this study offers a concrete, empirically validated step in that direction.</p>
<p><strong>Subject of Research:</strong> Adaptive decoupled knowledge distillation and meta-learning for federated learning under data heterogeneity</p>
<p><strong>Article Title:</strong> EMD-guided adaptive decoupled knowledge distillation for federated meta-learning under data heterogeneity</p>
<p><strong>Article References:</strong> Wang, D., &amp; Guan, S. (2026). EMD-guided adaptive decoupled knowledge distillation for federated meta-learning under data heterogeneity. <em>Cluster Computing, 29</em>(14), Article 786. <a href="https://doi.org/10.1007/s10586-026-06617-5" rel="noopener noreferrer">https://doi.org/10.1007/s10586-026-06617-5</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10586-026-06617-5" rel="noopener noreferrer">10.1007/s10586-026-06617-5</a></p>
<p><strong>Keywords:</strong> federated learning, knowledge distillation, meta-learning, data heterogeneity, Earth Mover&#x27;s Distance, client drift, non-IID data, CIFAR-100, decoupled knowledge distillation, distributed machine learning, privacy-preserving AI, model aggregation</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">236988</post-id>	</item>
	</channel>
</rss>
