<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Incomplete annotations in audio datasets &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/incomplete-annotations-in-audio-datasets/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 08 Oct 2026 16:08:13 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>Incomplete annotations in audio datasets &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Ontology Trick Boosts Weak Audio Labels by Up to 30 Percent Without Training</title>
		<link>https://scienmag.com/ontology-trick-boosts-weak-audio-labels-by-up-to-30-percent-without-training/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 08 Oct 2026 16:08:13 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[Audio dataset annotation]]></category>
		<category><![CDATA[audio event classification]]></category>
		<category><![CDATA[audio tagging]]></category>
		<category><![CDATA[AudioSet]]></category>
		<category><![CDATA[AudioSet label sparsity solutions]]></category>
		<category><![CDATA[Boosting audio classification without additional training]]></category>
		<category><![CDATA[CPU-only computing]]></category>
		<category><![CDATA[data-centric AI]]></category>
		<category><![CDATA[Hierarchical label propagation in audio analysis]]></category>
		<category><![CDATA[hierarchical learning]]></category>
		<category><![CDATA[Improving sound event detection accuracy]]></category>
		<category><![CDATA[Incomplete annotations in audio datasets]]></category>
		<category><![CDATA[label density]]></category>
		<category><![CDATA[label propagation]]></category>
		<category><![CDATA[Large-scale audio dataset labeling challenges]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Machine learning for audio label inference]]></category>
		<category><![CDATA[ontology]]></category>
		<category><![CDATA[Ontology-based sound label correction]]></category>
		<category><![CDATA[semantic hierarchy]]></category>
		<category><![CDATA[Semantic sound category hierarchy]]></category>
		<category><![CDATA[Sound event ontology structure]]></category>
		<category><![CDATA[Weak audio labels enhancement]]></category>
		<category><![CDATA[weak supervision]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=248557</guid>

					<description><![CDATA[Researchers at the University of Tehran have shown that a training-free, ontology-guided label propagation method can enrich weak AudioSet annotations by up to 30.7 percent and improve downstream audio classification while running entirely on CPUs.]]></description>
										<content:encoded><![CDATA[<p>Every large audio dataset carries a hidden flaw. AudioSet, Google&#8217;s sprawling collection of more than two million labeled sound clips, is annotated at the clip level, meaning a human listener tagged a ten-second recording with a handful of words such as guitar, dog, or traffic noise. Those tags are weak, sparse, and frequently incomplete. A clip containing a barking dog and passing cars may only carry one of those labels, and a machine learning model trained on such data inherits every omission. Researchers at the University of Tehran have now demonstrated a strikingly simple fix: instead of training bigger models or hiring more annotators, they let the structure of the sound ontology itself fill in the gaps, and they do it on an ordinary computer processor without touching a single neural network weight.</p>
<p>The study, published in Discover Artificial Intelligence by Alireza Bashooki and Hedieh Sajedi, builds on an idea called Hierarchical Label Propagation. AudioSet is not just a flat list of categories; it is a semantic tree, an ontology in which fine-grained events nest inside broader ones. A guitar pluck is a child of plucked string instrument, which in turn is a child of musical instrument. If a clip is labeled guitar, then logically it also contains a musical instrument. Propagating labels upward through this hierarchy therefore enriches annotations for free, using nothing more than the parent-child relations already encoded by the dataset&#8217;s designers. Earlier work, notably a 2025 study by Tuncay and colleagues, showed that this kind of propagation can improve audio tagging models, but it required substantial computational resources and was evaluated across the full 527-category ontology.</p>
<p>The Tehran team&#8217;s contribution is a lightweight, training-free refinement of this idea that they call Multi-hop Soft Hierarchical Label Propagation. The conventional baseline, which they term Hard-HLP, deterministically copies every active label to its immediate parent. If a clip contains drum, it gains a positive label for any parent categories that fall within the working label space. This single-hop strategy is effective but shallow: it cannot reach grandparents or more distant ancestors, and it treats every propagated label with equal certainty. The new method extends the traversal to three hops using a breadth-first search over the ontology graph, assigning each ancestor a fixed weight that decays with hierarchical distance: 1.0 for a direct parent, 0.7 for a grandparent, and 0.5 for a great-grandparent.</p>
<p>Two design choices make the soft variant robust rather than merely aggressive. First, when an ancestor can be reached along multiple paths, the algorithm keeps only the maximum propagated weight instead of summing contributions, which prevents score inflation when many descendants share the same broad ancestor. Second, the accumulated scores are binarized with a threshold parameter tau, giving practitioners a dial that controls how much semantic expansion they are willing to accept. The entire procedure involves only graph traversal and matrix-based label updates, with no trainable parameters, no iterative optimization, and no matrix inversions. The computational cost grows roughly linearly with the number of active labels and their reachable ancestors, which is why the whole pipeline runs comfortably on CPU-only hardware.</p>
<p>The experiments focused on the thirty most frequent classes in AudioSet&#8217;s balanced training split. After filtering, the working subset contained 20,847 audio clips, each represented as a multi-hot vector over the thirty selected classes, with an average of exactly one positive label per clip. The AudioSet ontology file used in the implementation contained 632 nodes and 670 child-to-parent relations, and 28 of the thirty selected classes had parent links. Crucially, only ancestors that themselves belong to the top-30 output space were retained as final labels; more distant ontology nodes were traversed but discarded, keeping the label space fixed so that the effect of propagation could be measured cleanly.</p>
<p>The headline numbers are compelling. Hard-HLP raised the average label density from 1.000 to 1.173 labels per clip, a relative gain of 17.3 percent. Multi-hop Soft-HLP pushed the density to 1.307 at a threshold of 0.3, an improvement of 30.7 percent over the baseline, and to 1.281 at a stricter threshold of 0.7, still a 28.1 percent gain. Importantly, the method proved stable across moderate threshold values, meaning practitioners do not need to fine-tune the parameter obsessively. At very high thresholds such as 0.9, the density falls back toward the Hard-HLP level, showing that aggressive filtering suppresses the multi-hop enrichment entirely. The sweet spot, according to the authors, lies between 0.3 and 0.5.</p>
<p>Denser labels are only valuable if they are correct, and the researchers were careful to separate completeness from reliability. They manually verified the newly propagated labels on a random subset of 200 clips, scoring each as correct, incorrect, or uncertain based on whether the sound event was actually audible. Precision climbed from 84.6 percent at the loosest threshold to 97.1 percent at the strictest, while recall fell from 76.8 percent to 44.5 percent, tracing the expected trade-off curve. The F1-score peaked at 80.7 percent with tau set to 0.3, which is why that value became the default operating point. Per-class analysis revealed that acoustically distinctive categories such as drum, guitar, dog, and train propagated more reliably than broad categories like music, speech, and vehicle, which are prone to over-generalization. Most of the added labels were high-level ancestors such as musical instrument, vehicle, and animal, confirming that the method activates meaningful concepts rather than injecting noise.</p>
<p>To test whether the enriched labels genuinely help learning, the team trained a small multilayer perceptron with a single hidden layer of 128 units on fixed pre-computed features, holding the architecture and training protocol identical across three supervision settings: baseline labels, Hard-HLP labels, and Multi-hop Soft-HLP labels at tau equal to 0.3. The classifier was deliberately kept outside the proposed method; it served purely as an independent yardstick. The result was decisive. The model trained on multi-hop enriched labels outperformed both alternatives across all metrics, posting absolute gains of 6.5 percent in F1-score and 6.6 percent in mean average precision over the baseline. Because the only variable was the label set, the improvement can be attributed to the quality of the supervision rather than any change in the model itself.</p>
<p>The work sits within a broader data-centric movement in machine learning that treats supervision quality, not just model architecture, as the key lever for performance. Recent efforts have attacked AudioSet&#8217;s annotation problems from other directions: Grzywalski and Botteldooren documented systematic inconsistencies in how annotators applied labels at different ontology levels, Dinkel and colleagues generated pseudo strong labels from weak ones, and Sun and colleagues built AudioSet-R using multi-stage reannotation with large language models. Those approaches are powerful but computationally demanding. The Tehran framework is complementary: it performs explicit, interpretable hierarchical reasoning at the label level, costs almost nothing to run, and can be applied as a preprocessing step before any downstream model is trained.</p>
<p>The authors acknowledge limitations. The evaluation covered a subset of AudioSet under a single experimental protocol, the propagation weights were fixed rather than learned, and the method relies exclusively on parent-child relations without adaptive or confidence-aware edge weighting. Future work will explore larger benchmarks, cross-modal and multilingual annotation scenarios, adaptive weighting schemes, and formal statistical testing across multiple random seeds, along with direct comparisons against graph neural network approaches. Even so, the message is clear and likely to resonate far beyond audio research: sometimes the cheapest way to improve a dataset is to read the taxonomy it already ships with. A three-hop walk through a semantic tree, executed on a laptop processor, delivered nearly a third more supervision than the original annotations, and measurably better classifiers as a result.</p>
<p><strong>Subject of Research:</strong> Ontology-guided multi-hop hierarchical label propagation for enriching weakly labeled audio datasets</p>
<p><strong>Article Title:</strong> Ontology guided multi hop label propagation for AudioSet refinement</p>
<p><strong>Article References:</strong> Bashooki, A., &amp; Sajedi, H. (2026). Ontology guided multi hop label propagation for AudioSet refinement. <em>Discover Artificial Intelligence, 6</em>(1), Article 1402. <a href="https://doi.org/10.1007/s44163-026-02365-y" rel="noopener noreferrer">https://doi.org/10.1007/s44163-026-02365-y</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44163-026-02365-y" rel="noopener noreferrer">10.1007/s44163-026-02365-y</a></p>
<p><strong>Keywords:</strong> AudioSet, label propagation, ontology, weak supervision, audio event classification, hierarchical learning, machine learning, data-centric AI, label density, CPU-only computing, semantic hierarchy, audio tagging</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">248557</post-id>	</item>
	</channel>
</rss>
