<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>temporal stability &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/temporal-stability/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 11:49:04 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>temporal stability &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>When Machine Learning Labels Lie: Auditing Crypto Risk Models Exposes Hidden Weaknesses</title>
		<link>https://scienmag.com/when-machine-learning-labels-lie-auditing-crypto-risk-models-exposes-hidden-weaknesses/</link>
		
		<dc:creator><![CDATA[Teresa Odom]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 11:49:04 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[auditing AI in digital asset markets]]></category>
		<category><![CDATA[auditing weakly supervised AI models]]></category>
		<category><![CDATA[Bitcoin]]></category>
		<category><![CDATA[challenges in cryptocurrency risk labeling]]></category>
		<category><![CDATA[cryptocurrency]]></category>
		<category><![CDATA[cryptocurrency market risk modeling]]></category>
		<category><![CDATA[financial machine learning reliability]]></category>
		<category><![CDATA[financial time series]]></category>
		<category><![CDATA[ground truth in crypto risk assessment]]></category>
		<category><![CDATA[label noise]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning label accuracy in finance]]></category>
		<category><![CDATA[model auditing]]></category>
		<category><![CDATA[proxy labels]]></category>
		<category><![CDATA[proxy labels in crypto risk prediction]]></category>
		<category><![CDATA[risk-state learning]]></category>
		<category><![CDATA[semantic regularization]]></category>
		<category><![CDATA[temporal stability]]></category>
		<category><![CDATA[TOPSIS]]></category>
		<category><![CDATA[trustworthy crypto market risk analysis]]></category>
		<category><![CDATA[uncovering biases in crypto risk models]]></category>
		<category><![CDATA[weak supervision]]></category>
		<category><![CDATA[weak supervision in financial AI]]></category>
		<category><![CDATA[weaknesses of market risk models]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=222454</guid>

					<description><![CDATA[A new audit protocol reveals that machine learning models trained on cheap proxy labels for cryptocurrency risk can excel at fitting their training labels while failing to capture real market danger.]]></description>
										<content:encoded><![CDATA[<p>Machine learning models that promise to detect danger in cryptocurrency markets may be learning something quite different from what their creators intended. A new study published in Applied Intelligence by Junwen Lou, Fukai Zhang, and Shuo Tian of Henan Polytechnic University in China does not offer yet another forecasting architecture for digital assets. Instead, it delivers something arguably more valuable: a rigorous audit protocol designed to expose what weakly supervised models actually learn when they are trained on cheaply constructed proxy labels rather than verified ground truth. The findings carry an uncomfortable message for the booming field of financial machine learning, because the model that best fits its training labels is not necessarily the model that produces the most trustworthy picture of market risk.</p>
<p>The core problem the researchers tackle is one that haunts much of modern artificial intelligence: labeled data is expensive, and in financial markets it is often simply unavailable. Nobody can observe a market&#8217;s true risk state directly the way a radiologist can confirm a tumor on a scan. The standard workaround is weak supervision, a family of techniques in which heuristic rules generate approximate labels from raw data. In cryptocurrency research, this typically means defining market states, such as calm, stressed, or crisis regimes, using rules built from observable quantities like volatility, drawdowns, or return thresholds. These proxy labels are easy to construct, but as the new study demonstrates, what they transfer into a learned model&#8217;s internal representation of the market remains disturbingly opaque.</p>
<p>To pierce that opacity, the team built an audit protocol that evaluates candidate state learners along three deliberately separated axes. The first axis is proxy-label fit, a measure of how well a trained model&#8217;s inferred states agree with the heuristic labels it was trained on. The second is temporal stability, which asks whether the model&#8217;s state path through time behaves sensibly, avoiding pathological flickering between states that would render the output useless for decision-making. The third, and arguably most important, axis is ex-post risk semantics: whether the states the model recovers actually stratify future risk in a meaningful way, for instance by identifying periods that precede large drawdowns. By keeping these three criteria apart rather than blending them into a single score, the protocol makes disagreements between them visible instead of hiding them inside an aggregate number.</p>
<p>The empirical setting was deliberately grounded in real trading conditions. The researchers used Binance spot one-hour candlestick data covering five cryptocurrency assets, and tested lightweight baseline learners alongside more demanding stress tests built on temporal backbones applied to BTCUSDT. The headline result is striking: across the tested assets and configurations, the state paths recovered by the models aligned more consistently with future risk stratification than with the direction of future returns. In other words, these weakly supervised learners appear to be capturing something about impending danger rather than impending profit, a distinction that matters enormously for anyone hoping to use such models for risk management rather than speculation. The authors are careful, however, to flag that this conclusion is conditional on the specific weak-label design used in their experiments, a caveat that turns out to be one of the paper&#8217;s most important contributions.</p>
<p>That conditionality centers on a single design choice with outsized consequences. The proxy labels in the study were anchored to a metric the authors call max_drawdown_72, a measure of maximum drawdown computed over a 72-hour window, and this downside-risk information was embedded both in the feature set fed to the models and in the label rule itself. The audit revealed that the risk layering recovered by the models was materially shaped by this anchor. This is a textbook example of label-design dependence: the model appears to discover risk structure in the market, but part of that structure was baked in by the very definition of the labels. An unwary practitioner could easily mistake this circularity for genuine predictive insight, which is precisely the kind of failure the audit protocol is designed to catch.</p>
<p>Perhaps the most practically consequential finding concerns model selection. In multi-criteria decision settings, practitioners commonly reach for established aggregation tools such as weighted scoring or the Technique for Order of Preference by Similarity to Ideal Solution, known as TOPSIS, to pick the best candidate from a pool of models. The study shows that these conventional selectors do not reliably detect the gap between proxy-label fit and semantic credibility once stronger label-fitting candidates enter the pool. A model can dominate on the metric that measures agreement with its own training labels while producing a state path that is less faithful to real future risk than a competitor with worse label fit. The implication is sobering: the standard machinery of model selection can systematically favor models that are best at memorizing their labels rather than best at understanding the market.</p>
<p>The researchers also tested whether the problem could be patched with more sophisticated machinery, and the answers were largely negative. Adding semantic regularization, a technique intended to nudge learned representations toward desired semantic properties, provided no measurable gain under the tested settings. Similarly, lightweight Transformer-based and PatchTST-style reference models, architectures drawn from the current generation of time-series deep learning, did not overturn the primary trade-off between label fit and semantic validity. This matters because a common reflex in applied machine learning is to assume that newer or more expressive architectures will resolve fundamental data problems. Here, the evidence suggests the bottleneck lies not in the model class but in the quality and design of the weak supervision itself, a conclusion consistent with the broader literature on learning with noisy labels.</p>
<p>One of the paper&#8217;s most original contributions is a diagnostic called the Semantic Orientation Ratio, which caught a failure before any model was even trained. On the XRP asset, the proxy labels were already misaligned with the target risk semantics at the label-audit stage, meaning the heuristic labeling rule itself was producing labels that contradicted the risk concept they were supposed to encode. The Semantic Orientation Ratio detected this misalignment early, flagging the problem before researchers wasted effort training models on corrupted supervision. This boundary case illustrates the value of auditing the labels rather than only the models: when the supervision signal is broken, no amount of architectural sophistication downstream can repair it, and early detection saves both computational resources and misleading conclusions.</p>
<p>The broader significance of this work extends well beyond cryptocurrency markets. Weak supervision is now a cornerstone of applied machine learning across domains where ground truth is scarce, from programmatic data labeling systems to medical and scientific applications, and the gap between what proxy labels reward and what practitioners actually need is a universal hazard. By formalizing an audit view in which label fidelity, path stability, semantic validity, and label-design dependence can all disagree, the study offers a template for interrogating any weakly supervised system before trusting its outputs. The authors have also committed to openness, releasing core preprocessing, model-training, and evaluation code publicly, with raw Binance data available from the Binance Public Data portal, enabling independent replication of their audit. In a field crowded with claims of predictive prowess, a protocol that asks not how well a model fits its labels but whether those labels mean anything at all may prove to be the most important contribution of all.</p>
<p><strong>Subject of Research:</strong> Weakly supervised machine learning for cryptocurrency market risk-state detection and model auditing</p>
<p><strong>Article Title:</strong> Auditing weakly supervised cryptocurrency risk-state learning: when label fit, stability, and semantics disagree</p>
<p><strong>Article References:</strong> Lou, J., Zhang, F., &amp; Tian, S. (2026). Auditing weakly supervised cryptocurrency risk-state learning: when label fit, stability, and semantics disagree. <em>Applied Intelligence, 56</em>(15), Article 467. <a href="https://doi.org/10.1007/s10489-026-07467-9" rel="noopener noreferrer">https://doi.org/10.1007/s10489-026-07467-9</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10489-026-07467-9" rel="noopener noreferrer">10.1007/s10489-026-07467-9</a></p>
<p><strong>Keywords:</strong> cryptocurrency, weak supervision, machine learning, risk-state learning, model auditing, proxy labels, Bitcoin, temporal stability, label noise, TOPSIS, financial time series, semantic regularization</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">222454</post-id>	</item>
		<item>
		<title>Four Global Change Drivers Reshape Grassland Stability in Surprising Ways</title>
		<link>https://scienmag.com/four-global-change-drivers-reshape-grassland-stability-in-surprising-ways/</link>
		
		<dc:creator><![CDATA[Gavin Prescott]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:25:55 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[biodiversity]]></category>
		<category><![CDATA[ecosystem services under climate change]]></category>
		<category><![CDATA[ecosystem stability]]></category>
		<category><![CDATA[effects of elevated CO2 on grasslands]]></category>
		<category><![CDATA[elevated CO2]]></category>
		<category><![CDATA[factorial field experiments in ecology]]></category>
		<category><![CDATA[functional traits]]></category>
		<category><![CDATA[global change drivers impact on grasslands]]></category>
		<category><![CDATA[global change factors]]></category>
		<category><![CDATA[grassland ecosystem stability]]></category>
		<category><![CDATA[grassland productivity]]></category>
		<category><![CDATA[long-term field experiment]]></category>
		<category><![CDATA[long-term grassland productivity studies]]></category>
		<category><![CDATA[multi-factor global change experiments]]></category>
		<category><![CDATA[nitrogen enrichment]]></category>
		<category><![CDATA[nitrogen enrichment and grassland productivity]]></category>
		<category><![CDATA[non-additive interactions in ecological stability]]></category>
		<category><![CDATA[reduced rainfall]]></category>
		<category><![CDATA[reduced rainfall and ecosystem resilience]]></category>
		<category><![CDATA[species asynchrony]]></category>
		<category><![CDATA[temporal stability]]></category>
		<category><![CDATA[temporal stability of grassland ecosystems]]></category>
		<category><![CDATA[warming]]></category>
		<category><![CDATA[warming effects on grassland stability]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=203884</guid>

					<description><![CDATA[A 13-year fully factorial experiment manipulating carbon dioxide, nitrogen, warming, and reduced rainfall shows that grassland stability is governed by shifting, non-additive driver interactions mediated by species asynchrony and trait composition.]]></description>
										<content:encoded><![CDATA[<p>In one of the longest and most ambitious experiments of its kind, researchers have shown that the stability of grassland ecosystems under human-driven environmental change cannot be predicted by studying one stressor at a time. A 13-year fully factorial field experiment, described in Nature Ecology &amp; Evolution, manipulated four major global change factors simultaneously—elevated atmospheric carbon dioxide, nitrogen enrichment, warming, and reduced rainfall—and tracked how they alone and in combination shaped the productivity of planted grassland communities and the consistency of that productivity through time. The results reveal a world of shifting, non-additive interactions that would remain invisible in the short-term, single-driver studies that have long dominated the field.</p>
<p>The central measure of interest was temporal stability, defined as the mean of aboveground net primary productivity divided by its temporal standard deviation. A stable ecosystem is one whose year-to-year output varies little relative to its average performance, and stability matters because it underpins forage supply, carbon storage, and the livelihoods that depend on productive landscapes. By quantifying both the mean and the variability of productivity across more than a decade, the team could disentangle whether a given driver changed stability by lifting average output, by amplifying or damping the swings between good years and bad ones, or by some interaction of both.</p>
<p>When each factor acted on its own, with all others held at ambient levels, the single-driver results were themselves instructive. Elevated carbon dioxide reduced stability, and nitrogen enrichment did so less markedly, because in both cases the temporal standard deviation of productivity increased more than the mean did. In other words, carbon dioxide and nitrogen made grasslands behave more erratically even when average productivity did not rise proportionally. Warming and reduced rainfall, by contrast, each increased stability when acting alone, but through opposite mechanisms: reduced rainfall suppressed the temporal standard deviation, dampening the fluctuations that destabilize communities, while warming enhanced mean productivity disproportionately, raising the denominator&#8217;s benefit relative to variability.</p>
<p>The deeper story, however, emerged from the combinations. Because the experiment was fully factorial, every permutation of the four drivers was replicated across independent experimental plots, allowing the team to compare observed combined effects against the additive expectations built from single-driver responses. Combined driver effects frequently shifted in magnitude over the thirteen years and, in some cases, reversed their expected additive direction entirely. Interactions were classified as synergistic when the combined effect exceeded the additive expectation and antagonistic when it fell below it, and both categories appeared across the treatment matrix. A driver that destabilized in the early years might stabilize later, or a destabilizing pair might be rescued by a third factor, with the balance of effects drifting as the plant community itself reorganized.</p>
<p>This time dependence carries a blunt message for the field: short-term experiments, typically two to five years long, can return conclusions that are directionally wrong for the long run. The authors show that rolling five-year windows within the same continuous experiment produce different interaction classifications depending on when the window is placed. Since most existing evidence about multi-driver effects comes from exactly such short windows, the meta-analyses and models built upon them may systematically misrepresent how real ecosystems will respond as carbon dioxide, nitrogen deposition, temperatures, and drought regimes continue to change together over decades.</p>
<p>Why do these interactions keep shifting? The study points to mechanisms operating through the structure and composition of the plant community itself. Across treatments, stability was governed primarily by species asynchrony—the degree to which different species fluctuate out of phase with one another, so that declines in one species are buffered by increases in another. Asynchrony is a classical insurance mechanism of biodiversity, but the experiment demonstrates that global change drivers remodel it continuously. As the relative abundances of planted species changed through time under the different treatment combinations, the degree of temporal compensation among them changed too, dragging stability up or down in ways no single year could capture.</p>
<p>Secondary contributions came from soil moisture and from functional composition, the suite of community-weighted average plant traits that describe how the community acquires resources and withstands stress. The team compiled eleven community-weighted mean traits spanning resource acquisition and stress resistance gradients, including specific leaf area, leaf nitrogen and phosphorus content per unit mass, leaf water content, leaf carbon content, vegetative spread rate, seed mass, plant height, root depth, and leaf dry matter content. Changes in this trait coordination—the coordinated shifts in which resource-use strategies dominate the community—formed a mechanistic bridge between the physical drivers and the demographic insurance captured by asynchrony. A trait-based principal component analysis separated axes of moisture usability and acquisitive versus conservative strategies, linking the drivers directly to the functional identity of the vegetation.</p>
<p>The individual species trajectories underline how dynamic the communities were. Each plot had been planted in 1997 with nine species drawn randomly from a pool of sixteen native grassland species, and by 2012 all four drivers were fully imposed. Species-specific cover records from 2012 to 2024 show divergent linear trends, with some lineages expanding under particular treatment combinations and others contracting toward local rarity or disappearance. Plot-level species variability declined with species richness, consistent with the averaging effect by which richer communities dilute the influence of any one fluctuating population. This compositional turnover is precisely the material through which the drivers acted: there is no fixed community responding mechanically to stress, only a continuously reshuffled assemblage whose functional and temporal properties evolve year by year.</p>
<p>Statistically, the team used piecewise structural equation modeling to trace pathways from the drivers and their interactions, through soil moisture, species asynchrony, and functional composition, to mean productivity, temporal variability, and ultimately stability. The path diagrams show that driver effects on stability are largely indirect, funneled through these intermediate variables rather than acting on stability alone. The framework explains why additive expectations fail: each driver modifies soil moisture, shifts trait composition, and alters asynchrony, and because those mediators are shared, the drivers inevitably interfere with one another in nonlinear ways. The approach also explains the counterintuitive single-driver results—for instance, warming can stabilize productivity by raising the mean enough to outweigh added variance, while carbon dioxide destabilizes by inflating variance faster than yield.</p>
<p>The implications extend well beyond the experimental plots. Grasslands cover vast areas of the terrestrial surface and supply a disproportionate share of the world&#8217;s forage and grazing capacity, so their temporal reliability is an economic and food-security variable, not merely an ecological abstraction. Models of terrestrial carbon cycling and land-surface feedbacks routinely scale up results from short-term, single-factor experiments; this study suggests such scaling inherits both a missing-interaction bias and a missing-time bias. The authors argue that predicting ecosystem behavior in the coming decades requires experiments that manipulate multiple drivers simultaneously over long enough horizons to capture the dynamic reorganization of species asynchrony and trait composition. Thirteen years was long enough for interactions to change magnitude and direction; the real world will not offer a shorter timeline. As global change factors continue to arrive together, the stability of the systems that feed and clothe humanity will be decided not by any single stressor, but by the shifting, non-additive choreography among them.</p>
<p><strong>Subject of Research:</strong> Long-term multifactor global change experiment on grassland productivity stability</p>
<p><strong>Article Title:</strong> Ecosystem stability is shaped by resource–trait coordination under multiple interacting global change factors</p>
<p><strong>Article References:</strong> Ding, X., Chen, H. Y. H., Isbell, F., &amp; Reich, P. B. (2026). Ecosystem stability is shaped by resource–trait coordination under multiple interacting global change factors. <em>Nature Ecology &amp;amp; Evolution</em>. <a href="https://doi.org/10.1038/s41559-026-03190-3" rel="noopener noreferrer">https://doi.org/10.1038/s41559-026-03190-3</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41559-026-03190-3" rel="noopener noreferrer">10.1038/s41559-026-03190-3</a></p>
<p><strong>Keywords:</strong> ecosystem stability, global change factors, grassland productivity, elevated CO2, nitrogen enrichment, warming, reduced rainfall, species asynchrony, functional traits, biodiversity, temporal stability, long-term field experiment</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">203884</post-id>	</item>
	</channel>
</rss>
