<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>time series &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/time-series/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 01 Oct 2026 23:30:34 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>time series &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Algorithm Repairs Timestamp Errors in Correlated Sensor Networks</title>
		<link>https://scienmag.com/new-algorithm-repairs-timestamp-errors-in-correlated-sensor-networks/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 23:30:34 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced algorithms for sensor data quality control]]></category>
		<category><![CDATA[air quality]]></category>
		<category><![CDATA[automated timestamp error detection in air quality monitoring]]></category>
		<category><![CDATA[big data]]></category>
		<category><![CDATA[big data solutions for sensor timestamp errors]]></category>
		<category><![CDATA[data repair]]></category>
		<category><![CDATA[Environmental Monitoring]]></category>
		<category><![CDATA[environmental sensor network data integrity]]></category>
		<category><![CDATA[handling displaced sensor measurements]]></category>
		<category><![CDATA[impact of timestamp errors on pollution pattern analysis]]></category>
		<category><![CDATA[improving accuracy of environmental data logging]]></category>
		<category><![CDATA[improving data accuracy in industrial and traffic monitoring]]></category>
		<category><![CDATA[mean absolute scaled error]]></category>
		<category><![CDATA[Mpumalanga]]></category>
		<category><![CDATA[multiple imputation]]></category>
		<category><![CDATA[public health data reliability in sensor networks]]></category>
		<category><![CDATA[real-time correction of sensor timestamp discrepancies]]></category>
		<category><![CDATA[sensor data]]></category>
		<category><![CDATA[sensor timestamp correction]]></category>
		<category><![CDATA[simulation study]]></category>
		<category><![CDATA[spatial correlation]]></category>
		<category><![CDATA[spatial relationship-based data repair algorithms]]></category>
		<category><![CDATA[time series]]></category>
		<category><![CDATA[timestamp error]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=224310</guid>

					<description><![CDATA[Researchers at the University of KwaZulu-Natal have developed an algorithm that detects and repairs timestamp-displaced sensor data by exploiting spatial correlation across air monitoring networks.]]></description>
										<content:encoded><![CDATA[<p>Every second counts in environmental monitoring, yet the clocks behind the sensors that watch our air are far less reliable than we might like to believe. When an automatic monitoring station records a measurement, it stamps that reading with a time. If that timestamp is wrong, the data itself is not lost but displaced, sliding along the timeline so that a temperature recorded at noon may appear in the record as if it belonged to midnight. A new study published in the Journal of Big Data by Natalie D. Benschop, Temesgen Zewotir and Rajen N. Naidoo of the University of KwaZulu-Natal tackles this deceptively simple problem with a deceptively elegant solution: an algorithm that detects and repairs displaced sequences of sensor data by exploiting the spatial relationships between monitoring stations scattered across a network.</p>
<p>The stakes are higher than they might first appear. Air quality networks feed decisions about public health warnings, traffic management and industrial regulation. When a station&#8217;s data arrives shifted by several hours, the daily rhythm of pollution, the diurnal pattern that peaks during rush hours and dips overnight, becomes scrambled in ways that standard quality-control checks may not catch. The South African authors focus on exactly this scenario, where timestamp error manifests as an irregular diurnal pattern rather than a uniform offset. Their work addresses a gap they identify directly: literature on correcting timestamp error in big sensor data streams remains scarce, particularly in the domain of environmental monitoring, even though the impact of such errors can be profound.</p>
<p>The core insight behind the new method, which the authors call SIMMI, short for Shift Identification Mechanism by Multiple Imputation, rests on a simple premise. In a spatial network of monitoring stations, nearby sensors measure correlated phenomena. Air temperature at one station tracks the temperature at its neighbours with a lag that should be close to zero if all clocks agree. If one station&#8217;s clock is wrong, its data will align best with its neighbours&#8217; data only after being shifted by some number of time steps. The algorithm therefore takes one selected variable from the multivariate set, gradually shifts the displaced measurements across a range of candidate lags, and compares each shifted version against multiple imputed representations of what the data should look like at corresponding time points.</p>
<p>Those imputed representations are where the spatial correlation does its work. Using multivariate imputation via chained equations, a standard technique for filling in missing values, the method builds several plausible reconstructions of the suspect series based on the surrounding stations. Because these reconstructions are anchored to correctly timed neighbours, they serve as a reference clock. The algorithm then evaluates each candidate shift using a pooled measure of the mean absolute scaled error, a metric that normalises forecast accuracy in a way that allows fair comparison across series with different variability. The shift that minimises this pooled error identifies how far the displaced sequence must be moved to restore it to its true position on the timeline.</p>
<p>To find out whether this approach actually works, the researchers stress-tested it in a univariate simulation study against a battery of baseline methods. The competitors included strategies familiar to anyone who has wrestled with broken time series: last observation carried forward, autoregressive integrated moving average models, and their extension with exogenous variables, along with several maximisation-based alternatives that relied on global optimisation, average correlation, or modal lag identification. The results were unambiguous. The new algorithm proved superior and highly reliable in the exact correction of contrived sequences of displaced air temperature data, and its advantage was greatest under conditions of strong cross-correlation between series recorded at different locations, precisely the situation that prevails in a dense network of air monitoring stations.</p>
<p>Strong cross-correlation is the algorithm&#8217;s fuel. When neighbouring stations record similar values at similar times, the imputed reference series become faithful mirrors of the true signal, and the correct shift stands out sharply against the alternatives. When correlation weakens, for instance between stations separated by large distances or different microclimates, the mirror blurs and the task becomes harder. This dependence on spatial structure is not a flaw so much as a design principle: the method is built for exactly the kind of spatially correlated, multivariate sensor networks that dominate modern environmental monitoring, from air quality grids to weather station arrays. The authors suggest the approach may generalise to other domains that share the same simple premise of correlated series recorded in parallel.</p>
<p>Simulations alone rarely settle the matter, so the team applied the algorithm to an empirical case of timestamp error in real multivariate data drawn from air monitoring stations dispersed across Mpumalanga province in South Africa during 2018. The data, drawn from the South African Air Quality Information System and the South African Weather Service, included pollutants such as nitrogen dioxide, ozone and sulphur dioxide alongside meteorological variables like air temperature and ambient pressure, each with its own patterns of missingness. In this real-world setting, where missing values, irregular patterns and timestamp displacement coexist, the algorithm was asked to do the full job: identify the shift and repair the sequence without contaminating the result with artifacts of the imputation process.</p>
<p>The empirical application was accompanied by a sensitivity analysis, and here the results proved robust and consistent. The corrected output did not hinge on fragile choices of imputation details or error-metric thresholds; varying the conditions left the identified shifts stable. For practitioners, this matters as much as raw accuracy. A repair tool that works only under laboratory conditions is of limited use in operational data pipelines, where the messy realities of sensor downtime, communication failures and clock drift arrive together. The authors&#8217; demonstration that their method holds up under sensitivity testing in genuine multivariate environmental data is a meaningful step from proof of concept toward practical deployment.</p>
<p>Why does this matter beyond the air quality community? The volume of timestamped sensor data is exploding, driven by Internet of Things deployments, smart city infrastructure and dense environmental networks. Every one of those systems inherits the same vulnerability: a sensor&#8217;s clock is a single point of failure that silently corrupts data without deleting it. Displaced data is arguably more dangerous than missing data, because it looks complete and plausible while being wrong. Downstream models, from pollution forecasts to machine learning systems trained on historical records, absorb the error and propagate it. A correction algorithm that can realign displaced sequences automatically, using the network&#8217;s own redundancy as its reference, offers a way to harden these data pipelines at the source.</p>
<p>The study also carries a quiet methodological lesson. Rather than treating imputation and error correction as separate chores, SIMMI fuses them: the multiple imputations that would normally fill gaps become the yardstick against which temporal displacement is measured. That fusion turns a weakness of correlated networks, their reliance on neighbours, into a strength. As sensor networks grow denser and the cost of bad timestamps grows with them, approaches that let stations check each other&#8217;s clocks, quietly and continuously, may become as standard a part of data infrastructure as the sensors themselves. For now, the Mpumalanga air monitoring network has provided the proving ground, and the results suggest that the era of silently scrambled environmental data may be drawing to a close.</p>
<p><strong>Subject of Research:</strong> An algorithm for correcting timestamp errors in spatially correlated environmental sensor data</p>
<p><strong>Article Title:</strong> A new algorithm for correcting irregular patterns associated with timestamp error in spatial correlated sensor data</p>
<p><strong>Article References:</strong> Benschop, N. D., Zewotir, T., &amp; Naidoo, R. N. (2026). A new algorithm for correcting irregular patterns associated with timestamp error in spatial correlated sensor data. <em>Journal of Big Data</em>. <a href="https://doi.org/10.1186/s40537-026-01571-w" rel="noopener noreferrer">https://doi.org/10.1186/s40537-026-01571-w</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s40537-026-01571-w" rel="noopener noreferrer">10.1186/s40537-026-01571-w</a></p>
<p><strong>Keywords:</strong> timestamp error, sensor data, spatial correlation, environmental monitoring, air quality, multiple imputation, mean absolute scaled error, time series, big data, data repair, Mpumalanga, simulation study</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">224310</post-id>	</item>
		<item>
		<title>AI Predicts Daily Strawberry Harvests on a Real Commercial Farm</title>
		<link>https://scienmag.com/ai-predicts-daily-strawberry-harvests-on-a-real-commercial-farm/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 10:06:38 +0000</pubDate>
				<category><![CDATA[Agriculture]]></category>
		<category><![CDATA[AI model evaluation in agriculture]]></category>
		<category><![CDATA[AI-based yield forecasting accuracy]]></category>
		<category><![CDATA[artificial intelligence in agriculture]]></category>
		<category><![CDATA[British strawberry farm data]]></category>
		<category><![CDATA[commercial farm crop forecasting]]></category>
		<category><![CDATA[commercial horticulture]]></category>
		<category><![CDATA[daily yield estimation]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[feature engineering]]></category>
		<category><![CDATA[forward-year validation]]></category>
		<category><![CDATA[global strawberry market trends]]></category>
		<category><![CDATA[impact of forecasting errors on harvest scheduling]]></category>
		<category><![CDATA[IoT sensors]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[precision agriculture]]></category>
		<category><![CDATA[precision farming technology]]></category>
		<category><![CDATA[short-season fruit crop logistics]]></category>
		<category><![CDATA[soft fruit crop management]]></category>
		<category><![CDATA[strawberry]]></category>
		<category><![CDATA[strawberry harvest prediction]]></category>
		<category><![CDATA[TabPFN]]></category>
		<category><![CDATA[time series]]></category>
		<category><![CDATA[UK farming]]></category>
		<category><![CDATA[yield forecasting]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=221894</guid>

					<description><![CDATA[A leakage-aware artificial intelligence framework trained on six seasons of commercial farm data forecasts daily strawberry yields with a pretrained tabular model outperforming conventional machine learning and deep neural networks.]]></description>
										<content:encoded><![CDATA[<p>Few crops punish forecasting errors quite like strawberries. The fruit is delicate, the season is short, and every punnet that ripens without a picker or a cold-chain slot waiting for it is money left rotting in the field. A new study published in Smart Agricultural Technology shows that daily strawberry yields on a working commercial farm can be predicted with useful accuracy using artificial intelligence, but only when the models are built and tested with unusual care. The research, led by Salman Ahmed and Nicholas Caldwell of the University of Suffolk, draws on six consecutive production years of real operational data from a British soft-fruit farm, and its central message is as much about how models are evaluated as it is about how they are built.</p>
<p>The commercial stakes are considerable. The global fresh strawberry market has been estimated at roughly 20.7 billion US dollars in 2024, with projections exceeding 30 billion dollars by the mid-2030s as year-round demand grows. In the United Kingdom alone, official statistics value the strawberry crop at 389 million pounds for 2024. At daily resolution, even modest forecasting errors ripple through the entire operation: harvest crews are scheduled against predicted volumes, packing lines and storage capacity are booked in advance, and supermarkets hold contracts that assume reliable supply. A model that misses a peak picking day by a wide margin creates labour inefficiencies on one hand and post-harvest waste on the other.</p>
<p>What makes strawberries so hard to forecast is that daily yield is not a string of independent numbers but a biologically structured time series. Production follows a recognisable lifecycle: output starts at zero before harvest begins, climbs as plants enter peak fruiting, and falls back to zero as the season ends. That rise-peak-decline shape is governed by cultivar, planting density, accumulated temperature and humidity, and the decisions of farm managers. The researchers argue that a realistic forecasting framework must satisfy three conditions at once. It must preserve the natural lifecycle, including genuine zero-yield periods at season boundaries. It must enforce strict temporal alignment so that only information available before harvest can be used as a predictor. And it must test models on unseen future seasons rather than interpolating within the data it has already seen.</p>
<p>That last condition turns out to be the study&#8217;s sharpest methodological point. In agricultural machine learning, it is dangerously easy to inflate performance through information leakage, for example by letting weather recorded after a harvest event act as a predictor, or by using random train-test splits on strongly seasonal data. The team deliberately excluded aggregate productivity indicators such as yield per hectare and percentage picked from the feature set, because those variables indirectly encode the answer. Environmental sensor readings from the farm&#8217;s SoilMoistureSense platform, covering temperature and relative humidity, were aggregated strictly within calendar-day boundaries and merged by date so that no future information could slip into the pipeline.</p>
<p>The dataset itself is a portrait of messy commercial reality rather than a tidy laboratory experiment. It was assembled from three sources: annual weekly planning spreadsheets, 189 individual daily forecast files recording both the farm&#8217;s own operational predictions and actual harvested weights, and sub-daily environmental sensor exports. After cleaning and harmonisation across years, the modelling dataset contained 2,280 daily observations, which rose to 3,540 after the researchers appended seven zero-yield days at the start and end of each harvest cycle. That padding step, an ablation analysis showed, mattered a great deal: removing it degraded the best model&#8217;s mean absolute error from about 340 kilograms to nearly 611 kilograms, while extending padding to 14 days lowered absolute error further but reduced explained variance. The seven-day configuration offered the best balance.</p>
<p>Feature engineering was deliberately biological. Rolling averages of prior yield over 1, 3, 7, 14 and 21 days captured what the authors call production memory, the tendency of strawberry harvests to persist across neighbouring days as overlapping cohorts of fruit mature and ripen. Rolling seven-day temperature averages and 14-day humidity averages represented accumulated environmental exposure, reflecting how heat and moisture stress modulate fruit development. Plant counts from the planning spreadsheets served as a static capacity indicator for each field. Together these features let the models learn whether production is rising, plateauing or declining, without ever violating temporal causality.</p>
<p>The model comparison spanned thirteen approaches, from simple ridge regression through gradient boosting frameworks such as XGBoost, LightGBM and CatBoost, to compact neural networks including a multilayer perceptron, a small LSTM and a temporal convolutional network, and finally TabPFN, a pretrained tabular foundation model. The results were striking. Under a conventional random 80/20 split, TabPFN achieved the lowest error with a mean absolute error of 340.18 kilograms and an R-squared of 0.733, while the neural sequence models performed worse than a constant mean predictor, producing negative R-squared values. Leave-one-out cross-validation told the same story, with TabPFN again leading at a mean absolute error of 338.09 kilograms. The authors note that with only six seasons of data, the strong inductive biases of tree ensembles and pretrained tabular models matter far more than architectural depth, which is precisely why the from-scratch deep learning baselines struggled.</p>
<p>The most deployment-relevant test, however, was forward-year validation: train on 2020 through 2024, then predict the entirely unseen 2025 season. Under this protocol TabPFN remained the strongest performer with a mean absolute error of 438.53 kilograms and an R-squared of 0.696, followed by Random Forest, Histogram Gradient Boosting and Extra Trees. The rise in error relative to the random split is not a failure, the researchers argue, but an honest measure of genuine interannual variation in weather, phenology and management. Permutation importance analysis on the 2025 test season confirmed that the models rely on biologically sensible variables: recent yield lags dominate the ranking, and interaction features combining recent production with temperature and humidity contribute measurably, suggesting that weather modulates an established production trajectory rather than acting as an independent driver.</p>
<p>The study is candid about its limits. All data come from a single farm, so the work demonstrates temporal generalisation to an unseen season within one commercial setting, not transfer across farms or regions. The available sensors captured only temperature and humidity, and the authors identify the absence of solar radiation, day length, irrigation records and growing degree days as a likely explanation for a recurring error mode: the systematic underestimation of peak harvest days, when similar recent conditions can conceal very different yield potential. Adding phenology-aware signals such as flower or truss counts, whether recorded manually or by computer vision, is flagged as a plausible route to better peak capture.</p>
<p>Even so, the implications are significant for an industry built on thin margins and perishable goods. The findings suggest that daily yield forecasting can be integrated into commercial decision-support systems using the data farms already collect, without exotic sensors or vast historical archives. In small, discontinuous seasonal datasets, the study concludes, strong inductive bias and regularisation beat model complexity, and rigorous forward-year evaluation is the only honest yardstick of what a forecasting system will actually deliver when the next season begins. For growers weighing the promise of agricultural AI, that combination of realism and rigour may be the most valuable harvest of all.</p>
<p><strong>Subject of Research:</strong> Daily strawberry yield forecasting with artificial intelligence in commercial horticulture</p>
<p><strong>Article Title:</strong> Daily strawberry yield forecasting using artificial intelligence in commercial horticulture</p>
<p><strong>Article References:</strong> Ahmed, S., &amp; Caldwell, N. H. (2026). Daily strawberry yield forecasting using artificial intelligence in commercial horticulture. <em>Smart Agricultural Technology, 15</em>, Article 102583. <a href="https://doi.org/10.1016/j.atech.2026.102583" rel="noopener noreferrer">https://doi.org/10.1016/j.atech.2026.102583</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.atech.2026.102583" rel="noopener noreferrer">10.1016/j.atech.2026.102583</a></p>
<p><strong>Keywords:</strong> strawberry, yield forecasting, machine learning, TabPFN, deep learning, precision agriculture, time series, commercial horticulture, feature engineering, forward-year validation, IoT sensors, UK farming</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">221894</post-id>	</item>
		<item>
		<title>Spiral Images Turn Time Series Into a Feast for Pretrained Vision Models</title>
		<link>https://scienmag.com/spiral-images-turn-time-series-into-a-feast-for-pretrained-vision-models/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 26 Sep 2026 21:27:50 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[anomaly detection applications]]></category>
		<category><![CDATA[Archimedean spiral]]></category>
		<category><![CDATA[autoencoders]]></category>
		<category><![CDATA[computer vision]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[generalization across data types]]></category>
		<category><![CDATA[geometric data representation]]></category>
		<category><![CDATA[innovative pattern recognition methods]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for time series]]></category>
		<category><![CDATA[pretrained models]]></category>
		<category><![CDATA[pretrained vision models]]></category>
		<category><![CDATA[SPIRAL]]></category>
		<category><![CDATA[spiral image transformation]]></category>
		<category><![CDATA[spiral-based data encoding]]></category>
		<category><![CDATA[time series]]></category>
		<category><![CDATA[time series anomaly detection]]></category>
		<category><![CDATA[time series to image conversion]]></category>
		<category><![CDATA[transfer learning]]></category>
		<category><![CDATA[TS2I techniques]]></category>
		<category><![CDATA[TS2I transformation]]></category>
		<category><![CDATA[TSB-AD benchmark]]></category>
		<category><![CDATA[visual analysis of time data]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=216417</guid>

					<description><![CDATA[A new spiral-based transformation converts time series into images that let pretrained vision models outperform specialized anomaly detectors across 23 benchmark datasets.]]></description>
										<content:encoded><![CDATA[<p>Anomaly detection in time series data underpins some of the most consequential applications of modern machine learning, from catching irregular heartbeats in hospital monitors to flagging failing sensors on spacecraft, spotting fraudulent trades in financial markets, and preventing outages in cloud infrastructure. Yet the field has long struggled with a stubborn problem: the algorithms that excel at one type of data often collapse on another. A new study published in the journal Machine Learning proposes an unexpectedly elegant solution that begins with a simple geometric idea — winding a stream of numbers around a spiral — and ends with a sweeping benchmark victory that could reshape how practitioners think about detecting the abnormal in the ordinary.</p>
<p>The method, called SPIRAL (Sequence ProjectIon via Radial Arrangement for anomaLy detection), was developed by Mateusz Smendowski, Kamil Faber, and Piotr Nawrocki of AGH University of Krakow together with Nathalie Japkowicz and Roberto Corizzo of American University in Washington, DC. It belongs to a family of techniques known as time series to image (TS2I) transformations, which convert one-dimensional signals into two-dimensional pictures. The appeal of this strategy is that it allows researchers to borrow the extraordinary power of computer vision models — networks pretrained on millions of natural images — and aim them at temporal data. Instead of inventing new architectures from scratch, the time series becomes something a ResNet or a Vision Transformer already knows how to look at.</p>
<p>Existing TS2I transformations, however, were designed mostly for forecasting and classification, not for anomaly detection, and they carry well-known liabilities. Methods such as Gramian Angular Summation Fields, Markov Transition Fields, and Recurrence Plots compute full pairwise interaction matrices, which means their computational cost grows quadratically with window length. They also tend to produce symmetric images in which large regions mirror each other, wasting pixel real estate on redundant information. Many require careful tuning of hyperparameters — embedding dimensions, quantile bins, normalization ranges — and several cause notorious training instability when paired with pretrained backbones. SPIRAL was engineered from the ground up to eliminate each of these weaknesses.</p>
<p>The core of the method is a two-arm Archimedean spiral. Each pixel in the output image is characterized by its radial distance from the center and its polar angle, computed with the two-argument arctangent so that all four quadrants are handled correctly. The transformation then unfolds the spiral by translating radial distance into angular progression, creating a continuous path that winds counter-clockwise from the center of the image to its edge. Along this path, the values of the time series window are painted as pixel intensities. The result is a spatially continuous representation in which temporally adjacent points remain geometrically close, and the radial dimension encodes the passage of time.</p>
<p>This geometry has a crucial consequence for anomaly detection. A sudden spike in the signal becomes an isolated bright spot on the spiral arm; a level shift appears as an abrupt transition; a change in variance shows up as a contrast variation; a frequency change alters the spacing of the bands. In other words, anomalies are converted into exactly the kinds of local visual distortions — edges, contours, and broken motifs — that convolutional filters and attention layers are structurally built to detect. The mapping is parameter-free, requires no hyperparameter tuning, and runs in linear time with respect to window length, a dramatic improvement over the quadratic cost of the Gramian family of methods. Its asymmetric layout also eliminates the mirror redundancy that plagues symmetric transforms, concentrating anomaly-relevant evidence into every pixel.</p>
<p>Beyond the transformation itself, the team formalized a standardized workflow for vision-based anomaly detection. Window sizes are chosen automatically by analyzing the autocorrelation function of the training data: the window length is set to the first lag at which autocorrelation falls below the 95 percent confidence bound derived from Bartlett&#8217;s theorem. This data-driven rule, applied only to the training split, produces windows that reflect the dominant temporal structure of each series without any manual calibration. Each window is then transformed into a single-channel image, replicated across RGB channels, and fed to an autoencoder trained to reconstruct normal patterns. Reconstruction error becomes the anomaly score, propagated point-wise so that results align with the benchmark&#8217;s fine-grained labels.</p>
<p>To test the approach, the researchers mounted what may be one of the most exhaustive evaluation campaigns in the field: 24,430 experiments across 23 datasets from the TSB-AD benchmark, spanning medical monitoring, aerospace sensors, server metrics, industrial facilities, human activity recognition, network traffic, environmental data, financial markets, and synthetic anomalies. SPIRAL was pitted against nine established TS2I transformations, three vision backbones (a lightweight CNN, an ImageNet-pretrained ResNet18, and a Pyramid Vision Transformer), three transfer learning strategies, and 32 time-domain baselines ranging from classical statistical methods to modern foundation models such as MOMENT, TimesFM, and Chronos, all judged across nine evaluation metrics.</p>
<p>The results were striking. On VUS-PR, a demanding threshold-independent metric that rewards early and precise detection in imbalanced data, the best SPIRAL configuration achieved the lowest average rank among all 102 evaluated methods, according to Friedman and post-hoc Nemenyi statistical tests. Notably, all fifteen top-ranked methods were TS2I-based configurations, and the strongest time-domain baseline, KShapeAD, trailed more than ten rank positions behind. The authors attribute this not to domination of individual datasets but to a fundamentally different performance profile: while time-domain methods like KShapeAD and Sub-PCA can post spectacular scores on particular signals and near-zero on others, SPIRAL-based detection never fell below 0.06 VUS-PR on any dataset while reaching 0.95 on Exathlon and 0.88 on NEK. That cross-domain consistency, achieved without any domain-specific tuning, is precisely what matters in manufacturing, healthcare, and cloud operations.</p>
<p>Efficiency and stability told an equally compelling story. Despite producing geometrically rich images, SPIRAL ranked second only to naive array reshaping in preprocessing throughput, statistically indistinguishable from it and dramatically faster than Gramian-based methods, which were up to 55 times slower on large-window datasets. More importantly, training dynamics analysis revealed that SPIRAL exhibits 40 percent lower training instability than its closest TS2I competitor, measured as the coefficient of variation of the loss in the final training stage — a property the authors argue is critical for continual and online learning scenarios where unstable convergence can lead to catastrophic forgetting. Perhaps the most provocative finding is data-centric: the gap between the best and worst TS2I transformations was nearly ten times larger than the gap between backbone architectures, and an order of magnitude larger than the difference among transfer learning strategies. The choice of input representation, in other words, governs detection performance far more than model capacity or fine-tuning strategy.</p>
<p>The study is candid about its limits. The evaluation covers univariate series, following the convention of all benchmarked TS2I methods, and extending the paradigm to multivariate data — with its questions of channel composition and cross-variable anomaly structure — remains an open challenge. The uniform propagation of window scores to individual points also smooths event boundaries, costing ground on segment-level metrics such as point-adjust F1. Even so, the broader message lands with force: sometimes the path to better machine intelligence is not a bigger model but a smarter picture. By winding a stream of numbers into a spiral, the researchers have shown that anomalies hidden in time can become patterns visible in space — and that the eyes of a pretrained vision network, given the right image, can see them.</p>
<p><strong>Subject of Research:</strong> Time series to image transformation for vision-based anomaly detection</p>
<p><strong>Article Title:</strong> SPIRAL: A Novel Time Series to Image (TS2I) Transformation Method for Vision-Based Anomaly Detection</p>
<p><strong>Article References:</strong> Smendowski, M., Faber, K., Nawrocki, P., Japkowicz, N., &amp; Corizzo, R. (2026). SPIRAL: A Novel Time Series to Image (TS2I) Transformation Method for Vision-Based Anomaly Detection. <em>Machine Learning, 115</em>(10), Article 231. <a href="https://doi.org/10.1007/s10994-026-07134-7" rel="noopener noreferrer">https://doi.org/10.1007/s10994-026-07134-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10994-026-07134-7" rel="noopener noreferrer">10.1007/s10994-026-07134-7</a></p>
<p><strong>Keywords:</strong> anomaly detection, time series, computer vision, machine learning, SPIRAL, TS2I transformation, autoencoders, transfer learning, TSB-AD benchmark, Archimedean spiral, deep learning, pretrained models</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">216417</post-id>	</item>
		<item>
		<title>AI Learns to Spot Broken Sensors in Smart Farms Before Crops Suffer</title>
		<link>https://scienmag.com/ai-learns-to-spot-broken-sensors-in-smart-farms-before-crops-suffer/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Fri, 25 Sep 2026 23:44:49 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-powered farm system maintenance]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[anomaly detection in irrigation systems]]></category>
		<category><![CDATA[automated irrigation]]></category>
		<category><![CDATA[automated sensor fault identification]]></category>
		<category><![CDATA[BiGRU-VAE]]></category>
		<category><![CDATA[crop health monitoring with AI]]></category>
		<category><![CDATA[data-driven farm irrigation control]]></category>
		<category><![CDATA[deep learning models for farm management]]></category>
		<category><![CDATA[drip irrigation]]></category>
		<category><![CDATA[IoT sensor reliability in agriculture]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning for agricultural sensors]]></category>
		<category><![CDATA[minimal labeled data sensor fault detection]]></category>
		<category><![CDATA[precision agriculture]]></category>
		<category><![CDATA[scalable anomaly detection for large farms]]></category>
		<category><![CDATA[semi-supervised learning]]></category>
		<category><![CDATA[sensor failure prediction in precision agriculture]]></category>
		<category><![CDATA[sensor faults]]></category>
		<category><![CDATA[Smart Agriculture]]></category>
		<category><![CDATA[Smart farm sensor fault detection]]></category>
		<category><![CDATA[sparse autoencoder]]></category>
		<category><![CDATA[TabNet]]></category>
		<category><![CDATA[time series]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=215393</guid>

					<description><![CDATA[Researchers at Anhui Agricultural University have developed DE-TabNet, a semi-supervised AI model that detects sensor faults in smart drip irrigation systems with up to 86.9 percent accuracy.]]></description>
										<content:encoded><![CDATA[<p>Smart farms run on a quiet promise: that the sensors watching soil moisture, temperature, and nutrient flow are telling the truth. When a sensor drifts, freezes, or fails outright, the automated systems that irrigate fields and regulate greenhouses can respond to fiction instead of fact, quietly overwatering crops or letting them dry out. A new study published in Complex &amp; Intelligent Systems tackles this vulnerability head-on with a machine learning model called DE-TabNet, which its developers at Anhui Agricultural University say can detect faulty sensor behavior in drip irrigation systems with up to 86.9 percent accuracy, even when almost none of the training data has been labeled by human experts.</p>
<p>The research, led by Jun Zhu and colleagues at the School of Information and Artificial Intelligence in Hefei, China, addresses a problem that has long frustrated engineers of agricultural control systems. Traditional anomaly detection in these settings has relied on manual inspection, which is slow, labor-intensive, and impractical across large farms. Purely supervised machine learning approaches, meanwhile, demand large volumes of labeled data, meaning human annotators must painstakingly mark which sensor readings are normal and which are faults. In real agricultural deployments, such labels are scarce because failures are rare events and expert annotation is expensive. Unsupervised methods sidestep the labeling problem but often struggle to distinguish genuine sensor faults from unusual but legitimate patterns, such as those caused by extreme weather or unusual irrigation schedules.</p>
<p>DE-TabNet&#8217;s central innovation is a semi-supervised architecture that blends the strengths of both worlds. The model first learns from vast quantities of unlabeled sensor data through what the authors describe as an unsupervised feature fusion network. This network combines two complementary components. The first is a bidirectional gated recurrent unit paired with a variational autoencoder, abbreviated BiGRU-VAE, which learns the temporal structure of sensor readings. The bidirectional design means the network reads each sequence of measurements both forward and backward in time, capturing how a reading relates to what came before and what follows. The variational autoencoder component learns a compressed, probabilistic representation of that temporal information, forcing the model to distill the essential dynamics of normal system behavior rather than memorizing raw values.</p>
<p>The second component is a sparse autoencoder, which operates on the feature level rather than the time level. While the BiGRU-VAE captures temporal sequences, the sparse autoencoder extracts dependencies among different sensor features, learning which measurements tend to move together in a healthy system. A soil moisture sensor, for example, should respond in characteristic ways to irrigation events and to changes in humidity readings from neighboring sensors. When those relationships break down, the sparse representation of the data shifts in ways the model can detect. By fusing temporal and feature-level representations, the unsupervised stage builds a rich internal picture of what normal operation looks like, using only the raw, unlabeled data streams that smart farms generate continuously.</p>
<p>Once this unsupervised foundation is in place, the second stage of DE-TabNet fine-tunes the learned representations using a small set of labeled examples. This stage employs TabNet, a neural network architecture designed for tabular data that uses a form of attention to select the most relevant features for each decision. Because the underlying representations were already learned from abundant unlabeled data, only a limited amount of labeled data is needed to teach the model where the boundary between normal and anomalous behavior lies. This is the essence of semi-supervised learning: leverage the cheap, plentiful unlabeled data for general understanding, and reserve scarce labeled data for precise calibration.</p>
<p>The team evaluated DE-TabNet on a real-world smart drip irrigation dataset, a demanding testbed because drip irrigation systems involve tightly coupled sensors and actuators whose interactions are subtle. The results were striking. DE-TabNet achieved an accuracy of up to 86.9 percent and an F1 score of 80.8 percent, a combined measure of precision and recall that is particularly informative when anomalies are rare. Compared against baseline models, these figures represent improvements of 11.6 percentage points in accuracy and 3.82 percentage points in F1 score. The baselines were not weak competitors: they included DAGMM, a deep learning approach that mixes autoencoders with Gaussian mixture models; iForest, or isolation forest, a widely used classical algorithm; VAE-LSTM, a hybrid of variational autoencoders and long short-term memory networks; OCSVM, the one-class support vector machine; and Semi-TabNet, a semi-supervised method closely related to the new model&#8217;s supervised component.</p>
<p>Beating such a diverse field of established methods suggests that the advantage comes from the architecture itself rather than from any single trick. The bidirectional temporal encoding appears to matter: anomalies in irrigation systems often manifest as sequences, not isolated points. A stuck sensor produces a flatline; a drifting sensor produces a slow, systematic bias; a failing actuator produces responses that lag behind commands. Detecting these patterns requires a model that understands context in time, which is precisely what the BiGRU component provides. Meanwhile, the sparse autoencoder&#8217;s feature dependency extraction catches faults that a purely temporal model might miss, such as a single sensor whose readings become inconsistent with the rest of the network even while its own time series looks plausible.</p>
<p>Importantly, the researchers did not stop at the irrigation dataset. They validated DE-TabNet on additional public datasets and found that the model generalized well, maintaining strong performance beyond the specific agricultural context in which it was developed. This generalization is a critical property for any anomaly detection system intended for deployment, because real-world data distributions shift over time and across sites. A model that only works on one farm&#8217;s sensors would be of limited value; one that transfers across time series sensor systems could find applications in industrial automation, environmental monitoring, and infrastructure management, wherever networks of sensors feed automated control loops.</p>
<p>The practical stakes are considerable. Agriculture is increasingly dependent on automated systems that make decisions without human intervention, from precision irrigation to climate control in protected cultivation. The study was supported by Chinese research programs including the Special Fund for Anhui Agriculture Research System and the National Natural Science Foundation of China, reflecting national investment in intelligent farming infrastructure. As these systems proliferate, the cost of undetected sensor faults grows: a misreporting moisture sensor in an automated drip system can waste water, leach fertilizer, or stress crops during critical growth stages. Reliable anomaly detection acts as a form of insurance, allowing control systems to flag suspicious data before acting on it, or to fall back on safe operating modes while faults are investigated.</p>
<p>The work also illustrates a broader trend in applied machine learning: the rise of architectures that respect the structure of their data. Sensor streams are simultaneously temporal, with meaning encoded in sequences, and relational, with meaning encoded in correlations among channels. Models that encode both dimensions bidirectionally, then refine their understanding with whatever labels are available, mirror how a human expert actually works, forming an intuition for normal behavior from experience and then applying judgment to the rare cases that stand out. If DE-TabNet&#8217;s reported performance holds up in field deployments, the silent failure of a single sensor may no longer be able to silently mislead the systems that feed us. The research is open access, and the authors report no competing financial interests, allowing other teams to build directly on the approach as smart agriculture continues its rapid, data-driven evolution.</p>
<p><strong>Subject of Research:</strong> Semi-supervised anomaly detection for smart agricultural sensor and control systems</p>
<p><strong>Article Title:</strong> Anomaly detection in smart agricultural control systems using semi-supervised bidirectional encoding</p>
<p><strong>Article References:</strong> Anomaly detection in smart agricultural control systems using semi-supervised bidirectional encoding. (n.d.). <a href="https://doi.org/10.1007/s40747-026-02512-z" rel="noopener noreferrer">https://doi.org/10.1007/s40747-026-02512-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s40747-026-02512-z" rel="noopener noreferrer">10.1007/s40747-026-02512-z</a></p>
<p><strong>Keywords:</strong> smart agriculture, anomaly detection, semi-supervised learning, sensor faults, drip irrigation, BiGRU-VAE, TabNet, sparse autoencoder, machine learning, automated irrigation, time series, precision agriculture</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">215393</post-id>	</item>
		<item>
		<title>AI Learns From History to Predict Ozone Pollution Across Chinese Cities</title>
		<link>https://scienmag.com/ai-learns-from-history-to-predict-ozone-pollution-across-chinese-cities/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 23:28:51 +0000</pubDate>
				<category><![CDATA[Climate]]></category>
		<category><![CDATA[AI in air quality modeling]]></category>
		<category><![CDATA[Air pollution forecasting]]></category>
		<category><![CDATA[air quality]]></category>
		<category><![CDATA[China]]></category>
		<category><![CDATA[Chinese cities air pollution analysis]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning limitations in pollution forecasting]]></category>
		<category><![CDATA[historical pollution data retrieval]]></category>
		<category><![CDATA[knowledge base]]></category>
		<category><![CDATA[knowledge-based AI for environmental monitoring]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[meteorology]]></category>
		<category><![CDATA[multi-source data integration for air quality]]></category>
		<category><![CDATA[neural network correction methods]]></category>
		<category><![CDATA[ozone forecasting]]></category>
		<category><![CDATA[ozone pollution episodes recognition]]></category>
		<category><![CDATA[ozone pollution prediction]]></category>
		<category><![CDATA[photochemical reactions in ozone formation]]></category>
		<category><![CDATA[pollution warning]]></category>
		<category><![CDATA[retrieval-augmented forecasting]]></category>
		<category><![CDATA[time series]]></category>
		<category><![CDATA[urban agglomerations]]></category>
		<category><![CDATA[weather influence on ozone levels]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=213355</guid>

					<description><![CDATA[Researchers in China have developed OzoneKBNet, a retrieval-augmented AI framework that mines historical pollution episodes to deliver more accurate 48-hour ozone forecasts across seven urban agglomerations.]]></description>
										<content:encoded><![CDATA[<p>Ground-level ozone is one of the most stubborn air pollution problems facing the world&#8217;s megacities, and forecasting it accurately has long frustrated scientists. Unlike particulate matter, ozone is not emitted directly but forms in complex photochemical reactions involving nitrogen oxides, volatile compounds, and sunlight, all modulated by weather patterns that shift from hour to hour and year to year. Now, a team of researchers in China has unveiled a new artificial intelligence framework that takes an unusual approach: instead of relying solely on a neural network&#8217;s learned parameters, it actively retrieves historical episodes of pollution from a curated knowledge base and uses them to correct its forecasts. The system, called OzoneKBNet, is described in the journal Air Quality, Atmosphere &amp; Health.</p>
<p>The research, led by Mingkun Zhu of the Department of Automation at Taiyuan University of Technology, together with Xiaoxia Han, Jinde Wu, Haonan Zhu, and Wenxin Chai, targets a problem that most data-driven forecasting models handle poorly. Conventional deep learning models encode historical patterns implicitly within their trained weights, which means that once training is complete, the model has no explicit memory of specific pollution episodes it can consult. If a city experiences a recurring type of ozone event, driven for example by a particular combination of stagnant air, high temperatures, and regional transport of precursors, a standard model must re-learn that pattern from scratch each time it appears in slightly different form. OzoneKBNet instead makes those recurring episodes searchable.</p>
<p>The framework works by forecasting station-level ozone concentrations over the next 48 hours based on the preceding 96 hours of observations. Its inputs include ozone itself, nitrogen dioxide, fine particulate matter with a diameter of 2.5 micrometers or less, temperature, and relative humidity, a set of variables chosen to capture both the chemical precursors of ozone formation and the meteorological conditions that govern it. When a forecast is needed, the system searches a knowledge base of historical situations for multi-scale analogs drawn from across the monitoring network, retrieving similar episodes not just from the station in question but from neighboring stations as well. This cross-station retrieval is crucial in urban agglomerations, where pollution transported from upwind cities can dominate local ozone behavior.</p>
<p>Retrieval alone is not enough, and the researchers built in a sophisticated filtering stage. Candidate analogs are re-ranked according to two criteria: representation similarity, which measures how closely the retrieved historical situation matches the current one in the model&#8217;s learned feature space, and local-trend consistency, which checks whether the historical episode&#8217;s near-term trajectory actually resembles what is happening now. Only analogs that pass both tests contribute to the forecast. Their known future trajectories are then used as bounded, reliability-aware corrections to a lightweight long short-term memory network, or LSTM, that provides the baseline prediction. The word bounded matters: rather than allowing retrieved evidence to override the neural forecast wholesale, the system applies corrections within limits calibrated to how trustworthy the retrieved analogs appear to be.</p>
<p>A particularly careful aspect of the study is its handling of temporal data leakage, a common pitfall in machine learning applications to time series. If a model is trained and tested on data drawn from overlapping periods, it can inadvertently peek at information from the future, inflating its apparent accuracy. The team partitioned their data strictly by year: observations from 2023 were used for normalization, for pretraining the retrieval encoder, and for constructing the knowledge base; data from 2024 were reserved for supervised training; and a fixed set of samples from 2025 served as the test set. This design means the model was evaluated on genuinely unseen future conditions, a far more demanding test than random splitting of the data.</p>
<p>The results span seven Chinese urban agglomerations, offering a broad regional assessment of the framework&#8217;s capabilities. Across all stations, OzoneKBNet achieved macro-average mean absolute error of 20.06 micrograms per cubic meter and root mean square error of 24.45 micrograms per cubic meter over the 48-hour forecast horizon. It posted the lowest regional mean absolute error in six of the seven regions evaluated, suggesting that the retrieval-based approach generalizes well across geographically and climatologically diverse areas rather than being tuned to a single city&#8217;s quirks. Compared with a station-specific LSTM baseline, the framework reduced mean absolute error by 1.6 percent and root mean square error by 1.2 percent.</p>
<p>Those percentage improvements may sound modest, but their significance lies in what they demonstrate about the architecture of forecasting systems. The baseline LSTM is already a strong performer, trained on the same data, and squeezing additional accuracy out of it through any independent mechanism is difficult. The gains from OzoneKBNet come from an entirely complementary source: explicit regional historical evidence that the parametric model cannot access on its own. In effect, the study shows that a model&#8217;s learned parameters and a searchable library of past episodes contain partially non-overlapping information, and combining the two yields forecasts more accurate than either alone. This principle echoes a broader trend in artificial intelligence, where retrieval-augmented generation has transformed language models by letting them consult external documents rather than relying purely on memorized knowledge.</p>
<p>The intellectual lineage of the analog approach stretches back decades in atmospheric science. Long-range weather forecasting based on analogs, matching current conditions to historical situations and borrowing their outcomes, was explored as early as the late 1980s, and modified k-nearest-neighbor methods have been used for real-time weather prediction. What is new is the marriage of this old idea with modern deep learning: an encoder trained to map multivariate pollution and meteorology sequences into a representation space where similarity search becomes meaningful, combined with a re-ranking scheme and a residual correction mechanism that treats retrieved evidence with appropriate caution. The framework also connects to recent work on retrieval-augmented time series forecasting in the machine learning community, where researchers have begun applying similar ideas to general forecasting foundation models.</p>
<p>Why does station-level forecasting matter so much for ozone? Regional averages can hide enormous local variation. Ozone concentrations at a single monitoring station depend on the interplay of local emissions, photochemistry, boundary-layer dynamics, and transport from surrounding areas, and public health warnings are ultimately issued for specific places where people live and breathe. A framework that can deliver accurate 48-hour forecasts at the station level, across an entire urban agglomeration, gives authorities the granularity needed to warn residents before pollution episodes peak, to time traffic restrictions or industrial curbs more effectively, and to study how ozone responds to control measures. The authors suggest their approach may support short-term ozone-pollution episode warning, and the 48-hour horizon is well matched to the timescale on which such interventions operate.</p>
<p>The team has also made an unusually strong commitment to openness. The source code of OzoneKBNet and the complete fixed 2025 test set are available in a public repository on GitHub, along with the model implementation, running scripts, data-format instructions, and fixed test samples that allow others to inspect the experimental pipeline and evaluation protocol. The raw air quality monitoring data and ERA5 reanalysis data analyzed in the study come from publicly available providers, although redistribution restrictions prevent the processed training datasets from being shared directly. This transparency matters in a field where evaluation practices vary widely and where, as the strict 2023-2024-2025 split shows, the details of how models are tested can determine whether reported accuracy reflects real forecasting skill or subtle leakage. For a problem as consequential as urban ozone, reproducible methods and honest evaluation are not luxuries but necessities, and this work offers both alongside a genuinely new architectural idea.</p>
<p><strong>Subject of Research:</strong> Retrieval-augmented deep learning for station-level ground-level ozone forecasting in urban agglomerations</p>
<p><strong>Article Title:</strong> A retrieval-augmented historical knowledge-base framework with residual correction for station-level ozone forecasting across urban agglomerations</p>
<p><strong>Article References:</strong> Zhu, M., Han, X., Wu, J., Zhu, H., &amp; Chai, W. (2026). A retrieval-augmented historical knowledge-base framework with residual correction for station-level ozone forecasting across urban agglomerations. <em>Air Quality, Atmosphere &amp;amp; Health, 19</em>(10), Article 215. <a href="https://doi.org/10.1007/s11869-026-02105-2" rel="noopener noreferrer">https://doi.org/10.1007/s11869-026-02105-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s11869-026-02105-2" rel="noopener noreferrer">10.1007/s11869-026-02105-2</a></p>
<p><strong>Keywords:</strong> ozone forecasting, air quality, machine learning, retrieval-augmented forecasting, LSTM, urban agglomerations, knowledge base, time series, China, pollution warning, deep learning, meteorology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">213355</post-id>	</item>
		<item>
		<title>Same Temperature, Different Microbes: 13-Year Baltic Sea Study Reveals Hidden Seasonal Hysteresis</title>
		<link>https://scienmag.com/same-temperature-different-microbes-13-year-baltic-sea-study-reveals-hidden-seasonal-hysteresis/</link>
		
		<dc:creator><![CDATA[Morgan Morrow]]></dc:creator>
		<pubDate>Thu, 24 Sep 2026 11:16:23 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[16S rRNA]]></category>
		<category><![CDATA[16S rRNA gene sequencing in marine ecosystems]]></category>
		<category><![CDATA[Baltic Sea]]></category>
		<category><![CDATA[Baltic Sea microbial communities]]></category>
		<category><![CDATA[Cyanobacteria]]></category>
		<category><![CDATA[free-living bacteria]]></category>
		<category><![CDATA[hidden seasonal microbial patterns]]></category>
		<category><![CDATA[hysteresis]]></category>
		<category><![CDATA[impacts of seasonal changes on ocean bacteria]]></category>
		<category><![CDATA[long-term microbial monitoring]]></category>
		<category><![CDATA[long-term ocean microbiology study]]></category>
		<category><![CDATA[marine microbial ecology]]></category>
		<category><![CDATA[marine microbial seasonal variation]]></category>
		<category><![CDATA[microbial community dynamics in freezing and summer conditions]]></category>
		<category><![CDATA[microbial diversity in Baltic Sea]]></category>
		<category><![CDATA[microbial ecological assumptions challenged]]></category>
		<category><![CDATA[microbial seasonal hysteresis]]></category>
		<category><![CDATA[microbiome]]></category>
		<category><![CDATA[particle-associated bacteria]]></category>
		<category><![CDATA[particle-associated vs free-living microbes]]></category>
		<category><![CDATA[prokaryoplankton]]></category>
		<category><![CDATA[rare biosphere]]></category>
		<category><![CDATA[seasonal succession]]></category>
		<category><![CDATA[time series]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=212354</guid>

					<description><![CDATA[A 13-year Baltic Sea time series shows that bacterial communities at identical temperatures differ radically between spring and autumn, revealing hysteresis-like dynamics in marine microbiomes.]]></description>
										<content:encoded><![CDATA[<p>For thirteen years, researchers at the Linnaeus Microbial Observatory in the Baltic Proper pulled seawater samples from the same spot, essentially every two weeks, through freezing winters, cyanobacterial summers, and everything in between. What they found challenges one of the most basic assumptions in ocean ecology: that if you know the temperature, you know roughly which microbes should be there. As it turns out, the Baltic Sea&#8217;s bacterial communities at 12 degrees Celsius in spring look strikingly different from those at 12 degrees in autumn, even though the thermometer reads the same. The way from winter to summer, it seems, is not the same as the way from summer to winter.</p>
<p>The study, published in the journal Ocean Microbiology, analyzed 726 samples of 16S rRNA gene amplicons collected between 2011 and 2023, detecting more than 67,000 distinct prokaryotic amplicon sequence variants, or ASVs. The team, led by Carina Bunse of Linnaeus University and the University of Gothenburg together with Jarone Pinhassi and colleagues, split each sample into two size fractions: particle-associated prokaryotes larger than 3 micrometers, and free-living cells in the 0.2 to 3 micrometer range. This dual-fraction approach, sustained over more than a decade, is rare in oceanography and allowed the researchers to ask questions that shorter studies simply cannot answer.</p>
<p>The sampling site itself experiences dramatic seasonal forcing. Winter water temperatures drop below 5 degrees Celsius while inorganic nutrients accumulate to concentrations of 2 to 3 micromolar nitrate plus nitrite and up to 1 micromolar phosphate. By summer, temperatures climb to 18 degrees or higher, and nutrients are stripped to near detection limits. Two major phytoplankton blooms punctuate the year, one typically in April dominated by dinoflagellates and diatoms, and a cyanobacterial bloom in July and August. Prokaryotic cell counts swing roughly six-fold, from about half a million cells per milliliter in winter to more than three million in summer.</p>
<p>One of the study&#8217;s clearest findings concerns biodiversity. Particle-associated communities consistently harbored higher diversity than free-living ones, with estimated and observed richness significantly elevated in the particle fraction across all seasons, peaking in winter. This makes ecological sense: marine particles, whether aggregates of living and dead plankton, fecal pellets, or condensed organic material, offer a patchwork of microscopic habitats, each with its own chemistry and resources. Free-living seawater, by contrast, selects for streamlined cells adapted to a more homogeneous environment. Yet the evenness of the communities, measured by Shannon and inverse Simpson indices, differed between fractions only in winter, suggesting that while particles host more species, the balance of abundance remains broadly similar.</p>
<p>The researchers went further, using a permutation-based particle-association niche index to determine which taxa genuinely prefer particles and which prefer open water. The results revealed remarkable fine-scale specialization. Among the Bacteroidia and Gammaproteobacteria, most families showed strong particle association, including Alteromonadaceae, Pseudomonadaceae, and Shewanellaceae. Among the Alphaproteobacteria, families split roughly evenly, with Rhizobiaceae and Caulobacteraceae favoring particles while Rhodospirillaceae and the famously abundant Pelagibacteraceae favored the free-living fraction. Most strikingly, some families such as Flavobacteriaceae, Burkholderiaceae, and Rhodobacteraceae contained species spread across the entire spectrum, meaning that even within a single family, closely related organisms have carved out opposite lifestyles.</p>
<p>The filamentous cyanobacteria provided a natural test case. Genera like Dolichospermum, Nodularia, and Aphanizomenon form the Baltic Sea&#8217;s notorious summer blooms, thriving on nitrogen limitation and residual phosphorus. In the dataset, Cyanobacteriia dominated the particle fraction in summer, sometimes exceeding 90 percent of relative abundance, while contributing at most around 30 percent to the free-living fraction, largely because their filaments are physically excluded from the smaller filter. Non-filamentous cyanobacterial families, by contrast, showed low particle-association indices, confirming that the fractionation captures real biology rather than an artifact of filtration.</p>
<p>Perhaps the most unexpected finding involves the rare biosphere. Most ASVs in the dataset were consistently rare, with mean relative abundances below 0.7 percent. But the abundant taxa, the so-called core species that appear in high numbers every year, turned out to be surprisingly ephemeral. A majority of the 50 most abundant ASVs in each fraction collapsed into the rare biosphere for extended stretches of the year, sometimes falling below detection entirely, before rebounding in their characteristic season. Recruitment from the rare biosphere, previously viewed as an occasional event following disturbance, appears here to be an annual routine for most dominant taxa in this strongly seasonal environment.</p>
<p>The dataset also captured rare microbes responding to dramatic events. Campylobacteria, formerly known as Epsilonproteobacteria, normally inhabit the oxic-anoxic interface in deeper Baltic waters. Yet in the winters of 2011/12 and 2014/15, coinciding with major North Sea inflow events that pushed saline, oxygenated water into the Baltic&#8217;s deep basins, a Sulfurimonas ASV surged to relative abundances exceeding 80 percent in surface waters in December 2011. These episodic blooms of typically deep-water specialists offer a microbial fingerprint of large-scale oceanographic disturbances, lingering for months before the taxa retreated back into rarity.</p>
<p>The centerpiece of the study, however, is the demonstration of direction-dependent dynamics, a phenomenon ecologists call hysteresis. Using generalized additive models that compared a single temperature-abundance relationship against models allowing separate smooths for warming and cooling phases, the researchers found that most co-occurring taxon clusters behaved differently depending on the direction of environmental change. One free-living cluster dominated by Bacteroidia reached roughly 50 percent relative abundance at 7 degrees Celsius in spring but only about 15 percent at the same temperature in autumn. Another cluster, rich in Acidimicrobiia and Actinomycetes, accounted for more than 40 percent of the community at 10 degrees in autumn but a mere 5 percent at 10 degrees in spring. The same held for day length: one cluster exceeded 60 percent abundance at 15 hours of daylight in spring but contributed less than 10 percent at 15 hours in fall.</p>
<p>What does this mean for how we understand the ocean&#8217;s microbial engines? Temperature undeniably shapes microbial metabolism, and day length defines the seasonal calendar at any latitude. But the study&#8217;s authors argue that neither variable directly determines community composition. Instead, the identity of the microbes at any moment reflects the community&#8217;s history, the direction of environmental change, the quantity and quality of organic and inorganic nutrients, and a web of biotic interactions with phytoplankton, grazers, and viruses. Spring waters carry fresh labile organic matter from the phytoplankton bloom; autumn waters carry a different cocktail of substrates after a summer of production and degradation. The microbes respond to this context, not just to the thermometer. Given that seasonal shifts in dominant microorganisms ripple through marine food webs and biogeochemical cycles, the finding suggests that climate-driven warming may alter ecosystems in ways that simple temperature correlations cannot predict. For a brackish inland sea already stressed by eutrophication and hypoxia, that is a sobering message, and one that only long-term, high-frequency, size-fractionated time series like this one could have delivered.</p>
<p><strong>Subject of Research:</strong> Seasonal succession and direction-dependent dynamics of particle-associated and free-living prokaryotic communities in the Baltic Sea</p>
<p><strong>Article Title:</strong> Comparable temperatures but different microbiomes: long-term time series exposes seasonally divergent Baltic Sea prokaryotic communities</p>
<p><strong>Article References:</strong> Bunse, C., Farnelid, H., Lindehoff, E., Martínez-García, S., Fridolfsson, E., Di Leo, D., Martínez, C. P., Pontiller, B., Lindh, M. V., Sjöstedt, J., Lundin, D., Legrand, C., &amp; Pinhassi, J. (2026). Comparable temperatures but different microbiomes: long-term time series exposes seasonally divergent Baltic Sea prokaryotic communities. <em>Ocean Microbiology, 2</em>(1), Article 5. <a href="https://doi.org/10.1186/s44375-026-00011-7" rel="noopener noreferrer">https://doi.org/10.1186/s44375-026-00011-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s44375-026-00011-7" rel="noopener noreferrer">10.1186/s44375-026-00011-7</a></p>
<p><strong>Keywords:</strong> Baltic Sea, prokaryoplankton, microbiome, seasonal succession, time series, rare biosphere, hysteresis, particle-associated bacteria, free-living bacteria, cyanobacteria, 16S rRNA, marine microbial ecology</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">212354</post-id>	</item>
		<item>
		<title>Hybrid AI Model Tames the Chaos of GDP Forecasting Across Seven Economies</title>
		<link>https://scienmag.com/hybrid-ai-model-tames-the-chaos-of-gdp-forecasting-across-seven-economies/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 23:03:06 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced time series prediction methods]]></category>
		<category><![CDATA[algorithmic tuning in deep learning models]]></category>
		<category><![CDATA[Bayesian optimization]]></category>
		<category><![CDATA[CNN]]></category>
		<category><![CDATA[convolutional neural networks for economic data]]></category>
		<category><![CDATA[data-driven GDP prediction techniques]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[Diebold-Mariano test]]></category>
		<category><![CDATA[econometrics]]></category>
		<category><![CDATA[economic forecasting across multiple countries]]></category>
		<category><![CDATA[GDP forecasting]]></category>
		<category><![CDATA[hybrid AI models for economic indicators]]></category>
		<category><![CDATA[hybrid deep learning models]]></category>
		<category><![CDATA[long short-term memory networks in economics]]></category>
		<category><![CDATA[LSTM]]></category>
		<category><![CDATA[machine learning in macroeconomic forecasting]]></category>
		<category><![CDATA[macroeconomic indicators]]></category>
		<category><![CDATA[multivariable data]]></category>
		<category><![CDATA[neural network architecture for GDP]]></category>
		<category><![CDATA[nonlinear relationships in GDP prediction]]></category>
		<category><![CDATA[Pearson correlation]]></category>
		<category><![CDATA[temporal convolutional network]]></category>
		<category><![CDATA[time series]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=211070</guid>

					<description><![CDATA[A hybrid CNN–LSTM deep learning model tuned by Bayesian optimization outperformed classical and standalone neural benchmarks in quarterly GDP forecasts for seven countries.]]></description>
										<content:encoded><![CDATA[<p>Gross domestic product is the single most watched number in economics, yet it remains one of the hardest to predict. Quarterly output figures wobble under the combined pressure of consumer spending, trade flows, industrial production, interest rates and dozens of other indicators, all interacting in stubbornly nonlinear ways. Traditional econometric tools, from ARIMA models to vector autoregression, have long dominated the field, but they struggle when relationships between variables bend and shift over time. A new study published in the International Journal of Data Science and Analytics argues that a carefully engineered hybrid of two deep learning architectures, tuned by an algorithmic search method borrowed from machine learning research, can forecast GDP more accurately than either classical benchmarks or standalone neural networks.</p>
<p>The research, conducted by Sana Hamiane and Youssef Ghanou of Moulay Ismail University of Meknes in Morocco together with Gabriella Casalino of the University of Bari Aldo Moro in Italy, combines a convolutional neural network with a long short-term memory network in a single forecasting pipeline. Each component plays a distinct role. The convolutional layers act as feature extractors, scanning windows of multivariable economic data to detect local patterns, such as simultaneous movements in exports, industrial production and inflation, that might signal an upcoming change in output. The LSTM layers then take over, using their gated memory cells to track how these patterns evolve across consecutive quarters and to decide which information from the past should be retained or forgotten when producing a prediction.</p>
<p>This division of labor matters because GDP time series carry two kinds of structure at once. Spatially, the many macroeconomic indicators feeding into the model form intricate cross-correlations that convolutional filters are well suited to capture. Temporally, the economy exhibits long-range dependencies, momentum and regime shifts that recurrent architectures like LSTM were explicitly designed to remember. Models that handle only one of these dimensions tend to leave predictive signal on the table. By stacking the two architectures, the hybrid model can, in principle, exploit both, and the new results suggest that this is exactly what happens in practice.</p>
<p>Before any training began, the researchers confronted a problem that plagues every multivariable forecasting effort: which indicators should be included? Feeding a neural network dozens of weakly related variables invites overfitting and noise. The team turned to the Pearson correlation coefficient, a classical statistical measure of linear association, to rank candidate macroeconomic variables by their correlation with GDP. Only the most relevant features survived this filter. The resulting indicator sets differed by country, reflecting the structure of each economy. For the United States, France, Italy, Spain and Germany the model drew on measures such as private and government consumption, imports and exports, gross fixed capital formation, manufacturing output, consumer price indices, exchange rates and long-term interest rates. For Morocco, the study used a detailed breakdown of sectoral value added, spanning agriculture, mining, construction, transport, finance and other industries.</p>
<p>Architecture alone, however, does not guarantee good forecasts. Deep learning models are notoriously sensitive to hyperparameters: the number of layers, the size of filters and hidden states, learning rates, batch sizes and dropout rates all shape how well a network generalizes. Manual tuning is slow and subjective, and grid searches become computationally explosive as the number of settings grows. The authors addressed this with Bayesian optimization, a sequential model-based search strategy that builds a probabilistic surrogate of the model&#8217;s performance and uses it to select the most promising hyperparameter configurations to test next. Instead of blindly exploring the space, the optimizer balances exploration of uncertain regions with exploitation of configurations already known to perform well, a technique that has become a standard tool for squeezing maximum accuracy out of neural networks.</p>
<p>The evaluation was unusually broad for a single study. Quarterly time series from seven countries were used, drawn from the Federal Reserve Bank of St. Louis database for the five advanced economies and from the Haut Commissariat au Plan for Morocco and the Central Bank of Nigeria for the African cases. This mix of large, mature economies, a middle-income economy and a major emerging market gave the model a demanding test: could one framework, adapted country by country, deliver reliable forecasts across radically different data environments? The authors benchmarked their hybrid against standalone LSTM, standalone CNN, a temporal convolutional network known as TCN, and the vector autoregressive model that remains a workhorse of macroeconomic forecasting.</p>
<p>The results were striking. Across all seven countries, the hybrid CNN–LSTM model achieved coefficients of determination between 0.81 and 0.97, meaning it explained the vast majority of the variance in quarterly GDP movements. Prediction errors were correspondingly low, with mean squared error ranging from 0.00017 to 0.014077, mean absolute error from 0.0117 to 0.0960, and mean absolute percentage error between 1.18 and 7.34 percent. Crucially, the model outperformed every benchmark it was compared against, including the individual deep learning models that shared its building blocks. The advantage was not merely a statistical fluke: the authors applied the Diebold–Mariano test, the standard procedure for comparing predictive accuracy between competing forecasts, and found that the hybrid&#8217;s improvements were statistically significant in several cases, particularly against the vector autoregressive and temporal convolutional baselines.</p>
<p>Why should the combination beat its parts? The answer likely lies in complementary inductive biases. A standalone CNN sees the input as a grid of values and excels at local pattern detection, but it has limited capacity to model long-term temporal dynamics. A standalone LSTM remembers sequences well but processes multivariable inputs without an explicit mechanism for detecting local interactions between indicators. The temporal convolutional network, though efficient, similarly emphasizes local receptive fields. The hybrid model, by contrast, routes raw multivariable history through convolutional feature extraction before the recurrent layers ever see it, handing the LSTM a cleaner, more informative representation of the recent past. Bayesian optimization then ensures that the architecture is neither too small to capture the dynamics nor so large that it memorizes noise.</p>
<p>The practical implications reach beyond academic leaderboards. Accurate quarterly GDP forecasts inform central bank decisions on interest rates, government budget planning, and private sector investment strategies, and the lag between the end of a quarter and the publication of official GDP estimates creates persistent demand for reliable predictions. A model with mean absolute percentage errors as low as roughly 1 percent in some countries offers a meaningful margin over conventional approaches, and the study&#8217;s multivariable design means it can incorporate the kind of high-frequency economic indicators that policymakers already monitor. The authors note that their findings highlight the potential of hybrid deep learning architectures to enhance the accuracy of quarterly GDP forecasting in particular, suggesting a template that could be extended to other macroeconomic aggregates such as inflation, unemployment or trade balances.</p>
<p>The study also illustrates a broader trend in computational economics. Over the past decade, researchers have applied ARIMA models, artificial neural networks, support vector machines and various deep learning algorithms to GDP prediction for countries including China, India, Indonesia, Bangladesh, Egypt and New Zealand, with mixed results. What distinguishes the new work is the systematic combination of feature selection, architectural hybridization and principled hyperparameter optimization, together with a rigorous multi-country evaluation and formal statistical testing of forecast differences. As national statistical agencies and financial institutions accumulate ever larger repositories of economic indicators, the message from this research is that the path to better macroeconomic forecasts may lie not in any single model, but in intelligently combining the strengths of several, and in letting an optimizer do the fine-tuning that human analysts cannot.</p>
<p><strong>Subject of Research:</strong> Hybrid deep learning models for multivariable macroeconomic time series forecasting</p>
<p><strong>Article Title:</strong> Hybrid CNN–LSTM model optimized by Bayesian method for forecasting gross domestic product based on multivariable data</p>
<p><strong>Article References:</strong> Hamiane, S., Ghanou, Y., &amp; Casalino, G. (2026). Hybrid CNN–LSTM model optimized by Bayesian method for forecasting gross domestic product based on multivariable data. <em>International Journal of Data Science and Analytics, 22</em>(1), Article 308. <a href="https://doi.org/10.1007/s41060-026-01303-6" rel="noopener noreferrer">https://doi.org/10.1007/s41060-026-01303-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s41060-026-01303-6" rel="noopener noreferrer">10.1007/s41060-026-01303-6</a></p>
<p><strong>Keywords:</strong> GDP forecasting, CNN, LSTM, Bayesian optimization, deep learning, macroeconomic indicators, time series, econometrics, Pearson correlation, Diebold-Mariano test, multivariable data, temporal convolutional network</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">211070</post-id>	</item>
		<item>
		<title>Wavelet-Enhanced Mamba and Graph Networks Revolutionize Time Series Anomaly Detection</title>
		<link>https://scienmag.com/wavelet-enhanced-mamba-and-graph-networks-revolutionize-time-series-anomaly-detection/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Wed, 23 Sep 2026 01:46:57 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[anomaly detection]]></category>
		<category><![CDATA[cyberattack detection in IoT networks]]></category>
		<category><![CDATA[deep learning]]></category>
		<category><![CDATA[deep learning models for time series]]></category>
		<category><![CDATA[Graph neural network]]></category>
		<category><![CDATA[graph neural networks for data modeling]]></category>
		<category><![CDATA[healthcare sensor data anomaly detection]]></category>
		<category><![CDATA[Internet of Things]]></category>
		<category><![CDATA[Internet of Things sensor data analysis]]></category>
		<category><![CDATA[IoT sensors]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[Mamba]]></category>
		<category><![CDATA[Mamba sequence architecture in machine learning]]></category>
		<category><![CDATA[multi-modal data fusion in anomaly detection]]></category>
		<category><![CDATA[multivariate data]]></category>
		<category><![CDATA[real-time anomaly detection in industrial systems]]></category>
		<category><![CDATA[sequence modeling]]></category>
		<category><![CDATA[smart city traffic monitoring]]></category>
		<category><![CDATA[spacecraft telemetry anomaly detection]]></category>
		<category><![CDATA[state-space models]]></category>
		<category><![CDATA[time series]]></category>
		<category><![CDATA[time series anomaly detection]]></category>
		<category><![CDATA[wavelet analysis for signal processing]]></category>
		<category><![CDATA[wavelet transform]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=209641</guid>

					<description><![CDATA[Researchers in Beijing have unveiled a wavelet-enhanced model combining Mamba sequence modeling and graph neural networks that outperforms prevailing methods for detecting anomalies in multivariate time series from IoT systems.]]></description>
										<content:encoded><![CDATA[<p>The relentless growth of the Internet of Things has transformed ordinary environments into vast webs of interconnected sensors, each streaming measurements around the clock. From smart factories to hospital wards, from city traffic systems to orbiting spacecraft, the world now generates rivers of time series data faster than human operators can watch them. Buried within those streams are the signals that matter most: the sudden voltage spike that precedes a machine failure, the erratic heartbeat that foreshadows a cardiac event, the subtle sensor drift that hints at a cyberattack. A research team in Beijing has now introduced a machine learning model designed to catch precisely those signals, and their approach combines three of the most powerful ideas in modern artificial intelligence into a single architecture.</p>
<p>The study, led by Hanjie Xu of Zhongguancun Lab and Beihang University together with Ju Ren of Tsinghua University, Jie Liu of Beijing Information Science and Technology University, Zhenya Ma, Wang Huang of Central South University, and Jianwei Niu of Beihang University, was published in the International Journal of Machine Learning and Cybernetics. The team&#8217;s model, which fuses wavelet analysis, graph neural networks, and the newly popular Mamba sequence architecture, outperformed a wide range of established baselines in extensive benchmark testing. The work addresses one of the most persistent challenges in applied machine learning: detecting anomalies in multivariate time series, where dozens or hundreds of correlated sensor channels evolve together and a fault in one can ripple through the others in complex ways.</p>
<p>Each of the three components of the new model plays a distinct role. Wavelet representations, the first ingredient, give the network a way to see a signal at multiple resolutions simultaneously. Traditional Fourier analysis tells an engineer which frequencies exist in a signal but says little about when they occurred, a limitation known as the time frequency trade off. Wavelet transforms solve this by decomposing a signal into localized oscillations of different scales, so the model can capture both a slow drift lasting minutes and a sharp transient lasting milliseconds. By enhancing the raw inputs with these wavelet features, the architecture extracts far richer representations of each moment in the stream than models that read the raw sensor values alone.</p>
<p>The second ingredient, graph neural networks, addresses the fact that sensor data is rarely a set of independent channels. In a water treatment plant, a pressure sensor upstream is physically and causally linked to a flow sensor downstream; in a server cluster, the temperature readings of adjacent racks move in concert. Graph neural networks encode these relationships explicitly, treating each sensor as a node and the learned or measured dependencies as edges. Message passing across the graph allows the model to build a picture of how the whole system should behave as a coherent network, and to flag moments when that coherence breaks down. This topology learning capability, the authors argue, is essential for detecting anomalies that only become visible when several channels are viewed together.</p>
<p>The third component, Mamba, is one of the most talked about architectures in artificial intelligence today. Introduced by Albert Gu and Tri Dao in late 2023, Mamba is a selective state space model that processes sequences in linear time, in contrast to the quadratic cost of transformer attention. For time series workloads, where context windows can stretch across thousands of timesteps and the model must run on resource constrained edge devices, that efficiency matters enormously. Mamba&#8217;s selective mechanism lets the network decide dynamically which parts of the past to remember and which to forget, giving it a powerful ability to follow long range sequential patterns without the memory burden that attends older recurrent designs.</p>
<p>What makes the new model distinctive is the way these three ideas operate together rather than in sequence. The wavelet enhanced representations give the sequence model a frequency aware view of each channel, the graph neural network binds those views into a coherent picture of system wide behavior, and the Mamba backbone explores how the entire enriched representation evolves through time. In effect, the model asks three questions at once: how is each sensor oscillating, how are the sensors related, and how are those patterns changing? An anomaly is registered when any of these three questions produces an answer that deviates from what the model has learned to expect from normal operation.</p>
<p>The evaluation was extensive. The researchers tested their architecture against a broad collection of time series anomaly detection baselines, spanning the main families of methods that have dominated the field over the past decade. Those baselines include long short term memory networks, which have been used for everything from spacecraft telemetry to rail transit monitoring; variational autoencoders and generative adversarial networks that learn to reconstruct normal data and flag what they cannot reconstruct; transformer models such as TranAD and the Anomaly Transformer, which use attention to model dependencies across time; and earlier graph based detectors, including graph attention networks and graph deviation networks, which pioneered the topological view of multivariate data. Across this competitive landscape, the wavelet enhanced Mamba and graph network model outperformed many prevailing approaches.</p>
<p>The significance of the result extends beyond leaderboard positions. Anomaly detection sits at the foundation of industrial safety, cybersecurity, financial monitoring, and healthcare. Earlier studies in the literature the team cites show detectors deployed in industrial IoT systems, smart city traffic networks, financial markets analyzed with principal component analysis and neural networks, and clinical settings where deep learning reviews promise earlier warnings. Each of these domains has its own constraints: latency requirements that forbid heavy computation, noisy training sets, and distribution shifts that arise as machines age or seasons change. A detector that combines strong accuracy with linear time sequence processing and explicit modeling of sensor topology is therefore well matched to deployment realities, particularly at the network edge where IoT devices typically live.</p>
<p>The Mamba component deserves particular attention because it represents a broader shift in the machine learning community. Since the original Mamba paper appeared, researchers have rapidly adapted the architecture to forecasting and detection tasks, producing variants such as bidirectional Mamba for time series prediction, spatial temporal Mamba for multivariate anomaly detection, and Mamba based foundation models for general forecasting. By embedding Mamba within a graph structure and enhancing its inputs with wavelets, the Beijing team positions their model at the intersection of several active research fronts. The graph element draws on a lineage that runs from graph convolutional networks through temporal graph convolutional networks for traffic prediction to recent topological analysis methods, while the wavelet element revives a classical signal processing tool whose value modern deep learning has increasingly rediscovered.</p>
<p>The study also acknowledges the practical ecosystems in which such models must be validated. The references underpinning the work include widely used benchmarks such as the MIT BIH arrhythmia database for cardiac signals, spacecraft telemetry datasets with expert labeled anomalies, and secure water treatment testbeds used in adversarial cyberattack research. Datasets of this kind have shaped the field by providing realistic fault patterns against which methods can be calibrated, and the competitive advantage the new model demonstrated against baselines on this landscape suggests it generalizes across heterogeneous data regimes rather than being tuned to a single domain. The authors declare no conflict of interest, and the work received support reflected in the supervisory and funding roles of the senior researchers on the team.</p>
<p>Looking ahead, the research points toward anomaly detectors that are simultaneously faster, more aware of physical structure, and more sensitive to the fine grained dynamics of the systems they watch. As billions more devices come online, the volume of time series data will only grow, and the cost of missing a fault will continue to rise, whether measured in damaged equipment, lost revenue, compromised security, or human health. A model that can read a system through its frequencies, its connections, and its history at linear cost offers a template for the next generation of monitoring infrastructure. The Beijing team&#8217;s results suggest that the combination of wavelets, graphs, and selective state space models is more than a fashionable assembly of components; it is a coherent answer to the question of how machines can learn what normal looks like in an increasingly connected world, and how quickly they can notice when that normality shatters.</p>
<p><strong>Subject of Research:</strong> A wavelet-enhanced Mamba and graph neural network model for detecting anomalies in multivariate time series data from IoT systems.</p>
<p><strong>Article Title:</strong> Wavelet-enhanced Mamba and graph network-based model for time series anomaly detection</p>
<p><strong>Article References:</strong> Xu, H., Ren, J., Liu, J., Ma, Z., Huang, W., &amp; Niu, J. (2026). Wavelet-enhanced Mamba and graph network-based model for time series anomaly detection. <em>International Journal of Machine Learning and Cybernetics, 17</em>(10), Article 475. <a href="https://doi.org/10.1007/s13042-026-03250-x" rel="noopener noreferrer">https://doi.org/10.1007/s13042-026-03250-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s13042-026-03250-x" rel="noopener noreferrer">10.1007/s13042-026-03250-x</a></p>
<p><strong>Keywords:</strong> time series, anomaly detection, Mamba, graph neural network, wavelet transform, Internet of Things, machine learning, state space models, multivariate data, deep learning, IoT sensors, sequence modeling</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">209641</post-id>	</item>
		<item>
		<title>Machine Learning Meets Its Match in the Simplest Crop Price Forecast</title>
		<link>https://scienmag.com/machine-learning-meets-its-match-in-the-simplest-crop-price-forecast/</link>
		
		<dc:creator><![CDATA[Alan Morgan]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 23:55:49 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AGMARKNET]]></category>
		<category><![CDATA[agricultural commodity prices]]></category>
		<category><![CDATA[agricultural economics]]></category>
		<category><![CDATA[Agricultural price forecasting]]></category>
		<category><![CDATA[challenges in applying machine learning to agriculture]]></category>
		<category><![CDATA[designing effective agricultural forecasting models]]></category>
		<category><![CDATA[ensemble algorithms vs simple price models]]></category>
		<category><![CDATA[ensemble learning]]></category>
		<category><![CDATA[feature engineering]]></category>
		<category><![CDATA[future directions for machine learning in agricultural economics]]></category>
		<category><![CDATA[impacts of market data leakage on forecasting]]></category>
		<category><![CDATA[Indian agricultural market data analysis]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning limitations in agriculture]]></category>
		<category><![CDATA[Maharashtra]]></category>
		<category><![CDATA[mandi-level crop price data]]></category>
		<category><![CDATA[naive benchmark]]></category>
		<category><![CDATA[naive benchmark in agricultural price prediction]]></category>
		<category><![CDATA[price forecasting]]></category>
		<category><![CDATA[Random Forest]]></category>
		<category><![CDATA[role of market memory in crop price prediction]]></category>
		<category><![CDATA[short-term price persistence in commodity markets]]></category>
		<category><![CDATA[time series]]></category>
		<category><![CDATA[XGBoost]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=204236</guid>

					<description><![CDATA[A five-year study of nine commodities in Maharashtra found that ensemble machine learning models beat seasonal benchmarks but could not surpass a simple one-month naive forecast of crop prices.]]></description>
										<content:encoded><![CDATA[<p>A meticulous, five-year experiment in India has delivered one of the most refreshingly honest results in agricultural data science: sophisticated machine learning models, trained on nearly a quarter-century of market records, were unable to beat the humble forecast that next month&#8217;s price will simply equal this month&#8217;s price. The study, published in the journal Discover Informatics, examined nine major agricultural commodities in the western state of Maharashtra and found that short-run price persistence is such a powerful force in monthly market data that even the best ensemble algorithms could not outperform a one-month naive benchmark. Yet the research is far from a defeat for machine learning. Instead, it offers a rigorous, leakage-free template for how forecasting studies should be designed, and it pinpoints precisely where future gains must come from.</p>
<p>The research team, led by Payal Mahajan and Aniket A. Muley of Swami Ramanand Teerth Marathwada University in Nanded, together with Madhav R. Fegade of Digambarrao Bindu Arts, Commerce and Science College in Bhokar, built their framework on daily mandi-level price records from the AGMARKNET portal, the official agricultural marketing database of the Government of India. The commodities spanned the breadth of Indian farming: Arhar, Bajra, Cotton, Maize, Paddy, Sesamum, Soyabean, Sunflower and Wheat. Records dating from 2001 through October 2025 were cleaned, validated and aggregated into monthly average modal prices, the most commonly observed market price for each commodity, producing long monthly price series that capture decades of market behaviour.</p>
<p>One of the study&#8217;s central contributions is methodological rather than predictive. In many published forecasting papers, preprocessing steps such as outlier removal are performed on the entire dataset before the model is tested, which silently leaks information from the future into the training process and inflates apparent accuracy. The Maharashtra team avoided this trap entirely. Extreme prices were capped at commodity-specific percentiles, but the quantile limits were estimated only from training data available at each sequential forecasting step and then applied to the following test month. In other words, the pipeline was designed so that no future observation could ever influence a past forecast, a discipline that many machine learning benchmarks in economics and agriculture still lack.</p>
<p>The feature engineering pipeline was deliberately rich. Lag features captured prices from one, two, three, six, nine, twelve, eighteen and twenty-four months earlier, allowing the models to weigh both recent momentum and annual cycles. Rolling means over three, six, twelve and twenty-four month windows summarised short-, medium- and long-term trends, while rolling standard deviations of the same lengths encoded volatility, telling the models whether the market had been calm or turbulent in the preceding months. Seasonal calendar variables were encoded with sine and cosine transformations of the month number, a standard technique that preserves the cyclical nature of the year so that December and January are treated as neighbours rather than as numerically distant values. All prices were log-transformed to stabilise variance and dampen the influence of sudden spikes.</p>
<p>Four tree-based ensemble algorithms were put to the test: Random Forest, which aggregates decision trees trained on bootstrapped samples; Extra Trees, which injects additional randomisation into split selection for robustness; Histogram Gradient Boosting, a fast boosting method optimised for structured numerical data; and XGBoost, a regularised gradient boosting system that sequentially corrects the errors of weak learners. Models were trained with fixed hyperparameters and evaluated with Mean Absolute Error, Root Mean Squared Error, the coefficient of determination and Mean Absolute Percentage Error, using an expanding-window protocol in which each month from January 2021 to October 2025 was forecast using everything known before it, with the training set growing as actual prices became available.</p>
<p>The headline result is stark. The one-month naive forecast, which simply predicts that the coming month&#8217;s price equals the previous month&#8217;s, achieved the lowest MAPE for every single commodity, ranging from just 2.30 percent for Paddy to 4.91 percent for Sunflower. The best machine learning models, however, consistently outperformed the twelve-month seasonal naive benchmark, which predicts that this month&#8217;s price will match the same calendar month one year earlier. This distinction matters. It shows that the engineered lag, rolling and seasonal features do capture genuine nonlinear temporal structure beyond annual repetition, but not enough extra information to improve on the sheer stubbornness of short-run price continuation in aggregated monthly data.</p>
<p>Why is the naive benchmark so formidable? Monthly agricultural prices at the state level adjust gradually. Supply arrives in bulk through regulated markets, demand shifts slowly, and administrative factors such as minimum support prices dampen sudden swings. In such a regime, the previous month&#8217;s price already embeds most of the information needed for a one-step-ahead forecast, leaving little room for algorithms to extract additional signal from historical prices alone. The study&#8217;s authors note that abrupt movements driven by rainfall deficits, market arrivals, input costs, policy interventions or export-import shocks simply cannot be anticipated from the price series itself, which caps what any historical-price-only model can achieve.</p>
<p>The experiments also revealed meaningful differences among commodities and validation strategies. Switching from fixed-window training to expanding-window retraining, in which models were periodically refreshed with newly observed prices, reduced forecast errors for all nine commodities, confirming that regular updating is worthwhile even when it does not overturn the naive benchmark. Paddy and Wheat proved relatively easy to forecast, reflecting their lower relative price variability, while Arhar, Cotton, Sesamum, Soyabean and Sunflower, commodities with higher volatility and more erratic price behaviour, remained challenging. Sensitivity analyses with different training start years showed that no single configuration was uniformly best, underscoring that commodity-specific market dynamics, rather than any universal algorithmic choice, shape forecastability.</p>
<p>The researchers are transparent about the limitations of their design. Exogenous drivers such as rainfall, temperature, market arrivals, fuel and transport costs, inflation, festival demand and government policy were deliberately excluded, making the study a controlled assessment of what historical prices alone can deliver. Deep learning architectures such as LSTM and GRU networks were also left out, since a fair comparison would require a separate chronological tuning and architecture-selection study under the same leakage-safe principles. The authors frame both omissions as clear directions for future work, alongside probabilistic forecasting that would quantify uncertainty rather than produce single point estimates.</p>
<p>For farmers, traders and policymakers, the practical message is double-edged but valuable. First, for month-ahead planning, the cheapest forecast on the table, carrying last month&#8217;s price forward, is remarkably hard to beat, and any proposed machine learning service should be benchmarked against it before being trusted. Second, meaningful improvements in agricultural price forecasting will not come from more elaborate algorithms chewing on the same price history, but from richer information: weather data, crop arrivals, production estimates, cost indices and policy signals integrated into leakage-safe evaluation frameworks like the one this study demonstrates. In an era when artificial intelligence is often oversold, this Maharashtra experiment stands out as a model of scientific candour, showing exactly where machine learning helps, where it does not, and how the next generation of forecasting systems should be built.</p>
<p><strong>Subject of Research:</strong> Machine learning models for forecasting monthly agricultural commodity prices in Maharashtra, India</p>
<p><strong>Article Title:</strong> Forecasting agricultural commodity prices using machine learning models</p>
<p><strong>Article References:</strong> Mahajan, P., Muley, A. A., &amp; Fegade, M. R. (2026). Forecasting agricultural commodity prices using machine learning models. <em>Discover Informatics, 1</em>(1), Article 15. <a href="https://doi.org/10.1007/s44564-026-00015-0" rel="noopener noreferrer">https://doi.org/10.1007/s44564-026-00015-0</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44564-026-00015-0" rel="noopener noreferrer">10.1007/s44564-026-00015-0</a></p>
<p><strong>Keywords:</strong> agricultural commodity prices, machine learning, price forecasting, time series, Random Forest, XGBoost, AGMARKNET, Maharashtra, feature engineering, naive benchmark, ensemble learning, agricultural economics</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">204236</post-id>	</item>
		<item>
		<title>Correlation-Based Clustering Reveals Hidden Patterns in Citrus Price Dynamics</title>
		<link>https://scienmag.com/correlation-based-clustering-reveals-hidden-patterns-in-citrus-price-dynamics/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:36:59 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[agricultural economics]]></category>
		<category><![CDATA[agricultural market segmentation]]></category>
		<category><![CDATA[Citrus price dynamics]]></category>
		<category><![CDATA[citrus prices]]></category>
		<category><![CDATA[clustering]]></category>
		<category><![CDATA[Comunitat Valenciana]]></category>
		<category><![CDATA[correlation distance]]></category>
		<category><![CDATA[correlation-based clustering in market analysis]]></category>
		<category><![CDATA[economic pattern classification of crop prices]]></category>
		<category><![CDATA[feature engineering]]></category>
		<category><![CDATA[feature-based clustering for crop prices]]></category>
		<category><![CDATA[fruit price fluctuation analysis]]></category>
		<category><![CDATA[hidden patterns in agricultural markets]]></category>
		<category><![CDATA[innovative tools for agricultural economics]]></category>
		<category><![CDATA[k-medoids]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[market behavior analysis for policymakers]]></category>
		<category><![CDATA[market monitoring]]></category>
		<category><![CDATA[price volatility]]></category>
		<category><![CDATA[regional citrus export market study]]></category>
		<category><![CDATA[time series]]></category>
		<category><![CDATA[unsupervised learning]]></category>
		<category><![CDATA[unsupervised machine learning in agriculture]]></category>
		<category><![CDATA[volatility monitoring in citrus farming]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201820</guid>

					<description><![CDATA[Spanish researchers have developed a correlation-based k-medoids clustering method that classifies citrus price dynamics into three stable, economically meaningful patterns across nine seasons.]]></description>
										<content:encoded><![CDATA[<p>A team of researchers in Spain has developed a new unsupervised machine learning framework that classifies the price behavior of citrus varieties into three economically meaningful patterns, offering farmers, cooperatives and policymakers a simple yet robust tool for monitoring volatile agricultural markets. The study, published in Machine Learning with Applications, analyzes more than 6,700 weekly farm-gate price records collected across the three provinces of the Comunitat Valenciana, the region that produces roughly half of Spain&#8217;s citrus and anchors the country&#8217;s position as the world&#8217;s leading exporter of citrus for fresh consumption.</p>
<p>The work, led by Roger Arnau, Jose M. Calabuig, Nuria Ortigosa and Luiza Petrosyan, addresses a stubborn problem in agricultural economics: price series for different crop varieties cover seasons of wildly different lengths, start at different times of year, and fluctuate on very different scales. Comparing such series directly with conventional clustering tools, which typically rely on Euclidean distance between raw data points, tends to group varieties that merely share similar absolute price levels while ignoring whether their prices rise, fall or oscillate in the same way. The researchers&#8217; solution is to abandon raw prices altogether and instead describe each variety-province-season combination with a compact set of five interpretable features: the duration of the marketing season, the variance of prices, the slope of a linear trend fitted to the weekly prices, the R-squared quality of that fit, and a novel Q-ratio that captures price amplitude per week of season.</p>
<p>Once each observation is encoded in this feature space, the team applies k-medoids clustering, also known as Partitioning Around Medoids, using correlation dissimilarity rather than Euclidean or Manhattan distance. Unlike k-means, which represents each cluster with a mean that may not correspond to any real data point, k-medoids selects actual observations as cluster representatives, making the method less sensitive to outliers and fully deterministic, requiring no random seed. The correlation distance, defined as one minus the Pearson correlation between feature vectors, considers two observations similar if their variables vary proportionally, even when their absolute magnitudes differ. In practice, this means two citrus varieties are grouped together when their prices rise and fall at the same times, regardless of whether one sells at twice the price of the other.</p>
<p>Determining the right number of clusters, a classic hyperparameter challenge in unsupervised learning, was handled with multiple lines of evidence. The Silhouette method and the Elbow method both pointed to three clusters, and subsample consensus analysis, in which the clustering was repeated 500 times on random 80 percent subsets of seasons, produced its lowest proportion of ambiguous clustering, about six percent, at exactly that value. Bootstrap resampling with 1,000 repetitions yielded 95 percent confidence intervals for the internal validation indices, while the GAP statistic was treated as non-diagnostic because it favored a single cluster under the 1-SE rule. The convergence of the other criteria on three groups gave the researchers confidence that the structure was genuine rather than an artifact of a single algorithm run.</p>
<p>The choice of distance metric proved decisive. When the correlation-based k-medoids was compared against Euclidean and Manhattan alternatives using standard internal validation indices, the correlation approach dominated: it achieved a Dunn2 index of 1.60 versus 0.62 for Euclidean and 0.56 for Manhattan, and a Calinski-Harabasz score of 393 versus 158 and 204 respectively. An ablation study further showed that neither the five-feature representation nor the correlation distance alone explains the improvement; the gain emerges from their synergy. When raw weekly prices were used instead of features, even dynamic time warping, a sophisticated technique for aligning time series of unequal length, failed to match the combined approach, partly because very short seasons of five or six weeks produce pathological alignments.</p>
<p>The three resulting clusters translate directly into market narratives. The first group, dominated by mandarins and early clementines, is characterized by short seasons, high price volatility and a clear downward drift, with prices falling by a median of 1.4 euro cents per week and a strong linear trend. The second group, populated largely by orange varieties with long marketing windows and prices often agreed in advance, shows remarkable stability: near-zero trend slopes, the lowest price variance and the lowest Q-ratio. The third group is the most erratic, with a median R-squared of only 0.167, indicating that prices swing up and down in ways no linear model can capture, a signature of external shocks such as weather disruptions or sudden demand shifts.</p>
<p>Crucially, the cluster assignments were validated against an independent external criterion derived from official Valencian agricultural sector reports, in which each season was labeled up, flat or down based on weekly price changes. The agreement was moderate but statistically significant, with accuracy of 0.52, a Macro-F1 of 0.52, Cohen&#8217;s kappa of 0.28 and a permutation p-value of 0.001, and most disagreements occurred between the flat cluster and its adjacent neighbors, exactly where borderline seasons would be expected. The clustering structure also aligned with documented market events: the 2018-2019 season of overproduction and delayed harvesting pushed most mandarins into the declining cluster, torrential rains in late 2016 preceded a sharp price drop for Navelina oranges, the COVID-19 pandemic in 2020 produced atypical fluctuating patterns as demand for vitamin C surged, and the farmer protests that blocked roads in Castellón in early 2024 drove varieties such as Ortanique and Clemenvilla into the declining cluster in that province while they remained stable in Valencia.</p>
<p>Perhaps the most sobering finding is the sheer instability of cluster membership over time. On average, about 63 percent of variety-province pairs shifted clusters between consecutive seasons, peaking at 70 percent in 2018-2019. Every one of the 111 varieties tracked across all seasons changed groups at least once, which the authors interpret as evidence that agricultural commodity prices are subject to sharp, recurring fluctuations driven by weather, logistics, demand shocks and policy events. This volatility complicates long-term forecasting based on historical prices alone, but it also underscores the value of a monitoring tool that can flag, in near real time, when a variety&#8217;s behavior departs from its usual regime.</p>
<p>The researchers emphasize that the framework is deliberately simple and portable. Five features suffice, the algorithm is deterministic, the processed dataset of 528 observations has been made publicly available for reproduction, and the same pipeline could be transferred to other crops, regions or time periods with modest domain-specific tuning, such as defining season boundaries or selecting derived variables. The main limitations are that the model uses only prices at origin, without weather, trade or production-volume data, and that publicly available prices are aggregated by province, limiting microeconomic resolution. Even so, the fact that the clusters independently recovered the fingerprints of floods, a pandemic and road blockades suggests that a shape-based view of price dynamics, built on correlation rather than magnitude, can extract real economic signal from nothing more than weekly price records, providing a reproducible template for market surveillance across the agri-food sector.</p>
<p><strong>Subject of Research:</strong> Correlation-based k-medoids clustering of weekly citrus price dynamics in the Comunitat Valenciana, Spain</p>
<p><strong>Article Title:</strong> Enhancing unsupervised learning with correlation-based k -medoids: A case study on citrus price dynamics</p>
<p><strong>Article References:</strong> Arnau, R., Calabuig, J. M., Ortigosa, N., &amp; Petrosyan, L. (2026). Enhancing unsupervised learning with correlation-based k-medoids: A case study on citrus price dynamics. <em>Machine Learning with Applications, 26</em>, Article 100988. <a href="https://doi.org/10.1016/j.mlwa.2026.100988" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.100988</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.100988" rel="noopener noreferrer">10.1016/j.mlwa.2026.100988</a></p>
<p><strong>Keywords:</strong> unsupervised learning, k-medoids, correlation distance, citrus prices, agricultural economics, clustering, time series, machine learning, Comunitat Valenciana, price volatility, feature engineering, market monitoring</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201820</post-id>	</item>
	</channel>
</rss>
