A team of researchers in Spain has developed a new unsupervised machine learning framework that classifies the price behavior of citrus varieties into three economically meaningful patterns, offering farmers, cooperatives and policymakers a simple yet robust tool for monitoring volatile agricultural markets. The study, published in Machine Learning with Applications, analyzes more than 6,700 weekly farm-gate price records collected across the three provinces of the Comunitat Valenciana, the region that produces roughly half of Spain’s citrus and anchors the country’s position as the world’s leading exporter of citrus for fresh consumption.
The work, led by Roger Arnau, Jose M. Calabuig, Nuria Ortigosa and Luiza Petrosyan, addresses a stubborn problem in agricultural economics: price series for different crop varieties cover seasons of wildly different lengths, start at different times of year, and fluctuate on very different scales. Comparing such series directly with conventional clustering tools, which typically rely on Euclidean distance between raw data points, tends to group varieties that merely share similar absolute price levels while ignoring whether their prices rise, fall or oscillate in the same way. The researchers’ solution is to abandon raw prices altogether and instead describe each variety-province-season combination with a compact set of five interpretable features: the duration of the marketing season, the variance of prices, the slope of a linear trend fitted to the weekly prices, the R-squared quality of that fit, and a novel Q-ratio that captures price amplitude per week of season.
Once each observation is encoded in this feature space, the team applies k-medoids clustering, also known as Partitioning Around Medoids, using correlation dissimilarity rather than Euclidean or Manhattan distance. Unlike k-means, which represents each cluster with a mean that may not correspond to any real data point, k-medoids selects actual observations as cluster representatives, making the method less sensitive to outliers and fully deterministic, requiring no random seed. The correlation distance, defined as one minus the Pearson correlation between feature vectors, considers two observations similar if their variables vary proportionally, even when their absolute magnitudes differ. In practice, this means two citrus varieties are grouped together when their prices rise and fall at the same times, regardless of whether one sells at twice the price of the other.
Determining the right number of clusters, a classic hyperparameter challenge in unsupervised learning, was handled with multiple lines of evidence. The Silhouette method and the Elbow method both pointed to three clusters, and subsample consensus analysis, in which the clustering was repeated 500 times on random 80 percent subsets of seasons, produced its lowest proportion of ambiguous clustering, about six percent, at exactly that value. Bootstrap resampling with 1,000 repetitions yielded 95 percent confidence intervals for the internal validation indices, while the GAP statistic was treated as non-diagnostic because it favored a single cluster under the 1-SE rule. The convergence of the other criteria on three groups gave the researchers confidence that the structure was genuine rather than an artifact of a single algorithm run.
The choice of distance metric proved decisive. When the correlation-based k-medoids was compared against Euclidean and Manhattan alternatives using standard internal validation indices, the correlation approach dominated: it achieved a Dunn2 index of 1.60 versus 0.62 for Euclidean and 0.56 for Manhattan, and a Calinski-Harabasz score of 393 versus 158 and 204 respectively. An ablation study further showed that neither the five-feature representation nor the correlation distance alone explains the improvement; the gain emerges from their synergy. When raw weekly prices were used instead of features, even dynamic time warping, a sophisticated technique for aligning time series of unequal length, failed to match the combined approach, partly because very short seasons of five or six weeks produce pathological alignments.
The three resulting clusters translate directly into market narratives. The first group, dominated by mandarins and early clementines, is characterized by short seasons, high price volatility and a clear downward drift, with prices falling by a median of 1.4 euro cents per week and a strong linear trend. The second group, populated largely by orange varieties with long marketing windows and prices often agreed in advance, shows remarkable stability: near-zero trend slopes, the lowest price variance and the lowest Q-ratio. The third group is the most erratic, with a median R-squared of only 0.167, indicating that prices swing up and down in ways no linear model can capture, a signature of external shocks such as weather disruptions or sudden demand shifts.
Crucially, the cluster assignments were validated against an independent external criterion derived from official Valencian agricultural sector reports, in which each season was labeled up, flat or down based on weekly price changes. The agreement was moderate but statistically significant, with accuracy of 0.52, a Macro-F1 of 0.52, Cohen’s kappa of 0.28 and a permutation p-value of 0.001, and most disagreements occurred between the flat cluster and its adjacent neighbors, exactly where borderline seasons would be expected. The clustering structure also aligned with documented market events: the 2018-2019 season of overproduction and delayed harvesting pushed most mandarins into the declining cluster, torrential rains in late 2016 preceded a sharp price drop for Navelina oranges, the COVID-19 pandemic in 2020 produced atypical fluctuating patterns as demand for vitamin C surged, and the farmer protests that blocked roads in Castellón in early 2024 drove varieties such as Ortanique and Clemenvilla into the declining cluster in that province while they remained stable in Valencia.
Perhaps the most sobering finding is the sheer instability of cluster membership over time. On average, about 63 percent of variety-province pairs shifted clusters between consecutive seasons, peaking at 70 percent in 2018-2019. Every one of the 111 varieties tracked across all seasons changed groups at least once, which the authors interpret as evidence that agricultural commodity prices are subject to sharp, recurring fluctuations driven by weather, logistics, demand shocks and policy events. This volatility complicates long-term forecasting based on historical prices alone, but it also underscores the value of a monitoring tool that can flag, in near real time, when a variety’s behavior departs from its usual regime.
The researchers emphasize that the framework is deliberately simple and portable. Five features suffice, the algorithm is deterministic, the processed dataset of 528 observations has been made publicly available for reproduction, and the same pipeline could be transferred to other crops, regions or time periods with modest domain-specific tuning, such as defining season boundaries or selecting derived variables. The main limitations are that the model uses only prices at origin, without weather, trade or production-volume data, and that publicly available prices are aggregated by province, limiting microeconomic resolution. Even so, the fact that the clusters independently recovered the fingerprints of floods, a pandemic and road blockades suggests that a shape-based view of price dynamics, built on correlation rather than magnitude, can extract real economic signal from nothing more than weekly price records, providing a reproducible template for market surveillance across the agri-food sector.
Subject of Research: Correlation-based k-medoids clustering of weekly citrus price dynamics in the Comunitat Valenciana, Spain
Article Title: Enhancing unsupervised learning with correlation-based k -medoids: A case study on citrus price dynamics
Article References: Arnau, R., Calabuig, J. M., Ortigosa, N., & Petrosyan, L. (2026). Enhancing unsupervised learning with correlation-based k-medoids: A case study on citrus price dynamics. Machine Learning with Applications, 26, Article 100988. https://doi.org/10.1016/j.mlwa.2026.100988
Image Credits: AI Generated
DOI: 10.1016/j.mlwa.2026.100988
Keywords: unsupervised learning, k-medoids, correlation distance, citrus prices, agricultural economics, clustering, time series, machine learning, Comunitat Valenciana, price volatility, feature engineering, market monitoring
Cite Scienmag News
Denise Maddox. (September 20, 2026). Correlation-Based Clustering Reveals Hidden Patterns in Citrus Price Dynamics. Scienmag. https://scienmag.com/correlation-based-clustering-reveals-hidden-patterns-in-citrus-price-dynamics/
Denise Maddox. "Correlation-Based Clustering Reveals Hidden Patterns in Citrus Price Dynamics." Scienmag, 20 September 2026, https://scienmag.com/correlation-based-clustering-reveals-hidden-patterns-in-citrus-price-dynamics/. Accessed 20 September 2026.
Denise Maddox. "Correlation-Based Clustering Reveals Hidden Patterns in Citrus Price Dynamics." Scienmag. September 20, 2026. https://scienmag.com/correlation-based-clustering-reveals-hidden-patterns-in-citrus-price-dynamics/

