<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>k-medoids &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/k-medoids/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 20 Sep 2026 19:36:59 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>k-medoids &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Correlation-Based Clustering Reveals Hidden Patterns in Citrus Price Dynamics</title>
		<link>https://scienmag.com/correlation-based-clustering-reveals-hidden-patterns-in-citrus-price-dynamics/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:36:59 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[agricultural economics]]></category>
		<category><![CDATA[agricultural market segmentation]]></category>
		<category><![CDATA[Citrus price dynamics]]></category>
		<category><![CDATA[citrus prices]]></category>
		<category><![CDATA[clustering]]></category>
		<category><![CDATA[Comunitat Valenciana]]></category>
		<category><![CDATA[correlation distance]]></category>
		<category><![CDATA[correlation-based clustering in market analysis]]></category>
		<category><![CDATA[economic pattern classification of crop prices]]></category>
		<category><![CDATA[feature engineering]]></category>
		<category><![CDATA[feature-based clustering for crop prices]]></category>
		<category><![CDATA[fruit price fluctuation analysis]]></category>
		<category><![CDATA[hidden patterns in agricultural markets]]></category>
		<category><![CDATA[innovative tools for agricultural economics]]></category>
		<category><![CDATA[k-medoids]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[market behavior analysis for policymakers]]></category>
		<category><![CDATA[market monitoring]]></category>
		<category><![CDATA[price volatility]]></category>
		<category><![CDATA[regional citrus export market study]]></category>
		<category><![CDATA[time series]]></category>
		<category><![CDATA[unsupervised learning]]></category>
		<category><![CDATA[unsupervised machine learning in agriculture]]></category>
		<category><![CDATA[volatility monitoring in citrus farming]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201820</guid>

					<description><![CDATA[Spanish researchers have developed a correlation-based k-medoids clustering method that classifies citrus price dynamics into three stable, economically meaningful patterns across nine seasons.]]></description>
										<content:encoded><![CDATA[<p>A team of researchers in Spain has developed a new unsupervised machine learning framework that classifies the price behavior of citrus varieties into three economically meaningful patterns, offering farmers, cooperatives and policymakers a simple yet robust tool for monitoring volatile agricultural markets. The study, published in Machine Learning with Applications, analyzes more than 6,700 weekly farm-gate price records collected across the three provinces of the Comunitat Valenciana, the region that produces roughly half of Spain&#8217;s citrus and anchors the country&#8217;s position as the world&#8217;s leading exporter of citrus for fresh consumption.</p>
<p>The work, led by Roger Arnau, Jose M. Calabuig, Nuria Ortigosa and Luiza Petrosyan, addresses a stubborn problem in agricultural economics: price series for different crop varieties cover seasons of wildly different lengths, start at different times of year, and fluctuate on very different scales. Comparing such series directly with conventional clustering tools, which typically rely on Euclidean distance between raw data points, tends to group varieties that merely share similar absolute price levels while ignoring whether their prices rise, fall or oscillate in the same way. The researchers&#8217; solution is to abandon raw prices altogether and instead describe each variety-province-season combination with a compact set of five interpretable features: the duration of the marketing season, the variance of prices, the slope of a linear trend fitted to the weekly prices, the R-squared quality of that fit, and a novel Q-ratio that captures price amplitude per week of season.</p>
<p>Once each observation is encoded in this feature space, the team applies k-medoids clustering, also known as Partitioning Around Medoids, using correlation dissimilarity rather than Euclidean or Manhattan distance. Unlike k-means, which represents each cluster with a mean that may not correspond to any real data point, k-medoids selects actual observations as cluster representatives, making the method less sensitive to outliers and fully deterministic, requiring no random seed. The correlation distance, defined as one minus the Pearson correlation between feature vectors, considers two observations similar if their variables vary proportionally, even when their absolute magnitudes differ. In practice, this means two citrus varieties are grouped together when their prices rise and fall at the same times, regardless of whether one sells at twice the price of the other.</p>
<p>Determining the right number of clusters, a classic hyperparameter challenge in unsupervised learning, was handled with multiple lines of evidence. The Silhouette method and the Elbow method both pointed to three clusters, and subsample consensus analysis, in which the clustering was repeated 500 times on random 80 percent subsets of seasons, produced its lowest proportion of ambiguous clustering, about six percent, at exactly that value. Bootstrap resampling with 1,000 repetitions yielded 95 percent confidence intervals for the internal validation indices, while the GAP statistic was treated as non-diagnostic because it favored a single cluster under the 1-SE rule. The convergence of the other criteria on three groups gave the researchers confidence that the structure was genuine rather than an artifact of a single algorithm run.</p>
<p>The choice of distance metric proved decisive. When the correlation-based k-medoids was compared against Euclidean and Manhattan alternatives using standard internal validation indices, the correlation approach dominated: it achieved a Dunn2 index of 1.60 versus 0.62 for Euclidean and 0.56 for Manhattan, and a Calinski-Harabasz score of 393 versus 158 and 204 respectively. An ablation study further showed that neither the five-feature representation nor the correlation distance alone explains the improvement; the gain emerges from their synergy. When raw weekly prices were used instead of features, even dynamic time warping, a sophisticated technique for aligning time series of unequal length, failed to match the combined approach, partly because very short seasons of five or six weeks produce pathological alignments.</p>
<p>The three resulting clusters translate directly into market narratives. The first group, dominated by mandarins and early clementines, is characterized by short seasons, high price volatility and a clear downward drift, with prices falling by a median of 1.4 euro cents per week and a strong linear trend. The second group, populated largely by orange varieties with long marketing windows and prices often agreed in advance, shows remarkable stability: near-zero trend slopes, the lowest price variance and the lowest Q-ratio. The third group is the most erratic, with a median R-squared of only 0.167, indicating that prices swing up and down in ways no linear model can capture, a signature of external shocks such as weather disruptions or sudden demand shifts.</p>
<p>Crucially, the cluster assignments were validated against an independent external criterion derived from official Valencian agricultural sector reports, in which each season was labeled up, flat or down based on weekly price changes. The agreement was moderate but statistically significant, with accuracy of 0.52, a Macro-F1 of 0.52, Cohen&#8217;s kappa of 0.28 and a permutation p-value of 0.001, and most disagreements occurred between the flat cluster and its adjacent neighbors, exactly where borderline seasons would be expected. The clustering structure also aligned with documented market events: the 2018-2019 season of overproduction and delayed harvesting pushed most mandarins into the declining cluster, torrential rains in late 2016 preceded a sharp price drop for Navelina oranges, the COVID-19 pandemic in 2020 produced atypical fluctuating patterns as demand for vitamin C surged, and the farmer protests that blocked roads in Castellón in early 2024 drove varieties such as Ortanique and Clemenvilla into the declining cluster in that province while they remained stable in Valencia.</p>
<p>Perhaps the most sobering finding is the sheer instability of cluster membership over time. On average, about 63 percent of variety-province pairs shifted clusters between consecutive seasons, peaking at 70 percent in 2018-2019. Every one of the 111 varieties tracked across all seasons changed groups at least once, which the authors interpret as evidence that agricultural commodity prices are subject to sharp, recurring fluctuations driven by weather, logistics, demand shocks and policy events. This volatility complicates long-term forecasting based on historical prices alone, but it also underscores the value of a monitoring tool that can flag, in near real time, when a variety&#8217;s behavior departs from its usual regime.</p>
<p>The researchers emphasize that the framework is deliberately simple and portable. Five features suffice, the algorithm is deterministic, the processed dataset of 528 observations has been made publicly available for reproduction, and the same pipeline could be transferred to other crops, regions or time periods with modest domain-specific tuning, such as defining season boundaries or selecting derived variables. The main limitations are that the model uses only prices at origin, without weather, trade or production-volume data, and that publicly available prices are aggregated by province, limiting microeconomic resolution. Even so, the fact that the clusters independently recovered the fingerprints of floods, a pandemic and road blockades suggests that a shape-based view of price dynamics, built on correlation rather than magnitude, can extract real economic signal from nothing more than weekly price records, providing a reproducible template for market surveillance across the agri-food sector.</p>
<p><strong>Subject of Research:</strong> Correlation-based k-medoids clustering of weekly citrus price dynamics in the Comunitat Valenciana, Spain</p>
<p><strong>Article Title:</strong> Enhancing unsupervised learning with correlation-based k -medoids: A case study on citrus price dynamics</p>
<p><strong>Article References:</strong> Arnau, R., Calabuig, J. M., Ortigosa, N., &amp; Petrosyan, L. (2026). Enhancing unsupervised learning with correlation-based k-medoids: A case study on citrus price dynamics. <em>Machine Learning with Applications, 26</em>, Article 100988. <a href="https://doi.org/10.1016/j.mlwa.2026.100988" rel="noopener noreferrer">https://doi.org/10.1016/j.mlwa.2026.100988</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.mlwa.2026.100988" rel="noopener noreferrer">10.1016/j.mlwa.2026.100988</a></p>
<p><strong>Keywords:</strong> unsupervised learning, k-medoids, correlation distance, citrus prices, agricultural economics, clustering, time series, machine learning, Comunitat Valenciana, price volatility, feature engineering, market monitoring</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201820</post-id>	</item>
		<item>
		<title>AI Clustering of Raw Eye-Tracking Data Reveals How Young Drivers Scan the Road</title>
		<link>https://scienmag.com/ai-clustering-of-raw-eye-tracking-data-reveals-how-young-drivers-scan-the-road/</link>
		
		<dc:creator><![CDATA[Blake Davidson]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 19:30:31 +0000</pubDate>
				<category><![CDATA[Psychology & Psychiatry]]></category>
		<category><![CDATA[AI-powered road safety assessment]]></category>
		<category><![CDATA[automated analysis of eye movement patterns]]></category>
		<category><![CDATA[autonomous clustering of eye movement data]]></category>
		<category><![CDATA[Behavior Research Methods]]></category>
		<category><![CDATA[driver safety]]></category>
		<category><![CDATA[driving simulation]]></category>
		<category><![CDATA[driving simulation eye-tracking analysis]]></category>
		<category><![CDATA[dynamic time warping]]></category>
		<category><![CDATA[eye tracking]]></category>
		<category><![CDATA[eye-tracking data analysis]]></category>
		<category><![CDATA[hazard perception]]></category>
		<category><![CDATA[k-medoids]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in driving behavior]]></category>
		<category><![CDATA[neural network applications in driver behavior studies]]></category>
		<category><![CDATA[raw eye-tracking data in driver research]]></category>
		<category><![CDATA[situational awareness]]></category>
		<category><![CDATA[situational awareness measurement in driving]]></category>
		<category><![CDATA[time-series clustering]]></category>
		<category><![CDATA[time-series clustering for driver behavior]]></category>
		<category><![CDATA[unsupervised machine learning for visual scanning]]></category>
		<category><![CDATA[visual search strategies in young drivers]]></category>
		<category><![CDATA[visual search strategy]]></category>
		<category><![CDATA[young drivers]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197912</guid>

					<description><![CDATA[Researchers have developed a scalable machine learning method that uses time-series clustering to automatically identify visual search strategies from raw eye-tracking data collected from young drivers in a simulated driving assessment.]]></description>
										<content:encoded><![CDATA[<p>Every time a driver glances toward a hidden crosswalk or sweeps the road ahead of a curve, the eyes are carrying out a strategy that may determine whether a crash happens or not. Researchers have long known that this visual search behavior, often called a visual search strategy, underpins situational awareness, the capacity to understand the surrounding environment and anticipate what will happen next. But measuring it has been painstakingly slow, limiting studies to small groups of participants and short stretches of driving. A new study published in Behavior Research Methods offers a way out of that bottleneck by letting machine learning algorithms do the heavy lifting on raw eye-tracking data.</p>
<p>The research, led by Thomas Seacrist, Elizabeth E. Walshe, David Grethlein, Megan S. Ryerson, Flaura K. Winston and colleagues at the Children&#8217;s Hospital of Philadelphia and partner institutions, demonstrates that time-series clustering, an unsupervised machine learning technique, can automatically identify distinct visual search strategies from eye-tracking recordings collected during simulated driving. Instead of requiring human coders to watch hours of video and label where each driver looked, the approach compares entire streams of gaze data directly and groups drivers whose scanning patterns resemble one another.</p>
<p>The significance of this shift is hard to overstate for the field. Traditionally, characterizing a visual search strategy has meant video coding: trained analysts review synchronized footage of the road scene and the driver&#8217;s eyes, mark fixations on areas of interest, and translate those marks into summary measures. The process is labor-intensive, subjective at the margins, and effectively caps the size of datasets a lab can analyze. That cap matters because the consequences of insufficient visual search are severe. A driver who fails to scan far enough ahead may miss a pedestrian stepping out from behind an obstructed view, or may enter a blind curve without the anticipatory glances that experienced drivers deploy.</p>
<p>To test their method, the team collected eye-tracking data from 36 young drivers aged 16 to 24 as they completed a virtual driving assessment, a validated simulated driving test used to gauge performance behind the wheel. The cohort is deliberately focused on a high-risk population. Crash rates among newly licensed and teenage drivers remain dramatically elevated compared with older, experienced motorists, and errors in scanning and hazard anticipation figure prominently among the mistakes that precede serious crashes involving young novices.</p>
<p>At the heart of the method is a measure called dynamic localized coordinate aligned warping, abbreviated DCLAW, an extension of the well-known dynamic time warping algorithm. Dynamic time warping, originally developed for aligning spoken word recordings, allows two time series that unfold at different speeds to be compared by stretching and compressing them along the time axis. DCLAW adapts this idea to eye-tracking data, which arrive as rapidly sampled coordinates of gaze position and rarely align neatly between one driver and another. Two drivers might sweep their eyes across the same sequence of road regions but at slightly different moments or paces; a naive comparison would call them dissimilar, while DCLAW can recognize the underlying strategy as essentially the same.</p>
<p>Once pairwise similarities between all drivers&#8217; raw gaze streams had been computed, the researchers applied k-medoids clustering, an unsupervised algorithm that partitions data into groups organized around actual representative examples called medoids. Because the method is unsupervised, it requires no preconceived categories of good or bad scanning behavior. The number and structure of the clusters emerge from the data itself, and the researchers then examined each cluster&#8217;s medoid, the driver whose time series sits at the center of the group, to characterize what defined that particular visual search strategy.</p>
<p>The results showed that time-series clustering successfully identified generalizable visual search strategies during a curved roadway scenario, one of the more demanding situations in driving. Curves demand anticipatory glances toward the tangent of the bend and disciplined checking of the lane ahead, and differences in how young drivers allocate attention there are linked to crash risk. The fact that clustering raw data recovered meaningful, generalizable strategy groups in this setting suggests the technique can capture behaviorally important variation without any manual preprocessing of the eye-tracking signal.</p>
<p>The practical implications extend well beyond the driving simulator. Because the method removes the need for labor-intensive manual coding, it opens the door to analyzing far larger and more diverse datasets, the kind of scale needed for findings to generalize across populations, driving environments and research questions. Larger samples could reveal how visual search strategies differ by age, experience, fatigue, distraction or neurological condition, and could support the evaluation of training interventions designed to teach novice drivers to scan more like experts. Similar approaches could prove valuable in aviation, air traffic control, construction safety, medicine and any domain where situational awareness depends on where people look and when.</p>
<p>The study also reflects a broader trend in behavioral science, in which methods developed in the data mining community, including time-series clustering, shapelet-based classification and related techniques, are being redeployed to make sense of rich, high-frequency behavioral recordings. Eye-trackers have become cheaper and more ubiquitous, generating torrents of gaze data that conventional analysis pipelines were never designed to handle. Techniques that operate directly on raw time series, rather than on heavily processed summaries, promise to preserve the temporal structure of behavior that those summaries often discard.</p>
<p>The authors have made their data and code publicly available through the Children&#8217;s Hospital of Philadelphia&#8217;s GitHub repository, lowering the barrier for other teams to adopt and extend the approach. For a research area long constrained by the slow economics of video coding, the demonstration that raw eye-tracking streams can be clustered into interpretable visual search strategies marks a genuine methodological milestone, one that could accelerate the science of how humans take in the visual world during complex, safety-critical tasks.</p>
<p><strong>Subject of Research:</strong> A scalable time-series clustering method for characterizing visual search strategies from raw eye-tracking data in young drivers</p>
<p><strong>Article Title:</strong> A scalable method for characterizing visual search strategies: A novel application of time-series clustering to raw eye-tracking data</p>
<p><strong>Article References:</strong> Seacrist, T., Walshe, E. E., Grethlein, D., Ryerson, M. S., &amp; Winston, F. K. (2026). A scalable method for characterizing visual search strategies: A novel application of time-series clustering to raw eye-tracking data. <em>Behavior Research Methods, 58</em>(10), Article 290. <a href="https://doi.org/10.3758/s13428-026-03154-2" rel="noopener noreferrer">https://doi.org/10.3758/s13428-026-03154-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.3758/s13428-026-03154-2" rel="noopener noreferrer">10.3758/s13428-026-03154-2</a></p>
<p><strong>Keywords:</strong> eye-tracking, visual search strategy, time-series clustering, machine learning, situational awareness, young drivers, driving simulation, k-medoids, dynamic time warping, hazard perception, Behavior Research Methods, driver safety</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197912</post-id>	</item>
	</channel>
</rss>
