<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>self-organizing map &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/self-organizing-map/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 17:10:26 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>self-organizing map &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Banking AI Gets Leaner: Self-Organizing Maps Slash Data by 99 Percent</title>
		<link>https://scienmag.com/banking-ai-gets-leaner-self-organizing-maps-slash-data-by-99-percent/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 17:10:26 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-driven banking data management]]></category>
		<category><![CDATA[banking]]></category>
		<category><![CDATA[banking data reduction]]></category>
		<category><![CDATA[credit scoring data optimization]]></category>
		<category><![CDATA[customer profiling data reduction]]></category>
		<category><![CDATA[data]]></category>
		<category><![CDATA[data reduction]]></category>
		<category><![CDATA[dimensionality reduction]]></category>
		<category><![CDATA[dimensionality reduction in financial datasets]]></category>
		<category><![CDATA[feature selection]]></category>
		<category><![CDATA[fraud detection data management]]></category>
		<category><![CDATA[high-dimensional financial data analysis]]></category>
		<category><![CDATA[k-nearest neighbors]]></category>
		<category><![CDATA[large-scale banking datasets]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in banking]]></category>
		<category><![CDATA[Neural Computing and Applications]]></category>
		<category><![CDATA[neural networks for banking data efficiency]]></category>
		<category><![CDATA[numerosity reduction techniques]]></category>
		<category><![CDATA[permutation importance]]></category>
		<category><![CDATA[prototype generation]]></category>
		<category><![CDATA[Reduction]]></category>
		<category><![CDATA[self-organizing map]]></category>
		<category><![CDATA[self-organizing maps for data pruning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217330</guid>

					<description><![CDATA[Researchers in Spain have unveiled a data reduction method combining self-organizing maps and permutation importance that shrinks banking datasets by up to 99.8 percent while preserving classification performance.]]></description>
										<content:encoded><![CDATA[<p>Machine learning has quietly become the backbone of modern banking, powering everything from credit scoring to fraud detection and customer targeting. Yet behind every sleek algorithm lies an increasingly awkward truth: the data feeding these systems has grown so vast and so cluttered that the models themselves are buckling under the weight. A new study published in Neural Computing and Applications by researchers at the University of Seville and the Madrid-based analytics firm GAMCO proposes an elegant escape route. The team, led by Sara Ruiz-Moreno, has developed a data reduction methodology that combines two complementary strategies, numerosity reduction and dimensionality reduction, into a single pipeline built around the k-nearest neighbors classifier. Tested on both a public benchmark and a real bank dataset, the approach shrank the number of stored instances by as much as 99.8 percent while trimming feature counts by up to 82.4 percent, all without sacrificing the classification quality that banks depend on.</p>
<p>The core insight of the work is that most of the records and most of the columns in a typical banking dataset are redundant. When a bank profiles a customer, it may hold hundreds of variables representing that individual, from transaction counts and account ages to demographic markers, and many of these variables overlap heavily. At the same time, thousands or millions of customer records often cluster into patterns that a handful of representative examples can capture. The computational cost of many classification algorithms grows with both the number of instances and the number of features, so every redundant record and every superfluous variable inflates training time, memory usage, and storage requirements. For an industry that runs classification tasks continuously across enormous customer bases, the savings from aggressive but careful reduction can be substantial.</p>
<p>To compress the number of instances, the researchers turned to a self-organizing map, an unsupervised neural network architecture introduced by Teuvo Kohonen in the 1980s. A SOM maps high-dimensional input data onto a usually two-dimensional grid of neurons, preserving the topology of the original data space: similar customers end up near each other on the map. Each input record activates a best matching unit, the neuron whose weight vector most closely resembles the record. By training the map and then using the neuron weight vectors as prototypes, the method replaces thousands of raw records with a far smaller set of synthetic representatives that summarize the structure of the data. In their experiments, this prototype generation step reduced the Census Income benchmark dataset by 99.3 percent and the proprietary banking dataset by 99.8 percent, meaning that fewer than one in a hundred original records needed to be retained to preserve the information the classifier requires.</p>
<p>Compressing instances is only half the battle. The second axis of reduction concerns the features themselves, and here the team proposed five weighting procedures that assign each variable a degree of relevance before the classification step. The most novel of these is a feature weighting technique based directly on the self-organizing map, dubbed FWSOM. Because the SOM has already learned a low-dimensional representation of the data, its internal structure encodes information about which variables drive the organization of the map. FWSOM exploits this by analyzing how strongly each feature contributes to the positioning of records on the trained map, producing weights that can then be folded into a weighted Euclidean distance metric used by the k-nearest neighbors classifier. Variables judged irrelevant receive low weights, effectively muting their influence without requiring a separate feature selection pass.</p>
<p>The remaining four procedures adapt permutation importance, a model-agnostic technique popularized by random forests and widely used in interpretability tools, to the peculiarities of SOM output. Permutation importance works by shuffling the values of one feature at a time and measuring how much the model&#8217;s performance degrades; a feature whose shuffling wrecks the predictions is clearly important. The catch is that the standard formulation assumes a conventional supervised model, not a topology-preserving map. The researchers therefore designed four tailored criteria: performance-based PI, which tracks changes in overall accuracy; error-based PI, which monitors shifts in classification error; prediction changes-based PI, which counts how often predictions flip when a feature is permuted; and F1-score-based PI, which focuses on the harmonic mean of precision and recall, a metric better suited to imbalanced datasets where one class, such as defaulters or fraud cases, is rare. Each variant offers a different lens on feature relevance, and the team evaluated all of them empirically.</p>
<p>The experimental design paired a public benchmark with real-world data. The Census Income dataset, drawn from the UCI Machine Learning Repository, is a classic classification problem in which the goal is to predict whether an individual earns above a certain threshold, a task structurally similar to many banking applications such as creditworthiness assessment. Alongside it, the researchers used a proprietary dataset from GAMCO covering real bank customers, giving the evaluation a dose of industrial reality that benchmark studies often lack. Performance was measured using standard metrics including accuracy, balanced accuracy, sensitivity, and false positive rate, ensuring that the compressed models were judged not merely on how fast they ran but on whether they still classified correctly, including on the minority classes that matter most in risk management.</p>
<p>The results were striking on both fronts. Permutation importance proved the more aggressive feature pruner, cutting 82.40 percent of features from the bank dataset and 21.49 percent from the income dataset. FWSOM, while more conservative, still removed 6.4 percent of features from the banking data and 14.29 percent from the income data, with the advantage that it integrates seamlessly into the SOM-based pipeline rather than requiring a separate model. Combined with the near-total prototype reduction, the full methodology delivered datasets that were orders of magnitude smaller than the originals. The practical consequences cascade: lower storage requirements, reduced memory footprints during training, faster classification of new customers, and simpler data analysis overall. For banks operating under tight latency constraints and rising cloud computing costs, a classifier that needs a fraction of the data to reach comparable decisions translates directly into money saved.</p>
<p>What makes the approach particularly appealing for the financial sector is its interpretability. Permutation importance produces an explicit ranking of which variables drive predictions, which aligns with regulatory expectations that automated decisions be explainable. Related research has applied permutation-based methods to default prediction and to identifying biomarkers in medicine, underscoring the technique&#8217;s versatility. By fusing that interpretability with the topology-preserving compression of the SOM, the Spanish team has built a pipeline in which the same structure that shrinks the data also reveals which customer attributes actually matter. The work builds on the group&#8217;s earlier prototype generation method using a growing self-organizing map, published in the same journal in 2023, extending it from instance reduction alone to a joint treatment of rows and columns.</p>
<p>The methodology also fits into a broader research landscape. Feature selection has been tackled with genetic algorithms, particle swarm optimization, Harris hawks optimization, sine-cosine algorithms, and variational autoencoders, each bringing computational overhead of its own. Prototype selection, meanwhile, has been explored through multi-armed bandits, geometric medians, and multilabel instance-based methods. The SOM-based approach distinguishes itself by addressing both dimensions of reduction within one unsupervised framework, and by being validated on proprietary banking data rather than benchmarks alone. The authors acknowledge limitations: the banking dataset and the code are proprietary and not publicly available, which restricts independent replication, and the balance between reduction and accuracy must be tuned per application. Still, the reported figures suggest that in data-heavy industries, the smartest model may be the one trained on far less. As machine learning spreads further into finance, healthcare, and beyond, techniques that strip away redundancy while preserving signal are likely to become as important as the algorithms they feed.</p>
<p><strong>Subject of Research:</strong> A data reduction methodology combining self-organizing maps and permutation importance for k-nearest neighbors classification in the banking sector</p>
<p><strong>Article Title:</strong> New data reduction methodology for classification applied to the banking sector</p>
<p><strong>Article References:</strong> Ruiz-Moreno, S., Núñez-Reyes, A., García-Cantalapiedra, A., &amp; Pavón, F. (2026). New data reduction methodology for classification applied to the banking sector. <em>Neural Computing and Applications, 38</em>(19), Article 757. <a href="https://doi.org/10.1007/s00521-026-12471-8" rel="noopener noreferrer">https://doi.org/10.1007/s00521-026-12471-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00521-026-12471-8" rel="noopener noreferrer">10.1007/s00521-026-12471-8</a></p>
<p><strong>Keywords:</strong> machine learning, self-organizing map, k-nearest neighbors, data reduction, permutation importance, feature selection, banking, prototype generation, dimensionality reduction, Neural Computing and Applications, data, reduction</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217330</post-id>	</item>
		<item>
		<title>Hyperspectral Camera and AI Map Hidden Microplastics in Sand Without Sampling</title>
		<link>https://scienmag.com/hyperspectral-camera-and-ai-map-hidden-microplastics-in-sand-without-sampling/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 19:31:38 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[advanced AI techniques for environmental monitoring]]></category>
		<category><![CDATA[AI-based microplastic identification in sand]]></category>
		<category><![CDATA[beach pollution]]></category>
		<category><![CDATA[chemical signature mapping of plastics in soil and sand]]></category>
		<category><![CDATA[Environmental Monitoring]]></category>
		<category><![CDATA[hyperspectral imaging]]></category>
		<category><![CDATA[hyperspectral imaging for microplastic detection]]></category>
		<category><![CDATA[innovative methods for microplastic pollution measurement]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[microplastics]]></category>
		<category><![CDATA[near-infrared hyperspectral imaging in environmental analysis]]></category>
		<category><![CDATA[non-destructive analysis]]></category>
		<category><![CDATA[non-invasive microplastic contamination assessment]]></category>
		<category><![CDATA[PET]]></category>
		<category><![CDATA[polyethylene]]></category>
		<category><![CDATA[polypropylene]]></category>
		<category><![CDATA[polystyrene]]></category>
		<category><![CDATA[real-time microplastic detection with hyperspectral imaging]]></category>
		<category><![CDATA[remote sensing of microplastics using hyperspectral cameras]]></category>
		<category><![CDATA[sand substrates]]></category>
		<category><![CDATA[self-organizing map]]></category>
		<category><![CDATA[self-organizing map neural networks for pollution detection]]></category>
		<category><![CDATA[unsupervised machine learning for pollutant mapping]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=197928</guid>

					<description><![CDATA[Researchers have combined near-infrared hyperspectral imaging with a self-organizing map neural network to identify and semi-quantitatively map microplastic contamination on sand surfaces without collecting or destroying samples.]]></description>
										<content:encoded><![CDATA[<p>Microplastics have become one of the most pervasive pollutants on the planet, turning up everywhere from the deepest ocean trenches to the air we breathe. Yet for all the alarm surrounding these tiny fragments, scientists still lack fast, reliable ways to measure how heavily a beach, a riverbank, or an agricultural soil is contaminated without scooping up samples and hauling them back to a laboratory. A new study published in the journal Microplastics and Nanoplastics offers a striking answer: a camera system paired with an unsupervised machine learning algorithm that can photograph a patch of sand and simultaneously identify which polymers are present and how much of the surface they cover, all without touching a single grain.</p>
<p>The research, led by Sureerat Makmuang and Kanet Wongravee of the Sensor Research Unit at Chulalongkorn University in Thailand, together with Simon Maher of the University of Liverpool and Sanong Ekgasit of Chulalongkorn University, combines near-infrared hyperspectral imaging (NIR-HSI) with a self-organizing map, or SOM, a type of artificial neural network that learns to organize complex data without being told what to look for. The result is a workflow that transforms the invisible chemical signatures of plastics into vivid color-coded maps, in which each polymer class—PET, polyethylene, polypropylene, or polystyrene—appears as its own distinct hue painted across the sandy terrain.</p>
<p>The physics behind the technique is elegant. When near-infrared light strikes a plastic fragment, the chemical bonds within the polymer absorb specific wavelengths in patterns as unique as fingerprints. A hyperspectral camera captures not just a conventional image but a full spectrum at every pixel, effectively recording hundreds of narrow wavelength bands simultaneously. Each pixel therefore carries a chemical identity waiting to be decoded. The challenge has always been interpretation: a sandy surface littered with fragments of different shapes, sizes, colors, and orientations produces spectra that are noisy, mixed, and difficult to separate with conventional statistical tools.</p>
<p>That is where the self-organizing map enters. An SOM is trained on the spectral data by repeatedly adjusting an internal grid of artificial neurons so that similar spectra cluster together, forming a topology that mirrors the chemical relationships in the data. In this study, the researchers enhanced the standard approach with a modified SOM and a novel percent-based expansion tolerance, or PBET, scheme that allows the model to estimate semi-quantitatively how much of a scanned surface is covered by each polymer type. The output is doubly informative: qualitative maps that show exactly where each class of microplastic sits, and quantitative coverage estimates expressed as percentages of the imaged area.</p>
<p>To test the system, the team prepared fragments of polyethylene terephthalate, polyethylene, polypropylene, and polystyrene from everyday household plastic materials, cut them into particles ranging from one to five millimeters, and distributed them on sand surfaces at controlled coverage levels spanning roughly 0.78 to 12.5 percent. After preprocessing the hyperspectral data to sharpen spectral quality, the modified SOMs classified the plastics with remarkable fidelity. Visually, individual fragments and mixed-polymer scenes alike were rendered as clean, color-separated maps. Quantitatively, the model&#8217;s predictions of surface coverage achieved coefficients of determination reaching as high as 1.00, with very low root-mean-square errors—performance figures that suggest the approach can rival labor-intensive reference methods.</p>
<p>Critically, the researchers did not stop at idealized laboratory conditions. They deliberately stressed the model with sources of real-world variability that plague field measurements: variations in particle size, differences in pigment color, and overlapping particles that stack atop one another and produce mixed spectra. The SOM workflow remained robust under these challenges, holding its classification accuracy where simpler methods would falter. The team then pushed the test further by imaging microplastic particles collected from natural beach samples—plastics the model had never seen before, weathered and coated by the environment. It identified and classified them correctly, a demonstration that the laboratory-trained system generalizes to the messy chemistry of the real world.</p>
<p>Perhaps the most striking finding concerns size. Conventional visual surveys of microplastic contamination rely on human eyes or standard photography, both of which routinely miss particles below about one millimeter. The hyperspectral approach proved capable of detecting and correctly classifying particles smaller than that threshold, underscoring a sensitivity that could close one of the largest blind spots in microplastic monitoring. Because smaller fragments are generally more bioavailable to organisms—and more likely to carry adsorbed toxins—this capability matters not just for counting pollution but for assessing its ecological risk.</p>
<p>The non-destructive nature of the method is its other defining advantage. Traditional microplastic analysis typically requires collecting sediment, transporting it to a lab, digesting organic matter, and running samples through spectroscopic instruments such as Fourier-transform infrared or Raman spectrometers—accurate techniques, but slow, costly, and destructive to the sample. Hyperspectral imaging flips that model: the sand stays in place, the measurement takes the form of a scan, and the same patch of ground can be revisited over time to track how contamination evolves. That opens the door to genuine longitudinal monitoring of beaches, dunes, and remediation sites, where repeated sampling has historically been impractical.</p>
<p>The researchers are careful to frame the achievement within the boundaries of their experiments. The workflow was developed and evaluated on sandy substrates with the four most common commodity polymers, under controlled illumination and geometry, and the coverage estimates are semi-quantitative rather than exhaustive particle counts. Wet sediments, dark soils, biofilms, and polymers beyond the tested four remain open challenges, and translating laboratory performance to drones or handheld field scanners will require further engineering. Yet the analytical foundation the study establishes—polymer-class mapping paired with surface-coverage estimation in a single rapid scan—is precisely the kind of groundwork needed before such instruments can be built.</p>
<p>If the approach matures as the results suggest, the implications reach far beyond sandy shores. Agricultural soils amended with plastic mulch fragments, construction sites receiving recycled aggregates, and coastal zones awaiting cleanup all demand the same basic information: which plastics are present, where, and in what abundance. By fusing hyperspectral imaging with self-organizing maps, this study demonstrates that answer can be rendered almost photographically—a colored chemical portrait of pollution that regulators, remediation engineers, and the public can read at a glance. In a world drowning in plastic fragments too small to see, a camera that makes them visible may prove one of the most consequential environmental tools of the decade.</p>
<p><strong>Subject of Research:</strong> Non-destructive detection and mapping of microplastic contamination in sandy substrates using hyperspectral imaging and self-organizing maps</p>
<p><strong>Article Title:</strong> Hyperspectral imaging and self-organizing map approach for non-destructive monitoring of microplastic contamination in sandy substrates</p>
<p><strong>Article References:</strong> Makmuang, S., Maher, S., Ekgasit, S., &amp; Wongravee, K. (2026). Hyperspectral imaging and self-organizing map approach for non-destructive monitoring of microplastic contamination in sandy substrates. <em>Microplastics and Nanoplastics</em>. <a href="https://doi.org/10.1186/s43591-026-00225-1" rel="noopener noreferrer">https://doi.org/10.1186/s43591-026-00225-1</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s43591-026-00225-1" rel="noopener noreferrer">10.1186/s43591-026-00225-1</a></p>
<p><strong>Keywords:</strong> microplastics, hyperspectral imaging, self-organizing map, machine learning, sand substrates, polyethylene, polypropylene, polystyrene, PET, environmental monitoring, non-destructive analysis, beach pollution</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">197928</post-id>	</item>
	</channel>
</rss>
