<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>banking &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/banking/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Wed, 30 Sep 2026 17:10:26 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>banking &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Banking AI Gets Leaner: Self-Organizing Maps Slash Data by 99 Percent</title>
		<link>https://scienmag.com/banking-ai-gets-leaner-self-organizing-maps-slash-data-by-99-percent/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Wed, 30 Sep 2026 17:10:26 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI-driven banking data management]]></category>
		<category><![CDATA[banking]]></category>
		<category><![CDATA[banking data reduction]]></category>
		<category><![CDATA[credit scoring data optimization]]></category>
		<category><![CDATA[customer profiling data reduction]]></category>
		<category><![CDATA[data]]></category>
		<category><![CDATA[data reduction]]></category>
		<category><![CDATA[dimensionality reduction]]></category>
		<category><![CDATA[dimensionality reduction in financial datasets]]></category>
		<category><![CDATA[feature selection]]></category>
		<category><![CDATA[fraud detection data management]]></category>
		<category><![CDATA[high-dimensional financial data analysis]]></category>
		<category><![CDATA[k-nearest neighbors]]></category>
		<category><![CDATA[large-scale banking datasets]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning in banking]]></category>
		<category><![CDATA[Neural Computing and Applications]]></category>
		<category><![CDATA[neural networks for banking data efficiency]]></category>
		<category><![CDATA[numerosity reduction techniques]]></category>
		<category><![CDATA[permutation importance]]></category>
		<category><![CDATA[prototype generation]]></category>
		<category><![CDATA[Reduction]]></category>
		<category><![CDATA[self-organizing map]]></category>
		<category><![CDATA[self-organizing maps for data pruning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=217330</guid>

					<description><![CDATA[Researchers in Spain have unveiled a data reduction method combining self-organizing maps and permutation importance that shrinks banking datasets by up to 99.8 percent while preserving classification performance.]]></description>
										<content:encoded><![CDATA[<p>Machine learning has quietly become the backbone of modern banking, powering everything from credit scoring to fraud detection and customer targeting. Yet behind every sleek algorithm lies an increasingly awkward truth: the data feeding these systems has grown so vast and so cluttered that the models themselves are buckling under the weight. A new study published in Neural Computing and Applications by researchers at the University of Seville and the Madrid-based analytics firm GAMCO proposes an elegant escape route. The team, led by Sara Ruiz-Moreno, has developed a data reduction methodology that combines two complementary strategies, numerosity reduction and dimensionality reduction, into a single pipeline built around the k-nearest neighbors classifier. Tested on both a public benchmark and a real bank dataset, the approach shrank the number of stored instances by as much as 99.8 percent while trimming feature counts by up to 82.4 percent, all without sacrificing the classification quality that banks depend on.</p>
<p>The core insight of the work is that most of the records and most of the columns in a typical banking dataset are redundant. When a bank profiles a customer, it may hold hundreds of variables representing that individual, from transaction counts and account ages to demographic markers, and many of these variables overlap heavily. At the same time, thousands or millions of customer records often cluster into patterns that a handful of representative examples can capture. The computational cost of many classification algorithms grows with both the number of instances and the number of features, so every redundant record and every superfluous variable inflates training time, memory usage, and storage requirements. For an industry that runs classification tasks continuously across enormous customer bases, the savings from aggressive but careful reduction can be substantial.</p>
<p>To compress the number of instances, the researchers turned to a self-organizing map, an unsupervised neural network architecture introduced by Teuvo Kohonen in the 1980s. A SOM maps high-dimensional input data onto a usually two-dimensional grid of neurons, preserving the topology of the original data space: similar customers end up near each other on the map. Each input record activates a best matching unit, the neuron whose weight vector most closely resembles the record. By training the map and then using the neuron weight vectors as prototypes, the method replaces thousands of raw records with a far smaller set of synthetic representatives that summarize the structure of the data. In their experiments, this prototype generation step reduced the Census Income benchmark dataset by 99.3 percent and the proprietary banking dataset by 99.8 percent, meaning that fewer than one in a hundred original records needed to be retained to preserve the information the classifier requires.</p>
<p>Compressing instances is only half the battle. The second axis of reduction concerns the features themselves, and here the team proposed five weighting procedures that assign each variable a degree of relevance before the classification step. The most novel of these is a feature weighting technique based directly on the self-organizing map, dubbed FWSOM. Because the SOM has already learned a low-dimensional representation of the data, its internal structure encodes information about which variables drive the organization of the map. FWSOM exploits this by analyzing how strongly each feature contributes to the positioning of records on the trained map, producing weights that can then be folded into a weighted Euclidean distance metric used by the k-nearest neighbors classifier. Variables judged irrelevant receive low weights, effectively muting their influence without requiring a separate feature selection pass.</p>
<p>The remaining four procedures adapt permutation importance, a model-agnostic technique popularized by random forests and widely used in interpretability tools, to the peculiarities of SOM output. Permutation importance works by shuffling the values of one feature at a time and measuring how much the model&#8217;s performance degrades; a feature whose shuffling wrecks the predictions is clearly important. The catch is that the standard formulation assumes a conventional supervised model, not a topology-preserving map. The researchers therefore designed four tailored criteria: performance-based PI, which tracks changes in overall accuracy; error-based PI, which monitors shifts in classification error; prediction changes-based PI, which counts how often predictions flip when a feature is permuted; and F1-score-based PI, which focuses on the harmonic mean of precision and recall, a metric better suited to imbalanced datasets where one class, such as defaulters or fraud cases, is rare. Each variant offers a different lens on feature relevance, and the team evaluated all of them empirically.</p>
<p>The experimental design paired a public benchmark with real-world data. The Census Income dataset, drawn from the UCI Machine Learning Repository, is a classic classification problem in which the goal is to predict whether an individual earns above a certain threshold, a task structurally similar to many banking applications such as creditworthiness assessment. Alongside it, the researchers used a proprietary dataset from GAMCO covering real bank customers, giving the evaluation a dose of industrial reality that benchmark studies often lack. Performance was measured using standard metrics including accuracy, balanced accuracy, sensitivity, and false positive rate, ensuring that the compressed models were judged not merely on how fast they ran but on whether they still classified correctly, including on the minority classes that matter most in risk management.</p>
<p>The results were striking on both fronts. Permutation importance proved the more aggressive feature pruner, cutting 82.40 percent of features from the bank dataset and 21.49 percent from the income dataset. FWSOM, while more conservative, still removed 6.4 percent of features from the banking data and 14.29 percent from the income data, with the advantage that it integrates seamlessly into the SOM-based pipeline rather than requiring a separate model. Combined with the near-total prototype reduction, the full methodology delivered datasets that were orders of magnitude smaller than the originals. The practical consequences cascade: lower storage requirements, reduced memory footprints during training, faster classification of new customers, and simpler data analysis overall. For banks operating under tight latency constraints and rising cloud computing costs, a classifier that needs a fraction of the data to reach comparable decisions translates directly into money saved.</p>
<p>What makes the approach particularly appealing for the financial sector is its interpretability. Permutation importance produces an explicit ranking of which variables drive predictions, which aligns with regulatory expectations that automated decisions be explainable. Related research has applied permutation-based methods to default prediction and to identifying biomarkers in medicine, underscoring the technique&#8217;s versatility. By fusing that interpretability with the topology-preserving compression of the SOM, the Spanish team has built a pipeline in which the same structure that shrinks the data also reveals which customer attributes actually matter. The work builds on the group&#8217;s earlier prototype generation method using a growing self-organizing map, published in the same journal in 2023, extending it from instance reduction alone to a joint treatment of rows and columns.</p>
<p>The methodology also fits into a broader research landscape. Feature selection has been tackled with genetic algorithms, particle swarm optimization, Harris hawks optimization, sine-cosine algorithms, and variational autoencoders, each bringing computational overhead of its own. Prototype selection, meanwhile, has been explored through multi-armed bandits, geometric medians, and multilabel instance-based methods. The SOM-based approach distinguishes itself by addressing both dimensions of reduction within one unsupervised framework, and by being validated on proprietary banking data rather than benchmarks alone. The authors acknowledge limitations: the banking dataset and the code are proprietary and not publicly available, which restricts independent replication, and the balance between reduction and accuracy must be tuned per application. Still, the reported figures suggest that in data-heavy industries, the smartest model may be the one trained on far less. As machine learning spreads further into finance, healthcare, and beyond, techniques that strip away redundancy while preserving signal are likely to become as important as the algorithms they feed.</p>
<p><strong>Subject of Research:</strong> A data reduction methodology combining self-organizing maps and permutation importance for k-nearest neighbors classification in the banking sector</p>
<p><strong>Article Title:</strong> New data reduction methodology for classification applied to the banking sector</p>
<p><strong>Article References:</strong> Ruiz-Moreno, S., Núñez-Reyes, A., García-Cantalapiedra, A., &amp; Pavón, F. (2026). New data reduction methodology for classification applied to the banking sector. <em>Neural Computing and Applications, 38</em>(19), Article 757. <a href="https://doi.org/10.1007/s00521-026-12471-8" rel="noopener noreferrer">https://doi.org/10.1007/s00521-026-12471-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s00521-026-12471-8" rel="noopener noreferrer">10.1007/s00521-026-12471-8</a></p>
<p><strong>Keywords:</strong> machine learning, self-organizing map, k-nearest neighbors, data reduction, permutation importance, feature selection, banking, prototype generation, dimensionality reduction, Neural Computing and Applications, data, reduction</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">217330</post-id>	</item>
	</channel>
</rss>
