Wednesday, September 30, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Banking AI Gets Leaner: Self-Organizing Maps Slash Data by 99 Percent

September 30, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
Banking AI Gets Leaner: Self-Organizing Maps Slash Data by 99 Percent

Banking AI Gets Leaner: Self-Organizing Maps Slash Data by 99 Percent

Banking AI Gets Leaner: Self-Organizing Maps Slash Data by 99 Percent

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Machine learning has quietly become the backbone of modern banking, powering everything from credit scoring to fraud detection and customer targeting. Yet behind every sleek algorithm lies an increasingly awkward truth: the data feeding these systems has grown so vast and so cluttered that the models themselves are buckling under the weight. A new study published in Neural Computing and Applications by researchers at the University of Seville and the Madrid-based analytics firm GAMCO proposes an elegant escape route. The team, led by Sara Ruiz-Moreno, has developed a data reduction methodology that combines two complementary strategies, numerosity reduction and dimensionality reduction, into a single pipeline built around the k-nearest neighbors classifier. Tested on both a public benchmark and a real bank dataset, the approach shrank the number of stored instances by as much as 99.8 percent while trimming feature counts by up to 82.4 percent, all without sacrificing the classification quality that banks depend on.

The core insight of the work is that most of the records and most of the columns in a typical banking dataset are redundant. When a bank profiles a customer, it may hold hundreds of variables representing that individual, from transaction counts and account ages to demographic markers, and many of these variables overlap heavily. At the same time, thousands or millions of customer records often cluster into patterns that a handful of representative examples can capture. The computational cost of many classification algorithms grows with both the number of instances and the number of features, so every redundant record and every superfluous variable inflates training time, memory usage, and storage requirements. For an industry that runs classification tasks continuously across enormous customer bases, the savings from aggressive but careful reduction can be substantial.

To compress the number of instances, the researchers turned to a self-organizing map, an unsupervised neural network architecture introduced by Teuvo Kohonen in the 1980s. A SOM maps high-dimensional input data onto a usually two-dimensional grid of neurons, preserving the topology of the original data space: similar customers end up near each other on the map. Each input record activates a best matching unit, the neuron whose weight vector most closely resembles the record. By training the map and then using the neuron weight vectors as prototypes, the method replaces thousands of raw records with a far smaller set of synthetic representatives that summarize the structure of the data. In their experiments, this prototype generation step reduced the Census Income benchmark dataset by 99.3 percent and the proprietary banking dataset by 99.8 percent, meaning that fewer than one in a hundred original records needed to be retained to preserve the information the classifier requires.

Compressing instances is only half the battle. The second axis of reduction concerns the features themselves, and here the team proposed five weighting procedures that assign each variable a degree of relevance before the classification step. The most novel of these is a feature weighting technique based directly on the self-organizing map, dubbed FWSOM. Because the SOM has already learned a low-dimensional representation of the data, its internal structure encodes information about which variables drive the organization of the map. FWSOM exploits this by analyzing how strongly each feature contributes to the positioning of records on the trained map, producing weights that can then be folded into a weighted Euclidean distance metric used by the k-nearest neighbors classifier. Variables judged irrelevant receive low weights, effectively muting their influence without requiring a separate feature selection pass.

The remaining four procedures adapt permutation importance, a model-agnostic technique popularized by random forests and widely used in interpretability tools, to the peculiarities of SOM output. Permutation importance works by shuffling the values of one feature at a time and measuring how much the model’s performance degrades; a feature whose shuffling wrecks the predictions is clearly important. The catch is that the standard formulation assumes a conventional supervised model, not a topology-preserving map. The researchers therefore designed four tailored criteria: performance-based PI, which tracks changes in overall accuracy; error-based PI, which monitors shifts in classification error; prediction changes-based PI, which counts how often predictions flip when a feature is permuted; and F1-score-based PI, which focuses on the harmonic mean of precision and recall, a metric better suited to imbalanced datasets where one class, such as defaulters or fraud cases, is rare. Each variant offers a different lens on feature relevance, and the team evaluated all of them empirically.

The experimental design paired a public benchmark with real-world data. The Census Income dataset, drawn from the UCI Machine Learning Repository, is a classic classification problem in which the goal is to predict whether an individual earns above a certain threshold, a task structurally similar to many banking applications such as creditworthiness assessment. Alongside it, the researchers used a proprietary dataset from GAMCO covering real bank customers, giving the evaluation a dose of industrial reality that benchmark studies often lack. Performance was measured using standard metrics including accuracy, balanced accuracy, sensitivity, and false positive rate, ensuring that the compressed models were judged not merely on how fast they ran but on whether they still classified correctly, including on the minority classes that matter most in risk management.

The results were striking on both fronts. Permutation importance proved the more aggressive feature pruner, cutting 82.40 percent of features from the bank dataset and 21.49 percent from the income dataset. FWSOM, while more conservative, still removed 6.4 percent of features from the banking data and 14.29 percent from the income data, with the advantage that it integrates seamlessly into the SOM-based pipeline rather than requiring a separate model. Combined with the near-total prototype reduction, the full methodology delivered datasets that were orders of magnitude smaller than the originals. The practical consequences cascade: lower storage requirements, reduced memory footprints during training, faster classification of new customers, and simpler data analysis overall. For banks operating under tight latency constraints and rising cloud computing costs, a classifier that needs a fraction of the data to reach comparable decisions translates directly into money saved.

What makes the approach particularly appealing for the financial sector is its interpretability. Permutation importance produces an explicit ranking of which variables drive predictions, which aligns with regulatory expectations that automated decisions be explainable. Related research has applied permutation-based methods to default prediction and to identifying biomarkers in medicine, underscoring the technique’s versatility. By fusing that interpretability with the topology-preserving compression of the SOM, the Spanish team has built a pipeline in which the same structure that shrinks the data also reveals which customer attributes actually matter. The work builds on the group’s earlier prototype generation method using a growing self-organizing map, published in the same journal in 2023, extending it from instance reduction alone to a joint treatment of rows and columns.

The methodology also fits into a broader research landscape. Feature selection has been tackled with genetic algorithms, particle swarm optimization, Harris hawks optimization, sine-cosine algorithms, and variational autoencoders, each bringing computational overhead of its own. Prototype selection, meanwhile, has been explored through multi-armed bandits, geometric medians, and multilabel instance-based methods. The SOM-based approach distinguishes itself by addressing both dimensions of reduction within one unsupervised framework, and by being validated on proprietary banking data rather than benchmarks alone. The authors acknowledge limitations: the banking dataset and the code are proprietary and not publicly available, which restricts independent replication, and the balance between reduction and accuracy must be tuned per application. Still, the reported figures suggest that in data-heavy industries, the smartest model may be the one trained on far less. As machine learning spreads further into finance, healthcare, and beyond, techniques that strip away redundancy while preserving signal are likely to become as important as the algorithms they feed.

Subject of Research: A data reduction methodology combining self-organizing maps and permutation importance for k-nearest neighbors classification in the banking sector

Article Title: New data reduction methodology for classification applied to the banking sector

Article References: Ruiz-Moreno, S., Núñez-Reyes, A., García-Cantalapiedra, A., & Pavón, F. (2026). New data reduction methodology for classification applied to the banking sector. Neural Computing and Applications, 38(19), Article 757. https://doi.org/10.1007/s00521-026-12471-8

Image Credits: AI Generated

DOI: 10.1007/s00521-026-12471-8

Keywords: machine learning, self-organizing map, k-nearest neighbors, data reduction, permutation importance, feature selection, banking, prototype generation, dimensionality reduction, Neural Computing and Applications, data, reduction

Cite Scienmag News

Denise Maddox. (September 30, 2026). Banking AI Gets Leaner: Self-Organizing Maps Slash Data by 99 Percent. Scienmag. https://scienmag.com/banking-ai-gets-leaner-self-organizing-maps-slash-data-by-99-percent/

Denise Maddox. "Banking AI Gets Leaner: Self-Organizing Maps Slash Data by 99 Percent." Scienmag, 30 September 2026, https://scienmag.com/banking-ai-gets-leaner-self-organizing-maps-slash-data-by-99-percent/. Accessed 30 September 2026.

Denise Maddox. "Banking AI Gets Leaner: Self-Organizing Maps Slash Data by 99 Percent." Scienmag. September 30, 2026. https://scienmag.com/banking-ai-gets-leaner-self-organizing-maps-slash-data-by-99-percent/

Tags: AI-driven banking data managementbankingbanking data reductioncredit scoring data optimizationcustomer profiling data reductiondatadata reductiondimensionality reductiondimensionality reduction in financial datasetsfeature selectionfraud detection data managementhigh-dimensional financial data analysisk-nearest neighborslarge-scale banking datasetsMachine learningmachine learning in bankingNeural Computing and Applicationsneural networks for banking data efficiencynumerosity reduction techniquespermutation importanceprototype generationReductionself-organizing mapself-organizing maps for data pruning
Share26Tweet16
Previous Post

Decade-Long Case Report Shows How Team-Based Care Tames Recurrent Prostate Cancer

Next Post

Chemical Looping Combustion Cuts Dioxin Emissions From Chlorinated Waste by 87 Percent

Related Posts

Neural Networks Spontaneously Split Into Context and Sensory Specialists
Technology and Engineering

Neural Networks Spontaneously Split Into Context and Sensory Specialists

September 30, 2026
Chemical Looping Combustion Cuts Dioxin Emissions From Chlorinated Waste by 87 Percent
Technology and Engineering

Chemical Looping Combustion Cuts Dioxin Emissions From Chlorinated Waste by 87 Percent

September 30, 2026
Smart Food Network: Tennessee Engineers Win Nearly $1 Million NSF Grant to Rewire Local Food Production
Technology and Engineering

Smart Food Network: Tennessee Engineers Win Nearly $1 Million NSF Grant to Rewire Local Food Production

September 30, 2026
New Tensor Method Cleans Up Messy Multi-View Data in One Efficient Step
Technology and Engineering

New Tensor Method Cleans Up Messy Multi-View Data in One Efficient Step

September 30, 2026
Humanoid Robot Walks Blind Across Grass, Gravel and Steep Ramps Using Only Its Own Senses
Technology and Engineering

Humanoid Robot Walks Blind Across Grass, Gravel and Steep Ramps Using Only Its Own Senses

September 30, 2026
Microrobots Learn to Navigate Blood Vessels in Under Ten Minutes of Training
Technology and Engineering

Microrobots Learn to Navigate Blood Vessels in Under Ten Minutes of Training

September 30, 2026
Next Post
Chemical Looping Combustion Cuts Dioxin Emissions From Chlorinated Waste by 87 Percent

Chemical Looping Combustion Cuts Dioxin Emissions From Chlorinated Waste by 87 Percent

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • From Honey to Healing Pathways: New Review Maps the Molecular Science of Wound Repair
  • Why Himalayan Springs Dry Up: New Study Maps the Hidden Geology of Nepal’s Vanishing Water
  • Water Steals the Best Parking Spots: How Moisture Reshapes Methane Storage in Shale
  • Neural Networks Spontaneously Split Into Context and Sensory Specialists

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading