<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>transparent machine learning models &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/transparent-machine-learning-models/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Thu, 10 Sep 2026 05:28:11 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>transparent machine learning models &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Two interpretable supervised algorithms reveal feature aggregation in nonlinear systems</title>
		<link>https://scienmag.com/two-interpretable-supervised-algorithms-reveal-feature-aggregation-in-nonlinear-systems/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Thu, 10 Sep 2026 05:28:08 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[climate science data interpretation]]></category>
		<category><![CDATA[data mining and knowledge discovery methods]]></category>
		<category><![CDATA[dimensionality reduction techniques]]></category>
		<category><![CDATA[feature aggregation in complex systems]]></category>
		<category><![CDATA[feature aggregation in nonlinear systems]]></category>
		<category><![CDATA[feature interpretability in high-dimensional data]]></category>
		<category><![CDATA[genomics data dimensionality reduction]]></category>
		<category><![CDATA[high-dimensional data visualization]]></category>
		<category><![CDATA[image recognition feature analysis]]></category>
		<category><![CDATA[interpretable dimensionality reduction]]></category>
		<category><![CDATA[interpretable feature extraction methods]]></category>
		<category><![CDATA[interpretable machine learning]]></category>
		<category><![CDATA[machine learning for massive datasets]]></category>
		<category><![CDATA[nonlinear data analysis in climate science]]></category>
		<category><![CDATA[nonlinear data analysis tools]]></category>
		<category><![CDATA[nonlinear dimensionality reduction]]></category>
		<category><![CDATA[nonlinear feature aggregation]]></category>
		<category><![CDATA[overcoming overfitting in machine learning]]></category>
		<category><![CDATA[supervised algorithms for high-dimensional data]]></category>
		<category><![CDATA[supervised machine learning algorithms]]></category>
		<category><![CDATA[transparent data analysis in genomics]]></category>
		<category><![CDATA[transparent machine learning models]]></category>
		<guid isPermaLink="false">https://scienmag.com/two-interpretable-supervised-algorithms-reveal-feature-aggregation-in-nonlinear-systems/</guid>

					<description><![CDATA[Researchers at Politecnico di Milano have unveiled a pair of new machine learning algorithms that shrink massive datasets while keeping every feature interpretable, extending dimensionality reduction into the nonlinear territory where most real-world data lives. The algorithms, called NonLinCFA and GenLinCFA, are described in a study published in the open-access journal Data Mining and Knowledge [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>Researchers at Politecnico di Milano have unveiled a pair of new machine learning algorithms that shrink massive datasets while keeping every feature interpretable, extending dimensionality reduction into the nonlinear territory where most real-world data lives. The algorithms, called NonLinCFA and GenLinCFA, are described in a study published in the open-access journal Data Mining and Knowledge Discovery, and they promise to make high-dimensional analysis more transparent in fields ranging from climate science to genomics and image recognition.</p>
<p>Dimensionality reduction is one of the oldest and most practical problems in machine learning. Modern datasets can contain thousands or even tens of thousands of variables, and feeding all of them into a model invites trouble: computations slow down, memory demands balloon, and models risk overfitting, memorizing noise rather than learning genuine patterns. The classical solutions come in two flavors. Feature extraction methods such as Principal Component Analysis compress many variables into a handful of latent components, while feature selection simply keeps the most informative variables and throws away the rest. Both come with a catch. PCA-style components are dense linear combinations of nearly all the original features, weighted by coefficients that a domain expert cannot easily interpret, and feature selection discards information that might have been valuable.</p>
<p>The new work, led by Paolo Bonetti together with Alberto Maria Metelli and Marcello Restelli, builds on an earlier technique the same group introduced, known as LinCFA, which took a third path: instead of extracting latent variables or discarding features, it merges correlated groups of features by replacing them with their average. Averaging is a transformation any scientist can understand without consulting a machine learning specialist, and no feature is ever lost entirely, because each one contributes to some mean. The catch with the original method was its assumption that the relationship between features and targets was linear, an assumption that rarely holds in practice and that reduced the technique to a heuristic whenever real data behaved nonlinearly.</p>
<p>NonLinCFA removes that restriction. The researchers consider a general nonlinear function linking the features to the target, corrupted by additive Gaussian noise, and they carry out a rigorous asymptotic bias-variance analysis of what happens when two features are replaced by a single aggregated variable. Their analysis yields precise expressions for both sides of the trade-off. On one side, merging features shrinks the hypothesis space, cutting the variance of the fitted model by exactly the noise variance divided by the number of samples. On the other side, aggregation increases bias, by an amount that depends on the correlations between the two features, their aggregation, and the true underlying function. Combining the two results produces an elegant criterion: merging two features is guaranteed not to increase the mean squared error if and only if the variance reduction, sigma squared divided by n minus one, exceeds the difference between how well the target is predicted by the two features jointly and how well it is predicted by the aggregated version alone. Intuitively, aggregation pays off when the noise is large, the sample is small, the aggregated feature shares plenty of information with the target, and the individual features contribute little unique information of their own.</p>
<p>The second algorithm, GenLinCFA, pushes the framework even further into generalized linear models, the statistical workhorses behind logistic regression, Poisson regression, and many other techniques. Here the target&#8217;s distribution is assumed to belong to the canonical exponential family, which covers the Normal, Exponential, Poisson, Bernoulli, and Binomial distributions, and the expected value of the target passes through a link function before being connected to the features. Because the mean squared error no longer provides a natural goodness-of-fit measure in this setting, the researchers instead analyze the deviance, a quantity that measures how far the fitted model&#8217;s likelihood falls short of the best achievable one. Through a second-order approximation, they derive an upper bound on how much the expected deviance can increase when two features are merged, and this bound becomes the merging criterion of GenLinCFA. Crucially, this makes the method applicable to classification problems, which the original linear approach could not handle at all.</p>
<p>In both algorithms the workflow is the same at heart. The procedure iteratively examines candidate pairs of features, applies a method-specific merge test derived from theory, and replaces a pair with its average whenever the test indicates this is beneficial. The process repeats until no further aggregation is worthwhile, producing a reduced set of features, each of which is the mean of a group of original variables. The computational cost remains quadratic in the number of features and linear in the number of samples, the same as the earlier linear method, and no more memory is required than storing the original dataset. A single hyperparameter called epsilon governs how aggressive the algorithms are: large values encourage more merging, small values keep the method conservative.</p>
<p>The theoretical guarantees were validated through an extensive experimental campaign. On synthetic regression problems with 100 and 1000 features, both algorithms achieved strong predictive scores using far fewer reduced features than a wrapper feature-selection baseline needed to match the same performance. The earlier LinCFA method, applied to the same data, produced 39 reduced features with an R-squared score of about 0.866 in the smaller setting, and nearly 194 reduced features with a score of about 0.707 in the larger, noisier one, confirming that the nonlinear extensions extract more compact and more accurate representations. When the sign of the target was applied to create binary classification tasks, GenLinCFA matched or beat the wrapper baseline while shrinking the dimensionality dramatically.</p>
<p>Real-world benchmarks reinforced the picture. The team tested their methods on four datasets from Kaggle and the UCI Machine Learning Repository, including a finance dataset, a corporate bankruptcy dataset, a Parkinson&#8217;s disease classification dataset, and a gene expression dataset with nearly 20,000 features but only about 800 samples, a notoriously challenging regime. Against ten established baselines, including PCA, Linear Discriminant Analysis, Kernel PCA, Isomap, Locally Linear Embedding, UMAP, t-SNE, autoencoders, supervised PCA, and Neighborhood Components Analysis, the new algorithms were competitive across the board and outperformed their linear predecessor in almost every case. More complex nonlinear projections such as Kernel PCA and autoencoders sometimes achieved better raw scores, but at the cost of producing features whose meaning is opaque.</p>
<p>Perhaps the most striking demonstration came from climate science, the motivating application behind the research. When predicting the vegetation state of a sub-basin of the Po River in Northern Italy from temperature and precipitation measurements taken at many locations, the algorithms automatically discovered spatial sub-regions within which measurements could be safely averaged. Because climatologists routinely aggregate neighboring measurements by hand, the data-driven partitions produced by NonLinCFA are immediately meaningful to them: a map of merged temperature readings over a sub-region reads like a familiar climatological summary rather than an inscrutable mathematical object. Additional experiments on MNIST and Fashion-MNIST image datasets, including a regression task predicting a central pixel from its neighbors and binary classification of visually similar digits, confirmed that the methods remain competitive with state-of-the-art dimensionality reduction while preserving interpretability.</p>
<p>The researchers acknowledge limitations, including the sensitivity of averaging to outliers, which can be mitigated with robust preprocessing, and the fact that their exact theoretical results currently cover linear models within the nonlinear framework. They also note that when features represent heterogeneous quantities, the average of a merged group may lack a concrete physical meaning, even though it remains transparent and empirically useful. Still, the study demonstrates that interpretability and nonlinear performance need not be opposing goals. By grounding a simple, human-readable operation, the mean, in formal bias-variance and deviance analysis, the Milan team has shown that it is possible to compress data aggressively without losing the ability to explain what the compressed features actually are. The authors suggest that future theoretical work could extend the guarantees to broader classes of supervised models, and that large-scale applications could deepen the empirical contribution, with the code and datasets made freely available to the community.</p>
<div class="scienmag-article-metadata"><strong>Subject of Research:</strong> Interpretable supervised dimensionality reduction through feature aggregation in nonlinear and generalized linear systems, via the NonLinCFA and GenLinCFA algorithms.</p>
<p><strong>Article Title:</strong> Feature aggregation in nonlinear systems: two interpretable supervised algorithms</p>
<p><strong>Article References:</strong> Bonetti, P., Metelli, A. M., &amp; Restelli, M. (2026). Feature aggregation in nonlinear systems: two interpretable supervised algorithms. <em>Data Mining and Knowledge Discovery, 40</em>(4), Article 59. <a href="https://doi.org/10.1007/s10618-026-01219-6" target="_blank" rel="noopener noreferrer">https://doi.org/10.1007/s10618-026-01219-6</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s10618-026-01219-6" target="_blank" rel="noopener noreferrer">10.1007/s10618-026-01219-6</a></p>
<p><strong>Keywords:</strong> dimensionality reduction, feature aggregation, interpretable machine learning, supervised learning, bias-variance tradeoff, generalized linear models, deviance analysis, nonlinear regression, feature selection, climate data, classification, machine learning algorithms</p>
</div>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">191295</post-id>	</item>
		<item>
		<title>AI Insights Uncover Causes of Injury Deaths</title>
		<link>https://scienmag.com/ai-insights-uncover-causes-of-injury-deaths/</link>
		
		<dc:creator><![CDATA[Phoebe Ingram]]></dc:creator>
		<pubDate>Sun, 24 May 2026 19:17:21 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI for policy-making in health]]></category>
		<category><![CDATA[AI in public health surveillance]]></category>
		<category><![CDATA[AI-driven health intervention strategies]]></category>
		<category><![CDATA[explainable AI for injury mortality]]></category>
		<category><![CDATA[geographic data in injury prevention]]></category>
		<category><![CDATA[intentional injury epidemiology analysis]]></category>
		<category><![CDATA[machine learning in epidemiology]]></category>
		<category><![CDATA[non-linear pattern detection in health data]]></category>
		<category><![CDATA[socioeconomic factors in injury deaths]]></category>
		<category><![CDATA[suicide and homicide risk prediction]]></category>
		<category><![CDATA[temporal analysis of injury mortality]]></category>
		<category><![CDATA[transparent machine learning models]]></category>
		<guid isPermaLink="false">https://scienmag.com/ai-insights-uncover-causes-of-injury-deaths/</guid>

					<description><![CDATA[In a groundbreaking development at the intersection of artificial intelligence and public health, researchers have unveiled a novel, explainable AI framework designed to address one of the most persistent and tragic crises in the Americas: intentional injury mortality, encompassing both suicide and homicide. This pioneering approach, detailed in a recent publication in Scientific Reports, harnesses [&#8230;]]]></description>
										<content:encoded><![CDATA[<p>In a groundbreaking development at the intersection of artificial intelligence and public health, researchers have unveiled a novel, explainable AI framework designed to address one of the most persistent and tragic crises in the Americas: intentional injury mortality, encompassing both suicide and homicide. This pioneering approach, detailed in a recent publication in Scientific Reports, harnesses the power of transparent machine learning models to dissect the complex epidemiology of intentional injuries, with a view toward enhancing surveillance, intervention, and policy-making.</p>
<p>The Americas have long grappled with high rates of intentional injuries, a public health concern with profound societal ramifications. Traditional surveillance methods have struggled to capture the nuanced socio-demographic and environmental factors that contribute to these mortalities. The newly introduced explainable AI model represents a paradigm shift, combining predictive power with interpretability, thereby enabling stakeholders not only to anticipate at-risk populations but also to understand the underlying drivers of risk.</p>
<p>At the core of this innovative work lies a sophisticated integration of heterogeneous data sources, ranging from socioeconomic indicators and health records to geographic and temporal variables. The researchers leveraged these multifaceted datasets to train AI algorithms capable of detecting subtle, non-linear patterns that elude conventional statistical techniques. Crucially, the explainability component of the model translates these complex associations into human-understandable insights, facilitating transparent decision-making that can earn public trust and inform targeted interventions.</p>
<p>The research team meticulously designed the AI framework to balance accuracy with interpretability, employing state-of-the-art explainable machine learning techniques such as SHAP (SHapley Additive exPlanations) and attention mechanisms. These methodologies enable the deconstruction of model predictions into feature contributions, allowing epidemiologists and policymakers to pinpoint which factors most significantly influence suicide and homicide rates in diverse populations and environments. This transparency is a vital advancement in AI ethics and accountability within public health domains.</p>
<p>By applying their AI model to data spanning multiple countries in the Americas, the investigators uncovered persistent regional disparities in intentional injury mortality that had previously been inadequately understood. Their approach illuminated the complex interplay between economic deprivation, mental health resource availability, urbanization, and demographic factors, revealing distinct profiles of vulnerability across different communities. These insights could catalyze more equitable allocation of resources and customized prevention strategies.</p>
<p>Furthermore, the study’s findings challenge some prevailing assumptions in the field. For example, while socioeconomic disadvantage is a well-documented risk factor for intentional injuries, the AI analysis highlighted that its impact is modulated by other contextual elements such as cultural attitudes toward mental health and the presence of community support structures. Such nuanced, data-driven revelations underscore the unmatched potential of explainable AI to redefine public health paradigms.</p>
<p>The implications of this research extend beyond academic circles and into practical implementation. The transparent nature of the AI model makes it a viable tool for public health agencies seeking to deploy real-time surveillance systems. These systems could dynamically monitor shifts in risk factors, enabling prompt public health responses to emerging crises and potentially saving lives by guiding timely interventions tailored to specific at-risk groups.</p>
<p>Moreover, the authors emphasize the ethical dimensions of their work, noting that explainability is crucial for mitigating biases inherent in AI models, particularly when dealing with sensitive data surrounding violence and mental health. By providing clear rationale for predictions, the framework enhances accountability and supports the development of fair, culturally sensitive health policies that consider the diverse contexts across the Americas.</p>
<p>This study also represents a significant technical achievement in handling missing or incomplete data frequently encountered in public health databases. The AI approach incorporates advanced imputation techniques coupled with uncertainty quantification, ensuring robust performance without sacrificing interpretability. Such resilience enhances the model’s applicability to real-world settings, where data imperfections are the norm rather than the exception.</p>
<p>Crucially, the interdisciplinary collaboration underpinning this research—spanning computer scientists, epidemiologists, sociologists, and public health officials—reflects the complexity of tackling intentional injury mortality. This team approach facilitates a holistic understanding that integrates technical innovation with social and behavioral insights, making the explainable AI framework not just a predictive tool but a strategic asset for comprehensive health planning.</p>
<p>The study also sets a precedent for future AI applications in public health surveillance worldwide, highlighting the necessity of balancing cutting-edge machine learning capabilities with transparency and ethical rigor. As AI technologies increasingly permeate health systems, frameworks such as the one presented here will be vital in ensuring that these tools foster trust, inclusivity, and measurable health benefits.</p>
<p>In summary, the development of explainable AI for understanding and mitigating the intentional injury mortality crisis marks a transformative milestone. By rendering complex data intelligible and actionable, this approach empowers stakeholders to confront a deeply entrenched public health challenge with unprecedented clarity and precision. The hope is that such advancements will lead to significant reductions in suicide and homicide rates, ultimately saving lives and improving well-being across the Americas.</p>
<p>As this research progresses, ongoing validation of the model’s predictions and continuous ethical oversight will be essential to maximize its impact and safeguard against unintended consequences. The integration of community voices and feedback mechanisms will further enhance the cultural sensitivity and relevance of AI-driven interventions in diverse populations.</p>
<p>Looking ahead, the researchers envision expanding their framework to incorporate emerging data streams, such as social media sentiment and wearable health devices, which could provide earlier warning signals and enrich risk stratification. By continually refining explainability and predictive accuracy, these AI systems hold promise for revolutionizing public health surveillance on a global scale.</p>
<p>This paradigm shift towards transparent, data-driven health intelligence exemplifies the profound ways in which AI can be harnessed to address some of humanity’s most urgent challenges. The insights generated by this explainable AI framework not only advance scientific understanding but also offer a tangible pathway towards safer, healthier communities.</p>
<hr />
<p><strong>Subject of Research</strong>: Application of explainable artificial intelligence (AI) for public health surveillance focused on intentional injury mortality (suicide and homicide) in the Americas.</p>
<p><strong>Article Title</strong>: Explainable AI for public health surveillance: investigating the persistent crisis of intentional injury mortality (suicide and homicide) in the Americas.</p>
<p><strong>Article References</strong>:<br />
Kularathne, S., Rathnayake, N., Jayathilaka, R. <em>et al.</em> Explainable AI for public health surveillance: investigating the persistent crisis of intentional injury mortality (suicide and homicide) in the Americas. <em>Sci Rep</em> (2026). <a href="https://doi.org/10.1038/s41598-026-51327-y">https://doi.org/10.1038/s41598-026-51327-y</a></p>
<p><strong>Image Credits</strong>: AI Generated</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">161164</post-id>	</item>
	</channel>
</rss>
