Machine learning systems that classify objects into multiple categories at once — a task known as multi-label learning — are everywhere in modern science and technology. They tag photographs, annotate biomedical literature, predict the functions of proteins, and power recommendation engines. Many of the most popular algorithms in this family, such as the widely used ML-kNN and radial-basis-function kernel machines, make their decisions by measuring how similar one data point is to another in a high-dimensional feature space. Yet a team of researchers at the University of Sassari in Italy has now shown that these distance-based classifiers suffer from a subtle but consequential blind spot, and they have proposed an elegant fix that could reshape how practitioners prepare data for such models. Their method, called Paired Signed-Deviation Feature Selection, or PSDFS, is described in an open-access paper published in the journal Data Mining and Knowledge Discovery.
The blind spot, which the authors call direction-blindness, arises from a property most people never think about when they compute a distance. Standard Euclidean distance is symmetric: it depends only on the absolute magnitude of the difference between two feature values, not on whether one value sits above or below the other. That means a feature whose informative signal for a label lies entirely on one side of its average value — say, a gene whose under-expression flags a disease, or a word whose absence marks a document topic — is, from the perspective of pairwise distances, indistinguishable from a feature whose signal is symmetrically distributed around the mean. The classifier sees proximity, but not direction, and can collapse two semantically very different patterns into a single similarity score.
Adaptive metric-learning techniques such as LMNN and NCA can address this problem, but they do so inside the classifier, downstream of any feature selection step. The Sassari team, Filippo Casu, Andrea Lagorio and Giuseppe A. Trunfio, took a different route: expose directional evidence at the representation level, before the classifier ever runs. Their approach is an embedded feature-selection method, meaning it learns which features matter while simultaneously fitting a predictive model, rather than scoring features one at a time with a filter statistic. Embedded methods are attractive because they account for interactions among features and labels, and because they return a ranking that can be plugged into any downstream algorithm.
The core idea of PSDFS is a mathematical trick the authors call a signed-deviation lift. Every original feature is split into two nonnegative channels: one measuring how far a value sits above the training-fold mean, and one measuring how far it sits below. This two-sided hinge-style expansion, which echoes classical constructions in multivariate adaptive regression splines and in convex nonnegative matrix factorization, has a powerful consequence. In a conventional nonnegative reconstruction model, the contribution of a feature is a monotonically increasing function of its value, so low feature values can never act as explicit evidence for a label. The lift breaks that monotonicity barrier while keeping every weight nonnegative — a constraint that preserves the stability of multiplicative updates and keeps feature contributions easy to interpret. The researchers prove formally that the class of functions representable after lifting strictly contains the class representable without it.
Splitting each feature into two channels, however, creates a new hazard. If a sparsity penalty acts independently on the lifted components, the optimizer may select only one half of a pair — an XOR-like artifact the authors quantify with a diagnostic called the SplitRatio. Under an unpaired penalty, this ratio often approaches one, meaning the reported selection fragments the identity of the original feature and produces rankings that are sensitive to an arbitrary decomposition. PSDFS solves this with a paired group-sparsity penalty that treats the above- and below-center channels of each feature as a single selection unit, ranking the original features directly by the combined norm of their two channels. The result is a single, coherent ranking in the original feature space, with no split artifacts by construction.
The optimization machinery is deliberately conservative. For fixed training-fold statistics, the objective — a Frobenius reconstruction loss plus the paired group penalty — is convex in the weight matrix, and the multiplicative updates descend monotonically with a sublinear O(1/T) bound on average suboptimality, properties the authors establish in a series of propositions with full proofs. Empirically, most of the objective decrease happens within roughly ten iterations, and a budget of forty updates proved safe across every benchmark tested. Additional propositions show that the optimum carries genuine directional semantics: an active above-center weight requires positive correlation between the above-center deviation and the label residual, and vice versa, so the relative magnitudes of the two channels reveal which side of the mean carries the signal.
The experimental evaluation is unusually thorough. The team compared PSDFS against seven recent embedded feature-selection methods — LRMFS, LSMFS, LRDG, GRRO, RFSFS, SRFS and SCNMF — on sixteen benchmark datasets drawn from the Mulan repository and the Yahoo topic collection, spanning text, image and bioinformatics domains with anywhere from 194 to nearly 14,000 instances and up to 374 labels. All methods ran on identical stratified folds with no per-dataset hyperparameter tuning, a discipline that guards against information leakage. Evaluated with ML-kNN as the downstream classifier, PSDFS achieved the best average rank on all seven reported metrics, including Micro-F1, Macro-F1, Hamming Loss, precision-recall areas and label-ranking measures. After Holm–Bonferroni correction, its improvements were statistically significant against every baseline on Micro-F1 and Macro-F1, and against all but one baseline — GRRO — on Hamming Loss.
Targeted ablations confirm that each design element earns its place. Removing the lift collapses performance back toward the monotonicity-limited baseline, with the lifted variant winning on twelve to fourteen of sixteen datasets across the core metrics. Replacing the nonnegative constraint with an unconstrained solver degrades results on nearly every metric, suggesting that nonnegativity regularizes the solution and prevents the two channels from canceling each other algebraically. A robustness check with a second fixed-metric classifier, binary relevance with an RBF-kernel logistic regression, preserved the ranking of methods, reducing the risk that the advantage is an artifact of ML-kNN’s neighborhood bias. The method is also fast: at an average of 1.37 seconds per fold, it outruns most of its competitors, only trailing the two baselines that solve simpler linear systems.
Perhaps the most practically valuable analysis concerns low-resource regimes. When the training split is subsampled to as little as five percent, PSDFS degrades gracefully and is never overtaken by any baseline, though its separation from the strongest reconstruction-based competitors narrows to a statistical tie. The authors identify the mechanism precisely: the lift estimates twice as many channel weights from disjoint subsets of the data, so its contribution is the first casualty of scarce labels. They offer concrete deployment guidance — use the method as-is when the sample size exceeds roughly twice the lifted dimensionality, and treat fine-grained directional annotations with caution below that threshold. The authors also note two honest limitations: the model remains additive and does not capture higher-order feature interactions, and it has not been validated inside end-to-end deep visual pipelines. A reference implementation is publicly available on GitHub, and given how widely fixed-metric distance classifiers are deployed, the idea of making feature selection direction-aware seems likely to travel far beyond the sixteen benchmarks where it was born.
Subject of Research: Direction-aware feature selection for distance-based multi-label classification
Article Title: Direction-aware multi-label feature selection via paired signed-deviation lifting
Article References: Direction-aware multi-label feature selection via paired signed-deviation lifting. (n.d.). https://doi.org/10.1007/s10618-026-01265-0
Image Credits: AI Generated
DOI: 10.1007/s10618-026-01265-0
Keywords: multi-label learning, feature selection, machine learning, distance-based classifiers, nonnegative reconstruction, group sparsity, ML-kNN, signed-deviation lift, metric learning, data mining, benchmark evaluation, interpretable AI
Cite Scienmag News
Denise Maddox. (September 30, 2026). New AI Method Teaches Distance-Based Classifiers to Read Feature Direction. Scienmag. https://scienmag.com/new-ai-method-teaches-distance-based-classifiers-to-read-feature-direction/
Denise Maddox. "New AI Method Teaches Distance-Based Classifiers to Read Feature Direction." Scienmag, 30 September 2026, https://scienmag.com/new-ai-method-teaches-distance-based-classifiers-to-read-feature-direction/. Accessed 30 September 2026.
Denise Maddox. "New AI Method Teaches Distance-Based Classifiers to Read Feature Direction." Scienmag. September 30, 2026. https://scienmag.com/new-ai-method-teaches-distance-based-classifiers-to-read-feature-direction/








