Outlier detection has become one of the quiet workhorses of the modern data economy. Every time a bank flags a suspicious transaction, a security system isolates an anomalous network event, or a quality-control pipeline catches a defective product, an algorithm is quietly deciding that a particular record does not belong. Yet a new large-scale study from the University of São Paulo suggests that for a huge class of real-world data, the field has been relying on guesswork. The research, published in Artificial Intelligence Review by Felippe Pires Ferreira and Robson L. F. Cordeiro of the Institute of Mathematics and Computer Sciences, offers the most systematic answer yet to a deceptively simple question: when your dataset contains categorical attributes, how should you actually detect the outliers?
The problem is more consequential than it might first appear. Most outlier detection algorithms were designed with numerical data in mind. They measure distances, densities, and deviations in spaces where every attribute is a number, and where concepts like proximity are mathematically straightforward. But real datasets are rarely so cooperative. Medical records contain diagnoses and blood types. Customer databases contain product categories, subscription tiers, and regions. Fraud datasets mix transaction amounts with merchant categories and card types. In these settings, the numerical machinery of classical anomaly detection either breaks down or silently distorts the data, and practitioners have been left to improvise without solid evidence about which improvisation works.
Ferreira and Cordeiro frame the field’s improvisations as three competing strategies. The first is to apply algorithms that can process categorical data directly, respecting the discrete nature of attributes such as color, brand, or diagnosis. The second is to convert categorical attributes into numerical ones before detection, transforming the data into a form that standard numerical detectors can consume. The third is the bluntest instrument: simply remove the categorical attributes and run detection on the numerical columns alone. Each strategy has intuitive appeal and intuitive drawbacks, but until now there has been no rigorous, large-scale comparison of how they actually perform across diverse datasets and detectors.
To settle the question, the researchers mounted an unusually comprehensive experimental campaign. They assembled 47 datasets and ran 14 different outlier detection algorithms across them, evaluating all three strategic approaches under controlled conditions. This scale matters. Many published comparisons in machine learning rest on a handful of benchmark datasets, which makes it dangerously easy to overfit conclusions to particular data quirks. By sweeping across dozens of datasets with varied characteristics, the study could distinguish between strategies that win consistently and strategies that win only under specific conditions, a distinction that turns out to be central to the paper’s findings.
The headline result is that the first strategy, applying algorithms that handle categorical data natively, is usually the preferred choice. And within that strategy, one detector stood out from the pack: CBRW, an algorithm that approaches categorical outliers through a weighted-graph perspective on attribute values. When datasets contain categorical attributes, the evidence indicates that reaching for a purpose-built categorical detector, particularly CBRW, is the most reliable default. This finding carries real practical weight, because it suggests that the common habit of forcing categorical data through numerical pipelines may often be an unnecessary detour that costs detection quality.
But the study resists the temptation of a one-size-fits-all verdict. The second strategy, converting categorical attributes into numbers before detection, proved superior in certain contexts, and the researchers found that which contexts depends on the characteristics of the data itself. Notably, when this conversion route was taken, detectors such as Isolation Forest, known as iForest, and KNN-outlier achieved better results in those favorable settings. Isolation Forest works by randomly partitioning the data and isolating records that are easy to separate from the rest, while KNN-outlier flags points that sit far from their nearest neighbors. Both are staples of the numerical anomaly detection toolkit, and the study shows they can still earn their keep on mixed data, provided the categorical-to-numerical conversion is done well.
That proviso led the authors to a second, finer-grained comparison: among the many techniques for converting categorical attributes into numerical representations, which one should practitioners choose? Here again the study produced a clear signal. Correspondence Analysis, a classical statistical method that maps categorical variables into a low-dimensional numerical space while preserving the associations between attribute values, often yielded the best detection results among the conversion approaches evaluated. For teams committed to running numerical detectors on mixed data, this finding offers a concrete, evidence-backed recommendation in place of an arbitrary choice of encoding technique.
Perhaps the most forward-looking contribution goes beyond ranking the strategies. Because the best approach depends on data characteristics, the researchers built a predictive model, a meta-level classifier trained on their experimental findings, that examines a new dataset and predicts which of the three strategies is most likely to perform best. The model achieves 80 percent accuracy in identifying the winning strategy. In a field where practitioners often have no principled basis for choosing an approach, an 80 percent reliable recommendation engine represents a meaningful step toward automated, data-driven method selection, analogous to the meta-learning systems that have begun to guide algorithm choice in supervised learning.
The study also delivers a quieter but equally important negative result: the third strategy, discarding categorical attributes altogether, does not emerge as a recommended path. Throwing away information is the kind of shortcut that survives in practice because it is easy, and part of the value of a systematic comparison is that it can retire such shortcuts with evidence. The findings collectively suggest that the information carried by categorical attributes is genuinely useful for anomaly detection, whether consumed directly by a categorical-aware algorithm or preserved through a careful numerical conversion such as Correspondence Analysis.
For the broader machine learning community, the significance of this work lies in its methodological honesty. Rather than crowning a single universal champion, it maps the conditions under which each strategy excels and then builds a bridge from those conditions to practical recommendations for new data. The authors have also committed fully to reproducibility: all code, detailed results, parameter values tested, and datasets used in the study are freely available for download on GitHub, allowing other researchers to scrutinize, extend, or challenge the conclusions. Supported in part by the Brazilian agency CAPES and published open access, the survey arrives as both a practical handbook for anyone deploying anomaly detection on messy real-world data and a benchmark against which future categorical outlier detection methods will now have to be measured. As fraud, cybersecurity, and data-quality applications continue to multiply across industries dominated by categorical information, the question this study answers is no longer an academic curiosity. It is an operational decision worth getting right, and for the first time, practitioners have a large-scale evidence base to guide it.
Subject of Research: Comparative evaluation of outlier detection methods for categorical and mixed data
Article Title: A comparative evaluation of outlier detection in categorical and mixed data
Article References: Ferreira, F. P., & Cordeiro, R. L. F. (2026). A comparative evaluation of outlier detection in categorical and mixed data. Artificial Intelligence Review. https://doi.org/10.1007/s10462-026-11687-3
Image Credits: AI Generated
DOI: 10.1007/s10462-026-11687-3
Keywords: outlier detection, categorical data, mixed data, anomaly detection, CBRW, Isolation Forest, KNN-outlier, Correspondence Analysis, data conversion, machine learning, data mining, benchmark evaluation
Cite Scienmag News
Blake Davidson. (October 2, 2026). New Study Reveals the Best Way to Spot Outliers in Categorical Data. Scienmag. https://scienmag.com/new-study-reveals-the-best-way-to-spot-outliers-in-categorical-data/
Blake Davidson. "New Study Reveals the Best Way to Spot Outliers in Categorical Data." Scienmag, 2 October 2026, https://scienmag.com/new-study-reveals-the-best-way-to-spot-outliers-in-categorical-data/. Accessed 2 October 2026.
Blake Davidson. "New Study Reveals the Best Way to Spot Outliers in Categorical Data." Scienmag. October 2, 2026. https://scienmag.com/new-study-reveals-the-best-way-to-spot-outliers-in-categorical-data/

