For nearly half a century, economists and operations researchers have relied on a deceptively simple question to judge how well organizations use their resources: given the inputs a firm consumes, how close does it come to producing the maximum feasible output? The dominant tool for answering that question, Data Envelopment Analysis, or DEA, has been applied to hospitals, banks, schools, airlines and entire national industries. But DEA has a well-known weakness. It is deterministic, meaning it delivers a blunt verdict, efficient or inefficient, based on a frontier constructed from the observed data itself. Now a team of Spanish researchers has unveiled an open-source software package that reimagines this classic technique through the lens of modern machine learning, converting efficiency measurement into a probability statement that comes with its own explanation.
The package, called PEAXAI, short for Probabilistic Efficiency Analysis using Explainable Artificial Intelligence, was developed by Ricardo González-Moyano, Juan Aparicio, José L. Zofío and Víctor J. España and described in the journal SoftwareX. Written in R and freely available under the GPL-3 license on CRAN and GitHub, the software implements a framework the authors introduced in a companion methodological paper. Its central idea is elegant: instead of asking whether a decision-making unit, or DMU, sits exactly on the efficient frontier, PEAXAI trains supervised classifiers to estimate the probability that a unit belongs to the efficient class. That single shift transforms efficiency analysis from a deterministic exercise into an uncertainty-aware one, allowing managers to set benchmarks at whatever confidence level they consider appropriate.
The workflow begins in familiar territory. PEAXAI first estimates a weighted additive DEA model, using the range-adjusted measure that normalizes slacks by the observed spread of each input and output. Units whose objective function evaluates to zero, meaning no input can be reduced or output expanded without worsening something else, are labeled efficient; all others are labeled not efficient. These binary labels then become the training data for machine learning classifiers, which can include neural networks, support vector machines and other algorithms implemented through the widely used caret infrastructure. Crucially, the package evaluates candidate models with stratified cross-validation in which the withheld observations are labeled by a frontier built only from the training subset, so the model’s generalization is tested against a technology it has never seen.
One of the thorniest practical problems the developers had to solve is class imbalance. In real datasets, the efficient frontier is typically sparsely populated. In the package’s own illustrative application to 1,267 meat-industry firms in Spain and Portugal, drawn from the SABI balance-sheet database, DEA under variable returns to scale flagged only 16 firms, roughly 1.26 percent of the sample, as efficient. A classifier trained on such lopsided data would struggle to learn where the boundary lies. PEAXAI addresses this with an adapted version of SMOTE, the synthetic minority over-sampling technique, redesigned to respect the polyhedral geometry of the DEA technology. When efficient units are scarce, synthetic firms are generated as convex combinations lying on maximal-dimensional facets of the estimated frontier, so they remain genuinely efficient. When inefficient units are scarce, synthetic observations are placed strictly inside the technology, close to but not on the frontier.
Once the classifiers are trained, the package turns to explainable artificial intelligence to open the black box. PEAXAI offers both global and local variable-importance tools, drawing on aggregated SHAP values, permutation feature importance and sensitivity analysis for the global picture, and local SHAP values, LIME explanations and sensitivity analysis for unit-specific insight. The practical payoff is that analysts can identify which inputs and outputs most strongly drive the probability of frontier membership, both across the whole sample and for any individual firm. In the meat-industry example, where personnel expenses and fixed assets serve as proxies for labor and capital and operating income is the output, such explanations can reveal whether workforce costs or capital stock dominate the efficiency classification, information that conventional DEA simply does not provide.
The most striking feature of the package may be its counterfactual machinery. Rather than reporting an abstract inefficiency score, PEAXAI quantifies inefficiency as the minimum adjustment a firm must make to cross the probabilistic efficiency boundary. Using the directional distance function, a framework that allows simultaneous input reduction and output expansion along a user-chosen direction, the software searches for the smallest change that pushes a firm’s predicted efficiency probability above a chosen threshold. The direction itself can be weighted by the data-driven variable-importance results, so the improvement path reflects the factors that actually matter to the classifier. Managers can thus ask a concrete question: how much would we need to cut costs or grow revenue to have, say, a 90 percent chance of being classified efficient?
That question has direct economic weight. Because the directional distance function permits improvements on both the input and output sides, the resulting targets translate into simultaneous cost savings and revenue gains rather than indiscriminate austerity. The authors argue this should caution managers against reflexive cost-cutting; in their example, expanding operating income may matter as much as trimming personnel expenses or fixed assets. The package complements this with proximity-based peer selection, pairing each firm with the observed high-probability unit closest to its counterfactual target, either by plain Euclidean distance or by a distance weighted by feature importance. The result is a benchmark that is not merely statistically nearest but managerially relevant, a comparable firm whose practices could plausibly be emulated.
Ranking is also rethought. Traditional DEA rankings rest on a single deterministic score, which can be unstable in high-dimensional settings where the curse of dimensionality erodes the method’s discriminating power and inflates efficiency estimates. PEAXAI offers two ordering schemes: one based purely on predicted efficiency probability, and a hierarchical attainable ranking that prioritizes the maximum probability a unit could reach, then the magnitude of the required input reductions and output expansions, and finally the current predicted probability. This blends where a firm stands today with how far it could plausibly climb, producing a prioritization that is more informative for policy and sectoral planning than a static league table.
The developers are candid about limitations. The DEA labeling stage requires both classes to be populated, which is precisely why the synthetic balancing step exists, and identifying all facets of the estimated technology becomes the computational bottleneck for very large or high-dimensional datasets. The package also does not yet accept fuzzy or interval-valued data, and the authors stress that the predicted probability reflects classifier uncertainty about the DEA label, not the true probability of technical efficiency, so users should rely on the outputs only when discrimination and calibration metrics are satisfactory. Future work, they suggest, includes post-hoc probability calibration, contextual environmental variables, robust frontier labels from order-m and order-alpha models, and faster counterfactual algorithms. Even so, by uniting frontier estimation, machine learning, explainability and counterfactual target setting in a single reproducible R workflow, PEAXAI signals where efficiency analysis is heading: away from verdicts and toward quantified, explained and actionable probabilities.
Subject of Research: Probabilistic technical efficiency analysis combining data envelopment analysis, machine learning classifiers and explainable AI in R
Article Title: PEAXAI: A probabilistic efficiency analysis using explainable artificial intelligence in R
Article References: González-Moyano, R., Aparicio, J., Zofío, J. L., & España, V. J. (2026). PEAXAI: A probabilistic efficiency analysis using explainable artificial intelligence in R. SoftwareX, 36, Article 103089. https://doi.org/10.1016/j.softx.2026.103089
Image Credits: AI Generated
DOI: 10.1016/j.softx.2026.103089
Keywords: data envelopment analysis, machine learning, explainable AI, R package, technical efficiency, SHAP values, counterfactual explanations, directional distance function, benchmarking, class imbalance, SMOTE, open-source software
Cite Scienmag News
Denise Maddox. (October 3, 2026). New R Package Turns Efficiency Analysis Into a Probability Problem AI Can Explain. Scienmag. https://scienmag.com/new-r-package-turns-efficiency-analysis-into-a-probability-problem-ai-can-explain/
Denise Maddox. "New R Package Turns Efficiency Analysis Into a Probability Problem AI Can Explain." Scienmag, 3 October 2026, https://scienmag.com/new-r-package-turns-efficiency-analysis-into-a-probability-problem-ai-can-explain/. Accessed 3 October 2026.
Denise Maddox. "New R Package Turns Efficiency Analysis Into a Probability Problem AI Can Explain." Scienmag. October 3, 2026. https://scienmag.com/new-r-package-turns-efficiency-analysis-into-a-probability-problem-ai-can-explain/

