For decades, the world’s most widely cited democracy indices have been built on a narrow foundation: whether people can vote freely, whether parties can compete, whether courts and electoral bodies act independently. Political variables have dominated the measurement of democratic quality so thoroughly that the possible contribution of everything else — education, health, energy, economic structure — has remained largely unexamined. A new study published in SN Social Sciences by researchers at the IMDEA Software Institute, IMDEA Networks Institute and the Universidad de Los Andes now challenges that orthodoxy with an unusually rigorous machine learning analysis, and its central finding is striking: socioeconomic and contextual variables alone can predict democracy classifications about as well as, and in some cases better than, the political indicators that define the indices themselves.
The research team, led by Diego Benito, Jose Aguilar, Juan Marcos Ramírez and Antonio Fernández Anta, set out to answer a deceptively simple question. If democracy indices are supposed to capture the quality of democratic life in a country, how much of that quality is actually encoded in the broader social and economic conditions in which democracies operate? Rather than running a handful of regressions, the group constructed classification models for three distinct democracy indices using multiple machine learning techniques and three different dataset configurations: one containing only political variables, one combining political and context variables, and one built exclusively from context variables such as education, health and social development indicators drawn from open repositories including the V-Dem dataset, Our World in Data and World Bank Open Data.
The methodological design is where the study distinguishes itself. The authors trained classifiers on each configuration and then subjected them to a battery of cutting-edge explainability methods: Local Interpretable Model-Agnostic Explanations (LIME), which probes individual predictions by perturbing inputs; SHapley Additive exPlanations (SHAP), which distributes credit for each prediction among features using game-theoretic principles; and Feature Importance Based on Random Values (FIRV), which benchmarks feature scores against randomly generated variables to filter out spurious attributions. Crucially, the team added a fourth layer of validation using the RemOve And Retrain (ROAR) technique, in which the supposedly most important features are removed and the model is retrained from scratch. If performance collapses, the explanation was real; if it barely changes, the explanation was likely noise. This combination of explanation and adversarial validation is still rare in social science applications of machine learning.
The results were consistent across all three democracy indices studied. Adding context variables to the political ones consistently improved or at least matched classification performance, meaning the models could distinguish more democratic from less democratic countries more reliably when socioeconomic information was included. More provocatively, models trained on context variables alone achieved predictive performance comparable to — and in some cases superior to — models built exclusively from the political variables that traditionally define these indices. In practical terms, a machine learning model that never sees a single measure of electoral freedom or party competition can still classify countries on a democracy index with remarkable accuracy, simply by examining patterns in education, health, energy and development data.
The explainability analyses then went a step further, identifying which specific context variables carried the most weight. Among the most influential was renewable energy consumption, which emerged as a highly important feature for both the Electoral Democracy Index and the Political Corruption Index. That finding is likely to raise eyebrows in both political science and energy policy circles: it suggests a measurable statistical association between a country’s transition toward renewable energy and its standing on indices of democratic quality and corruption, independent of the political variables usually used to explain those standings. The authors are careful to frame such relationships as insights into the structure of the data rather than proof of causation, but the associations offer new hypotheses about how socioeconomic conditions and democratic quality intertwine.
The study builds on a growing but still contested literature. Previous work had used machine learning to construct new democracy measures, including datasets covering 186 countries from 1919 onward, and a substantial body of research has explored links between democracy and outcomes ranging from economic growth and health to emissions, energy efficiency, public media funding and transparent lobbying. What most of that work shared was a framing in which democracy was the explanatory variable and social outcomes were the consequences. The new analysis inverts the perspective, asking instead whether the social and economic context can reconstruct the democracy measurement itself — and finding that it largely can.
Technically, the researchers faced the standard obstacles of cross-national social data: missing values, class imbalance and the risk of overfitting. Their pipeline drew on established remedies, including synthetic minority over-sampling (SMOTE) to balance classification categories, hyperparameter optimization with the Optuna framework, and careful treatment of missing data following best practices in statistical analysis. Variance inflation factors were used to guard against multicollinearity among features. By applying the same disciplined pipeline across three different indices and three dataset configurations, the team could check whether any single result was an artifact of one index’s particular construction — and the consistency of the findings across indices is one of the study’s strongest points.
The implications extend beyond academic measurement. Democracy indices are not just descriptive instruments; they inform country risk assessments, development aid allocations, governance conditionality and countless comparative studies. If a substantial share of the information in those indices is statistically recoverable from socioeconomic context, then index designers may need to think harder about what their measures actually capture and whether political and social dimensions are being conflated. The authors argue that their findings provide a basis for new data-driven approaches to understanding and strengthening democracy, and specifically for incorporating relevant context variables into the development of future democracy indices and, more importantly, into the design of evidence-based public policies.
There are, of course, limits to what classification models can establish. Machine learning identifies predictive structure, not causal mechanisms, and the study itself does not claim that investing in renewable energy or education will mechanically raise a country’s democracy score. The direction of causality between social development and democratic quality has been debated for generations, and the honest reading of this work is that the two are so deeply entangled that political variables alone were never going to tell the full story. What the study does deliver is a transparent, validated demonstration of that entanglement, using explainability methods that allow other researchers to inspect exactly which features drive which predictions rather than accepting a black box.
For a field that has measured democracy almost exclusively through political lenses for the better part of a century, the message of this research is quietly radical: the context in which democracies exist is not background noise but signal. Education systems, health outcomes, energy choices and broader social development appear to encode enough information to reproduce the judgments of the world’s democracy indices — and sometimes to do so better than the political indicators themselves. As machine learning and explainability tools continue to mature, the study suggests that the next generation of democracy measurement may look less like a checklist of institutional features and more like a multidimensional portrait of the societies in which those institutions either flourish or fail. All data underlying the work comes from open repositories, which means other teams can immediately test, replicate and extend these findings — a fittingly democratic way to advance the science of measuring democracy itself.
Subject of Research: Machine learning analysis of how socioeconomic context variables influence democracy indices
Article Title: Analysis of the impact of context variables on democracy indices using machine learning approaches
Article References: Benito, D., Aguilar, J., Ramírez, J. M., & Anta, A. F. (2026). Analysis of the impact of context variables on democracy indices using machine learning approaches. SN Social Sciences, 6(10), Article 508. https://doi.org/10.1007/s43545-026-01802-0
Image Credits: AI Generated
DOI: 10.1007/s43545-026-01802-0
Keywords: democracy indices, machine learning, explainability, SHAP, LIME, ROAR, political science, socioeconomic variables, renewable energy, classification models, V-Dem, public policy
Cite Scienmag News
Courtney Benton. (October 7, 2026). Machine Learning Reveals That Schools, Health and Energy Shape Democracy Scores. Scienmag. https://scienmag.com/machine-learning-reveals-that-schools-health-and-energy-shape-democracy-scores/
Courtney Benton. "Machine Learning Reveals That Schools, Health and Energy Shape Democracy Scores." Scienmag, 7 October 2026, https://scienmag.com/machine-learning-reveals-that-schools-health-and-energy-shape-democracy-scores/. Accessed 7 October 2026.
Courtney Benton. "Machine Learning Reveals That Schools, Health and Energy Shape Democracy Scores." Scienmag. October 7, 2026. https://scienmag.com/machine-learning-reveals-that-schools-health-and-energy-shape-democracy-scores/

