A new study published in the open-access journal Heliyon offers one of the most detailed statistical portraits yet of how prepared Europe’s regions are for the fourth industrial revolution, and it does so by borrowing a technique from an unexpected place: genomics. A team of Hungarian researchers led by Z.T. Kosztyán of the University of Pannonia combined biclustering, a method originally developed to find patterns in gene-expression data, with classic spatial autocorrelation analysis to sort 320 European regions into what they call industrial leagues. The result is a framework that tells policymakers not only which regions share the same strengths and weaknesses, but also whether those shared problems are geographically concentrated or scattered across the continent.
The analysis covers the EU-27 member states plus Norway and Iceland from the European Economic Area, along with NUTS-compatible statistical regions in Switzerland, Turkey, the United Kingdom, Montenegro, North Macedonia, and Serbia. The researchers worked at the NUTS 2 level, the standardized territorial units of roughly 800,000 to 3 million inhabitants that form the backbone of EU cohesion policy. Their raw material was the I4.0+ indicator system, a dataset of 101 indicators compiled from European statistical sources between 2008 and 2019, spanning education levels, youth employment, Erasmus student mobility, graduate specialization, scientific and technological capacity, and even the tone of media coverage of industry drawn from the GDELT database.
Because the underlying data were incomplete, the team faced a substantial data-cleaning challenge. Nearly 32 percent of the values in the initial 2018 table were missing, so the researchers applied a temporally ordered substitution rule, filling gaps with the nearest available year’s observation. After screening out regions and indicators with excessive missingness, the final analysis matrix contained 71 indicators for 320 regions, with only 3.26 percent of cells still unavailable. The authors are transparent about the limits of these choices, noting that zero-encoding of residual gaps could affect marginal league memberships, particularly in the low-performance group, and they have released a full audit identifying every affected cell.
The methodological core of the study is a two-stage sequence. First, seriation reorders the region-by-indicator matrix so that similar cells sit close together, and biclustering algorithms then carve out homogeneous submatrices, each pairing a set of regions with a set of indicators on which those regions perform similarly. These submatrices are the industrial leagues. League A captures comparatively high-performing region-indicator combinations, League C captures low-performing ones, and League B, identified by a variance-minimizing algorithm called BicARE, captures an intermediate group of emerging regions. Crucially, the biclustering is deliberately blind to geography. Only in the second stage does the team test whether league membership is spatially clustered, using Moran’s I, a statistic that measures whether nearby regions have more similar values than expected by chance.
This separation is what gives the framework its policy punch. Spatially constrained clustering methods, which force only adjacent regions into the same group, cannot distinguish a genuinely concentrated pattern from a dispersed one. By first identifying leagues without spatial constraints and then testing their spatial dependence, the researchers can tell a policymaker whether a problem calls for a place-based intervention, such as a regional innovation hub, or a sector-wide program, such as a pan-European education initiative. The authors formalize this as a two-part decision rule: league membership identifies which regions and indicators need attention, and the corrected Moran evidence determines the appropriate spatial scale of action.
The results reveal a striking asymmetry. League A, the high-performance group, shows strong and statistically significant spatial clustering across most threshold settings, with Moran statistics between roughly 0.40 and 0.57 after false-discovery-rate correction. Employment and population-education indicators dominate this league and carry the highest global autocorrelation values in the entire dataset, with education indicators P01 and P02 topping the list at 0.7716. This is the familiar core-periphery geography of European development: advantages concentrate in places like Inner London, Bratislava, and Prague, which emerged as the highest-performing capital regions, and spill over to their neighbors.
League C tells a very different story. Its regional membership pattern is spatially diffuse, failing significance tests under every nonconstant threshold setting, even though many of its component indicators, particularly in science and education, still show significant individual spatial structure. The low-performing league draws disproportionately from Turkey, Greece, North Macedonia, southern Italy, and northern England, and its characteristic weaknesses include the number of industry-related research studies, R&D personnel in the business sector, research institutes, and their reflection in media coverage. In other words, the continent’s innovation deficit is not a regional problem in the traditional sense; it is a shared weakness distributed across distant and dissimilar places, which argues for broad framework programs rather than geographically targeted ones, supplemented where indicator-level evidence warrants local action.
Perhaps the most counterintuitive finding concerns overlap. At the baseline threshold, 87.5 percent of regions belong to both the high-performing and low-performing leagues simultaneously, meaning most European regions excel on some indicators while lagging on others. Only four capital regions sit exclusively in the top league. Indicators, by contrast, are never assigned to both League A and League C, a mutual exclusivity built into the method’s construction. This region-indicator distinction dissolves the simplistic picture of rich and poor regions and replaces it with a more granular one in which nearly every territory has a specific, identifiable pathway of indicators to improve in order to advance to a higher league.
The timing of such a framework is significant. The EU’s Recovery and Resilience Facility is channeling 723.8 billion euros into member states with explicit mandates for digital and green transitions, and the effectiveness of that spending depends on knowing which regions can absorb and deploy digital-transformation investments. The authors emphasize that their 2018 snapshot should be read as a historical baseline rather than a current ranking, since structural gaps in digital infrastructure, R&D capacity, and workforce skills tend to persist over multiyear periods. The principal contribution is methodological and reusable: the biclustering-plus-spatial-testing pipeline can be reapplied as newer I4.0+ data become available.
The study is candid about its limitations, including the unrecorded random seed of the original stochastic runs, the unequal country representation of between 1 and 41 regions, and the fact that the 320 regions, not the 22,720 region-indicator cells, are the true statistical units. The spatial weight matrix required deterministic fallback connections for five remote regions, including an unusually long edge linking the French overseas territory of Réunion to Cyprus. Yet the released reproduction archive, with archived memberships, seed manifests, and full spatial diagnostics, sets a high bar for transparency. For regional development agencies, the message is practical: benchmark against comparable regions in your league, identify the specific indicators that gate advancement, and let the geography of the evidence, not administrative habit, decide whether the remedy should be local, continental, or both.
Subject of Research: A combined biclustering and spatial autocorrelation framework for assessing Industry 4.0 readiness across European NUTS 2 regions
Article Title: Combining spatial and aspatial methods to identify European industrial leagues: Development perspectives
Article References: Kosztyán, Z., Banász, Z., Kurbucz, M., Katona, A., Pribojszki-Németh, A., & Jakobi, Á. (2026). Combining spatial and aspatial methods to identify European industrial leagues: Development perspectives. Heliyon, 12(15), Article e45501. https://doi.org/10.1016/j.heliyon.2026.e45501
Image Credits: AI Generated
DOI: Not provided
Keywords: Industry 4.0, regional development, biclustering, spatial autocorrelation, Moran's I, NUTS 2 regions, European Union, cohesion policy, I4.0+ indicators, industrial leagues, digital transformation, core-periphery
Cite Scienmag News
Drew Townsend. (October 6, 2026). New Mapping Method Sorts Europe’s Regions Into Industrial Leagues. Scienmag. https://scienmag.com/new-mapping-method-sorts-europes-regions-into-industrial-leagues/
Drew Townsend. "New Mapping Method Sorts Europe’s Regions Into Industrial Leagues." Scienmag, 6 October 2026, https://scienmag.com/new-mapping-method-sorts-europes-regions-into-industrial-leagues/. Accessed 6 October 2026.
Drew Townsend. "New Mapping Method Sorts Europe’s Regions Into Industrial Leagues." Scienmag. October 6, 2026. https://scienmag.com/new-mapping-method-sorts-europes-regions-into-industrial-leagues/

