From the DNA of a bacterium to the streets of a sprawling metropolis, from the vocabulary of a novel to the modules of a software project, an astonishing range of complex systems appears to obey strikingly similar statistical laws. Word frequencies in texts, gene family sizes in genomes, and population distributions across cities all tend to follow heavy-tailed patterns that look suspiciously alike, no matter how different the underlying systems are. A new theoretical study published in PLOS Complex Systems by Andrea Mazzolini, Mattia Corigliano, Rossana Droghetti, Matteo Osella, and Marco Cosentino Lagomarsino asks a provocative question: if the same patterns show up everywhere, are they really telling us something deep about each system, or are they simply what mathematics forces any collection of modular parts to look like?
The researchers focus on what they call component systems: ensembles of realizations built from a shared repertoire of modular parts. A genome is assembled from genes, a book from words, a city from neighborhoods and functions, a codebase from software modules. When scientists count how often each part appears, how many distinct parts show up, and how these quantities scale with system size, the resulting curves often satisfy what physicists would recognize as bona fide laws: they are simple, general, and remarkably robust across domains. Heaps’ law, Zipf’s law, and related scaling relationships have been documented for decades in linguistics, genomics, ecology, and urban science, and their persistence has encouraged many researchers to interpret them as fingerprints of specific generative mechanisms.
The core argument of the new work is more sobering, and more interesting. The very generality and simplicity of these laws may be a consequence of basic combinatorial or sampling constraints rather than evidence of system-specific dynamics. When you draw parts from a finite repertoire and aggregate them into larger and larger ensembles, certain statistical trends are almost inevitable. The authors suggest that many of the celebrated regularities observed across component systems fall into this category: they are null trends, patterns that would emerge even in the absence of any particular biological, cultural, or technological mechanism. This reframing echoes a long-standing methodological debate in ecology, where null models have been used for decades to test whether observed community patterns exceed what random assembly would produce.
To make this argument precise, the team developed a unifying mathematical framework that allows modular systems from different fields to be compared on common ground. The framework treats each system as a collection of components drawn from a shared vocabulary and characterizes the ensemble through its counting statistics: how the number of distinct parts grows with the number of sampled elements, how part frequencies are distributed, and how fluctuations behave across realizations. By expressing these quantities in a shared formal language, the researchers can separate the contributions that any component system would display from those that are genuinely distinctive to genomes, texts, cities, or software.
The practical payoff of this framework is twofold. First, it explains why the general regularities emerge in the first place, identifying the constraints that link them together and the general principles from which they originate. Rather than treating each scaling law as an independent empirical discovery, the framework shows how they are mathematically intertwined: knowing one trend often fixes the others, given only weak assumptions about sampling. This kind of constraint analysis is familiar from statistical mechanics, where macroscopic regularities such as universal equations of state follow from counting microscopic configurations rather than from any detailed knowledge of the interactions involved.
Second, and perhaps more consequentially, the framework offers a way to subtract the null trends from real data. Once the generic combinatorial background is removed, what remains are the deviations: the informative features that cannot be explained by sampling alone. These residuals, the authors argue, are the true signatures of the underlying generative dynamics, carrying information about hidden mechanistic and causal processes. In a genome, such deviations might reflect the action of natural selection, duplication events, or functional organization. In a text, they might reveal stylistic choices, topic structure, or the idiosyncrasies of an author. In a city, they could point to planning decisions, economic forces, or historical contingencies that no neutral aggregation process would produce.
This subtraction strategy connects the work to two of the most powerful toolkits in modern quantitative science: statistical mechanics and machine learning. Statistical mechanics provides the conceptual apparatus for defining ensembles, computing expectations under null hypotheses, and quantifying how unlikely an observed configuration is under those hypotheses. Machine learning, meanwhile, supplies flexible generative models that can be trained on real data and then interrogated for the structure they capture beyond the null. The authors envision using these tools in combination: first establishing what a generic component system must look like, then deploying learned models to isolate and interpret the departures from that baseline.
The implications reach across disciplines. In genomics, where the statistics of gene families and functional categories have been studied for decades, the framework offers a principled way to decide which patterns demand biological explanation and which are artifacts of counting. In linguistics and the digital humanities, it provides a rigorous baseline for claims about authorship, complexity, and cultural evolution. In urban science and technology studies, it can help distinguish the generic scaling of cities and software systems from the specific interventions of planners and engineers. In every case, the message is the same: a pattern must first be compared against what combinatorics alone would produce before it can be credited to a mechanism.
The study also carries a cautionary lesson for the popular narrative that big data reveals universal laws of complexity. Universality can be seductive, but it can also be cheap: patterns that appear in everything may explain nothing in particular. By making the null expectations explicit and quantitative, the new framework turns that criticism into a research program. The goal is not to dismiss the regularities but to use them as a filter, sharpening the search for the features that genuinely distinguish one system from another and that therefore encode the physics, biology, or culture of that system.
What emerges is a methodological vision for the study of complex systems in the era of abundant data. Component systems everywhere share a common statistical skeleton imposed by the mathematics of parts and ensembles; the science of interest lies in the flesh that grows around that skeleton. By combining a unifying null framework with statistical mechanics and modern machine learning, Mazzolini and colleagues propose a disciplined path from pattern to process: explain the general, subtract it, and then read what remains as evidence of the hidden generative dynamics that make a genome a genome, a text a text, and a city a city.
Subject of Research: Null models and statistical regularities in component systems across biological, ecological, technological, and socio-cultural domains
Article Title: Component systems: Do null models explain everything?
Article References: Mazzolini, A., Corigliano, M., Droghetti, R., Osella, M., & Cosentino Lagomarsino, M. (2026). Component systems: Do null models explain everything?. PLOS Complex Systems, 3(6), e0000114. https://doi.org/10.1371/journal.pcsy.0000114
Image Credits: AI Generated
DOI: 10.1371/journal.pcsy.0000114
Keywords: component systems, null models, complex systems, statistical mechanics, scaling laws, genomics, linguistics, urban science, machine learning, combinatorics, heavy-tailed distributions, generative processes
Cite Scienmag News
Reid Dalton. (October 9, 2026). Null Models May Explain the Laws of Genomes, Cities, and Language. Scienmag. https://scienmag.com/null-models-may-explain-the-laws-of-genomes-cities-and-language/
Reid Dalton. "Null Models May Explain the Laws of Genomes, Cities, and Language." Scienmag, 9 October 2026, https://scienmag.com/null-models-may-explain-the-laws-of-genomes-cities-and-language/. Accessed 9 October 2026.
Reid Dalton. "Null Models May Explain the Laws of Genomes, Cities, and Language." Scienmag. October 9, 2026. https://scienmag.com/null-models-may-explain-the-laws-of-genomes-cities-and-language/

