Where does one bacterial species end and another begin? For most of the twentieth century, microbiologists answered that question with petri dishes and biochemical test strips, watching how organisms fermented sugars or reacted to stains. The genomic era promised something sharper: a hard numerical threshold that would carve the microbial world into clean, defensible units. A new study from researchers at Lodz University of Technology in Poland shows just how slippery that promise remains, even for some of the best-studied bacteria on Earth.
In research published in BMC Genomics, Tomasz Grzyb, Małgorzata Wlaźlak and Justyna Szulc took six bacterial strains isolated from soils in post-maize cultivation fields and subjected them to a battery of genome-based species assignment methods. Their targets belonged to the order Caryophanales, a group that includes the enormously consequential genera Bacillus, Priestia and Paenibacillus, organisms used in agriculture, industry and biotechnology, and close relatives of dangerous pathogens. What they found was a taxonomic landscape riddled with contradiction: species boundaries that shift depending on which metric you trust, genomes that intermix across supposedly distinct species names, and a widely used identity threshold that may be significantly miscalibrated.
The team compared several complementary approaches. Average Nucleotide Identity, or ANI, was computed with two different tools, FastANI and skani, which estimate the overall similarity between two genomes by aligning shared DNA sequence. Digital DNA-DNA hybridization, dDDH, a computational descendant of the wet-laboratory hybridization experiments that once defined bacterial species, was calculated using formula 2. The researchers also employed tetranucleotide Z-score distance, TZMD, which measures differences in short-word DNA composition, and single-copy gene phylogenomics, reconstructing evolutionary trees from sets of genes present exactly once in each genome. Crucially, rather than comparing each strain against only its nearest named neighbors, the authors performed comprehensive pairwise comparisons among all available RefSeq genomes at both complete and chromosome assembly levels for each taxonomic neighborhood, an unusually thorough sweep designed to reveal the true structure of variation around each strain.
For two of the six strains, the answer came back clean. Both could be assigned without ambiguity, one to Bacillus subtilis and one to Bacillus licheniformis, with every ANI comparison against their respective species exceeding the conventional 95 percent threshold and with clear phylogenomic separation from neighboring taxa. These cases show that the classical framework still works when species boundaries are genuinely well separated. The complications began with the remaining four strains, each of which landed in a different taxonomic minefield.
Strain Bac2 fell within the Operational Group Bacillus amyloliquefaciens, and here the analysis documented what the authors call fundamental boundary inconsistency. ANI values measured between recognized species within this group exceeded the ANI values measured among genomes assigned to Bacillus amyloliquefaciens itself. In other words, genomes bearing different species names were more similar to each other than genomes carrying the same name, a direct inversion of what a coherent species concept requires. Single-copy gene phylogenomics confirmed the chaos, revealing extensive intermixing of named species across the tree. A name assigned under these conditions conveys little about evolutionary relatedness.
Strain zielonkawy was assigned to the Bacillus cereus sensu stricto genomospecies, but its placement highlighted a familiar and stubborn problem: the deep intermixing of Bacillus cereus sensu stricto and Bacillus thuringiensis. These two names describe bacteria with dramatically different ecological roles, one an opportunistic pathogen and the other an insecticidal biocontrol agent, yet their genomes remain so entangled that no genomic metric reliably separates them. The new data add one more well-documented instance to a debate that has persisted since whole genomes first became available.
The two remaining strains, assigned to Priestia megaterium and Paenibacillus amylolyticus, produced perhaps the most novel findings. Both ANI multi-comparison analysis and single-copy gene phylogenomics suggested possible species intermixing or mislabelling within the public reference databases themselves. To the authors’ knowledge, this is the first study to quantitatively document species delimitation problems between Priestia megaterium and Priestia aryabhattai, and between Paenibacillus amylolyticus and Paenibacillus xylanexedens. The Paenibacillus analysis carried the caveat of a small available sample size, but the Priestia result points to a quietly widespread issue: reference databases, which thousands of labs treat as ground truth, may contain genomes whose species labels do not survive close genomic scrutiny.
Beyond the individual assignments, the study delivers a quantitative contribution to the methodology of bacterial taxonomy itself. Concordance analysis between FastANI and dDDH formula 2 revealed a non-linear relationship, well described by a quadratic fit with an R-squared of 0.991. From this relationship, the researchers derived a striking number: the traditional 70 percent dDDH species threshold, inherited from the pre-genomic era of DNA reassociation experiments, corresponds not to the conventionally assumed 95 percent ANI but to approximately 96.16 percent ANI. In practical terms, dDDH formula 2 is the more conservative of the two metrics, meaning that genomes judged to be the same species by the 95 percent ANI rule could still fail the dDDH test. Laboratories relying on ANI alone may be lumping together organisms that a stricter standard would split.
The comparison between alignment-free and tree-based approaches added further nuance. Correlations between skani genomic distances and single-copy gene phylogenomic patristic distances, the branch-length distances separating genomes on the reconstructed trees, were high across all datasets, with R-squared values ranging from 0.900 to 0.997. Overall genomic structure, in other words, is consistent between methods. But Spearman rank correlations, while still strong, were consistently lower, at 0.780 to 0.916, indicating that the precise identity of a genome’s closest neighbors can shift depending on whether you measure raw sequence similarity or reconstructed evolutionary distance. For taxonomists deciding whether a strain belongs to one species or its nearest rival, that rank-order disagreement is exactly where decisions get made.
As a constructive response, the authors propose a complementary diagnostic tool: within- and between-species ANI multi-comparison analysis paired with PERMANOVA, a non-parametric statistical test for differences among groups, followed by post-hoc pairwise testing. The idea is to treat species boundaries not as a single threshold but as a statistical question: do the ANI distributions within a named species differ significantly from the distributions between it and its relatives? Applied alongside single-copy gene phylogenomics, this framework could flag boundary inconsistencies that single-threshold assignment silently passes over, giving taxonomists an explicit, reproducible way to test whether a species name still carves nature at its joints.
The broader significance of the work extends well beyond six soil isolates from Polish maize fields. The Caryophanales taxa examined here anchor industries from probiotics to pest control and include the closest relatives of the anthrax bacillus. If species boundaries in these groups are inconsistent, then everything built on top of those names, from safety assessments of biocontrol strains to regulatory definitions of pathogenic species, inherits the uncertainty. The study also lands amid an ongoing, sometimes contentious community effort to redefine prokaryotic species entirely, with competing proposals for genome-based circumscriptions. By showing that even gold-standard tools disagree at the margins, and that the canonical thresholds themselves encode hidden conservatism, the Lodz team’s results argue for pluralism: no single number can settle species questions in difficult groups, but a disciplined combination of metrics, statistics and phylogenetics can at least make the disagreements visible, quantifiable and, ultimately, resolvable.
Subject of Research: Genome-based species delimitation and taxonomic boundary evaluation in Caryophanales bacteria using ANI, dDDH and phylogenomics
Article Title: Challenges in Caryophanales species delimitation: comparative evaluation of genomic similarity metrics and phylogenomics based on six environmental strains
Article References: Grzyb, T., Wlaźlak, M., & Szulc, J. (2026). Challenges in Caryophanales species delimitation: comparative evaluation of genomic similarity metrics and phylogenomics based on six environmental strains. BMC Genomics. https://doi.org/10.1186/s12864-026-13320-7
Image Credits: AI Generated
DOI: 10.1186/s12864-026-13320-7
Keywords: Caryophanales, species delimitation, phylogenomics, Average Nucleotide Identity, digital DNA-DNA hybridization, Bacillus subtilis group, Bacillus cereus group, Priestia, Paenibacillus, bacterial taxonomy, BMC Genomics, environmental strains
Cite Scienmag News
Juliet Wilcox. (September 12, 2026). Genomic Metrics Expose Fuzzy Species Boundaries in Bacillus-Like Bacteria. Scienmag. https://scienmag.com/genomic-metrics-expose-fuzzy-species-boundaries-in-bacillus-like-bacteria/
Juliet Wilcox. "Genomic Metrics Expose Fuzzy Species Boundaries in Bacillus-Like Bacteria." Scienmag, 12 September 2026, https://scienmag.com/genomic-metrics-expose-fuzzy-species-boundaries-in-bacillus-like-bacteria/. Accessed 12 September 2026.
Juliet Wilcox. "Genomic Metrics Expose Fuzzy Species Boundaries in Bacillus-Like Bacteria." Scienmag. September 12, 2026. https://scienmag.com/genomic-metrics-expose-fuzzy-species-boundaries-in-bacillus-like-bacteria/

