Cement is everywhere. It is the most widely produced material on Earth, the silent backbone of bridges, towers, dams, and cities, and it has been in continuous use for more than two millennia. Yet the very substance that gives concrete its strength remains one of the most stubbornly opaque materials in science. Calcium silicate hydrate, known to researchers as C–S–H, is the nanocrystalline glue that binds cement paste together, and it is responsible for nearly all of the cohesion, strength, and durability of concrete. It is also a major reason why cement production accounts for roughly eight percent of global carbon dioxide emissions, because understanding it well enough to design greener alternatives has proven extraordinarily difficult. Now, a team reporting in Advanced Science has unveiled a machine learning interatomic potential, dubbed CSH-MLP, that promises to bring quantum-mechanical accuracy to simulations of this enigmatic gel at scales and speeds never before possible.
The difficulty with C–S–H begins with its structure. Unlike the tidy crystals that fill mineralogy textbooks, C–S–H is X-ray amorphous, meaning its atomic arrangement is too disordered to resolve with conventional diffraction. Its calcium-to-silicon ratio varies widely, its silicate chains are broken at unpredictable points, and its water content shifts with the surrounding environment. This heterogeneity has frustrated experimentalists for decades and has equally hampered computationalists. Density functional theory, the gold standard of atomistic simulation, delivers benchmark accuracy but only for systems of a few hundred atoms over picoseconds. Classical molecular dynamics can reach larger scales, but its reliability rests entirely on the quality of empirical force fields, and the force fields developed for cement each carry serious compromises.
ClayFF, for example, covers a broad range of elements cheaply but was parameterized for clay minerals and struggles to reproduce the mechanics of C–S–H. The CSH-FF improves on crystalline tobermorite but cannot capture reactive chemistry. ReaxFF introduces bond breaking, useful for hydration reactions, yet suffers from poor mechanical accuracy and punishing computational demands. The newer ERICA FF extends elemental coverage and agrees better with quantum calculations, but it remains anchored to fixed functional forms. Machine learning potentials offer a way out of this dilemma: trained on quantum-mechanical data, they reproduce near-DFT accuracy while running orders of magnitude faster than ab initio molecular dynamics. Previous efforts had built such potentials for tobermorite or for narrow slices of C–S–H compositional space, but none had confronted the full chemical and structural wildness of the real material.
The heart of the new work is not the neural network itself but the dataset behind it. Because no public database of experimentally derived C–S–H atomic structures exists, the researchers constructed one from scratch using the pyCSH structure-generation program. They began by generating roughly 500,000 candidate supercells spanning calcium-to-silicon ratios from 0.83 to 2.5, then screened them on four key structural parameters, including silanol density, calcium hydroxylation, and mean silicate chain length, to eliminate redundancy while preserving diversity. The result was 1,926 distinct models covering 688 separate compositions. These were equilibrated with the ERICA force field, refined at the density functional theory level, and then augmented in two stages: deliberate structural perturbations of up to twenty percent generated 9,450 high-energy configurations, and configurations harvested from one hundred molecular dynamics trajectories added another 39,057. In total, more than 50,000 DFT-labeled configurations now constitute what the authors describe as the largest and most comprehensive openly accessible C–S–H dataset to date, a first step toward what they call a cement genome.
The training data were validated with structural fingerprinting. Using the smooth overlap of atomic positions method and principal component analysis, the team mapped the configurational landscape of their dataset and confirmed that it spreads across a genuinely broad space rather than clustering around a few crystal-like basins. The optimization sub-dataset sits near the origin, providing stable energy minima, while the perturbed and trajectory-sampled configurations fan outward, covering the non-equilibrium territory a simulation must navigate. This deliberate coverage of metastable and high-energy states is precisely what separates a robust potential from one that fails the moment atoms stray from perfect crystals, and it explains much of what follows.
The accuracy numbers are striking. CSH-MLP reproduces DFT atomic energies with root mean squared errors of 1.18 and 1.50 meV per atom on training and test sets respectively, and forces with errors of roughly 0.1 eV per angstrom, placing it toward the low end of the accepted range for machine learning potentials. During production simulations, the model maintained consistent accuracy across compositions spanning Ca/Si ratios from 0.91 to 2.06, and it stably equilibrated large-scale models it had never seen, including an 8,456-atom bulk structure and a 10,109-atom gel containing a water-filled nanopore, both built with methods entirely different from the training pipeline. When benchmarked on lattice constants, CSH-MLP deviated from DFT by an average of only 0.14 percent for the C–S–H benchmark model, and it even predicted tobermorite structures, which were absent from training, with errors mostly below one percent.
The comparisons with rival methods were revealing. The general-purpose foundation model MACE-MP-0, despite covering more than eighty elements, performed worse than CSH-MLP on both C–S–H and tobermorite, running at only 0.8 nanoseconds per day on a single GPU versus 21.8 for CSH-MLP. The earlier csh_cp potential, trained largely on crystalline configurations, produced spurious oxygen clustering and outright simulation failure on the gel system, while ClayFF showed unphysical water diffusion and ReaxFF showed artifacts of its own. CSH-MLP, by contrast, reproduced experimental pair distribution functions measured by two independent research groups, matched ab initio molecular dynamics on individual atomic pair correlations, and predicted a water diffusion coefficient in confined nanopores squarely within the range reported by prior studies.
Perhaps the most scientifically satisfying result concerns a long-standing contradiction. High-pressure X-ray diffraction experiments suggest that C–S–H stiffens monotonically as the calcium-to-silicon ratio rises, while decades of simulation work predict that higher calcium content breaks silicate chains and softens the material. CSH-MLP resolved the tension with a non-monotonic prediction: the bulk modulus climbs from 45.57 gigapascals at Ca/Si of 1.3 to a peak of 53.47 gigapascals at 1.5, then falls to 48.29 gigapascals at 1.9. Below the peak, added calcium densifies the interlayer and strengthens cohesion faster than chain scission weakens it; above it, progressive chain fragmentation and interlayer disorder take over. The low-calcium value agrees almost exactly with the experimental figure of 45.6 gigapascals, and the two regimes reconcile the densification story told by X-ray experiments with the defect story told by atomistic simulation as complementary halves of one physical picture.
Speed is the final piece of the puzzle. At the conservative 0.2 femtosecond timestep required by ReaxFF, CSH-MLP runs 3.5 to 6 times faster, and because it remains stable at timesteps up to 1.0 femtosecond, its effective advantage reaches 16.8 to 29.9 times. That means simulations of more than 10,000 atoms over nanosecond timescales at near-quantum accuracy, something previously infeasible for cementitious systems. The dataset and model are openly available on Figshare and the AIS Square platform, and the authors envision fine-tuning the framework toward carbonation, sulfate attack, alkali uptake, and interfaces. If the cement genome project succeeds, the humble glue holding the built environment together may finally yield its secrets, and with them, a credible path toward concrete that builds the future without overheating the planet.
Subject of Research: A machine learning interatomic potential for calcium silicate hydrate, the binding phase of cement, trained on a diverse DFT-labeled dataset
Article Title: Machine Learning Potential for Calcium Silicate Hydrates With Broad Compositional and Structural Diversity
Article References: Li, Y., Chen, C., Pellenq, R. J.-M., Li, Z., & Liu, J. (2026). Machine Learning Potential for Calcium Silicate Hydrates With Broad Compositional and Structural Diversity. Advanced Science, Article e77942. https://doi.org/10.1002/advs.77942
Image Credits: AI Generated
DOI: 10.1002/advs.77942
Keywords: calcium silicate hydrate, machine learning potential, cement, molecular dynamics, density functional theory, concrete, nanocrystalline materials, force fields, mechanical properties, carbon emissions, tobermorite, materials genome
Cite Scienmag News
Denise Maddox. (September 30, 2026). AI Cracks the Atomic Secrets of Cement’s Most Elusive Ingredient. Scienmag. https://scienmag.com/ai-cracks-the-atomic-secrets-of-cements-most-elusive-ingredient/
Denise Maddox. "AI Cracks the Atomic Secrets of Cement’s Most Elusive Ingredient." Scienmag, 30 September 2026, https://scienmag.com/ai-cracks-the-atomic-secrets-of-cements-most-elusive-ingredient/. Accessed 30 September 2026.
Denise Maddox. "AI Cracks the Atomic Secrets of Cement’s Most Elusive Ingredient." Scienmag. September 30, 2026. https://scienmag.com/ai-cracks-the-atomic-secrets-of-cements-most-elusive-ingredient/

