Granites are the quiet archivists of our planet. Locked inside their crystals and glassy intergrowths is a chemical record of the forces that assembled continents, opened oceans, and drove mountain belts skyward. For decades, geologists have tried to decode that record using hand-drawn discrimination diagrams, plotting a handful of elements against one another to infer whether a granite formed above a subduction zone, within a stable continental interior, or along a rift tearing a landmass apart. Now a research team has handed that task to machine learning, and the results suggest that the chemical fingerprints of tectonic settings are far richer, and far more multivariate, than any two-dimensional diagram can capture.
In a study published in Earth Science Informatics, Huanbao Zhang of the Hunan Institute of Technology and colleagues, working with collaborators at the University of South China, compiled an extraordinary dataset: 13,397 granite samples drawn from the global GEOROC database, each characterized by its major and trace element composition. Rather than relying on a few curated elements, the team let the full chemical spectrum speak. Every sample was assigned to one of five tectonic-setting classes: intraplate volcanics, convergent margins, Archean cratons, continental flood basalt provinces, and rift volcanics. This taxonomy spans the dramatic range of environments in which granitic magmas are generated, from the crushing pressures above sinking ocean slabs to the hot, stable roots of the oldest continental nuclei on Earth.
The scale of the compilation matters. Traditional discrimination diagrams, such as the influential trace element schemes developed by Julian Pearce and colleagues in the 1980s, were built on comparatively small reference suites and depend on subjective empirical criteria for drawing field boundaries. A single sample plotted near a dividing line can be interpreted one way or another depending on the eye of the beholder. By contrast, a dataset of more than thirteen thousand samples allows an algorithm to learn the statistical texture of each tectonic environment, including the overlapping zones where settings genuinely blur into one another. It is the difference between memorizing a few textbook examples and having read the entire library.
Before any learning could happen, the team invested heavily in data preprocessing and feature engineering, the craft of transforming raw measurements into inputs that algorithms can use effectively. Geochemical data are notoriously messy: detection limits vary between laboratories, elements are reported over many orders of magnitude, and some measurements correlate so strongly with others that they carry redundant information. Feature engineering addresses these problems by creating derived variables, such as element ratios, that often encode petrological meaning more directly than raw concentrations. The care taken at this stage is frequently what separates a machine learning result that works from one that merely memorizes noise.
The heart of the method is a stacking ensemble, an architecture in which multiple different learners are trained on the same problem and a higher-level model learns how best to combine their votes. The team integrated three complementary algorithms: XGBoost and LightGBM, two highly efficient implementations of gradient-boosted decision trees, and a multilayer perceptron, a classic neural network capable of capturing smooth nonlinear relationships. Gradient boosting excels at tabular data like geochemical compositions, building sequences of decision trees that each correct the errors of the last, while the neural network contributes a different inductive bias. Stacked generalization, a technique pioneered by David Wolpert in 1992, lets a meta-learner discover which base model to trust under which circumstances, often squeezing out performance that no single model could achieve alone.
The numbers are striking. After optimization, the ensemble achieved an accuracy of 0.91, a balanced accuracy of 0.90, and a macro-F1 score of 0.89 across the five tectonic classes. Balanced accuracy and macro-F1 are particularly important here because they prevent a model from looking good simply by favoring the most common classes; they require the algorithm to perform well even on rare tectonic settings. In practical terms, roughly nine out of ten granites can be assigned to their correct tectonic environment from chemistry alone, a level of reliability that approaches, and in some respects exceeds, what expert-driven diagram-based classification typically delivers, especially for ambiguous samples.
But raw accuracy is only half the story, and arguably the less interesting half. The real advance lies in the team’s use of SHAP analysis, a technique from explainable artificial intelligence developed by Scott Lundberg and Su-In Lee that attributes each prediction to the contributions of individual input features. SHAP values, grounded in cooperative game theory, reveal not just which elements matter most overall but how each element pushes a particular sample toward or away from a particular classification. For a discipline that has long depended on interpretable, if crude, graphical tools, this transparency is what makes machine learning a genuine scientific partner rather than a black-box oracle.
The SHAP results are geochemically telling. The most predictive features were associated with holmium, yttrium, the heavy rare earth elements, thorium, neodymium, and dysprosium. That lineup is not arbitrary. Yttrium and the heavy rare earths are sensitive to the presence or absence of garnet and other minerals in the source region or in the residue left behind during melting, and their behavior therefore tracks the depth and temperature conditions of magma generation. Thorium enrichments speak to crustal contributions and the degree of differentiation, while neodymium participates in the well-known geochemical systems used to trace mantle versus crustal sources. In other words, the algorithm independently rediscovered, through pure pattern recognition, element groups that petrologists consider diagnostic, and it did so across all five tectonic settings with distinct contribution patterns for each.
That last finding carries a conceptual punch. The analysis emphasized the multivariate nature of granite tectonic discrimination: no single element or pair of elements cleanly separates the settings, and the discriminative information is distributed across a web of interacting concentrations. This helps explain why conventional diagrams, which by necessity flatten high-dimensional chemistry into two or three axes, so often produce ambiguous results. The machine learning framework does not replace those diagrams so much as reveal their limits, offering an interpretable complement that can quantify how much each chemical variable contributes to a decision and expose where the boundaries between tectonic environments genuinely overlap in nature.
The implications reach well beyond academic taxonomy. Reconstructing ancient tectonic settings is foundational to understanding crustal evolution and geodynamic processes, and it feeds directly into mineral exploration, since many ore deposits are intimately tied to specific magmatic environments. A tool that can reliably fingerprint a granite’s birthplace from a routine whole-rock analysis could sharpen exploration targeting, refine reconstructions of vanished supercontinents, and help test competing models of how the earliest continental crust, including the Archean cratons included in this study, was assembled. The work also joins a growing wave of machine learning applications across the geosciences, from basalt discrimination to tracking crustal thickness variations, that suggest the era of big data in Earth science is moving from promise to practice. For a rock that has been silent for billions of years, granite is suddenly talking, and algorithms are learning to listen.
Subject of Research: Machine learning classification of granite tectonic settings using global geochemical data and SHAP interpretation
Article Title: Machine learning-based discrimination of granite tectonic settings: Integrating global geochemical datasets, feature engineering and SHAP interpretation
Article References: Zhang, H., He, H., Shi, Y., He, Y., Zeng, T., Tao, Y., Wang, Z., Xie, Y., Huang, X., & Li, X. (2026). Machine learning-based discrimination of granite tectonic settings: Integrating global geochemical datasets, feature engineering and SHAP interpretation. Earth Science Informatics, 19(11), Article 199. https://doi.org/10.1007/s12145-026-02256-x
Image Credits: AI Generated
DOI: 10.1007/s12145-026-02256-x
Keywords: granite, tectonic settings, machine learning, geochemistry, SHAP, XGBoost, LightGBM, stacking ensemble, rare earth elements, crustal evolution, GEOROC, petrology
Cite Scienmag News
Violet Maxwell. (October 3, 2026). AI Reads Ancient Granite Chemistry to Reconstruct Earth’s Tectonic Past. Scienmag. https://scienmag.com/ai-reads-ancient-granite-chemistry-to-reconstruct-earths-tectonic-past/
Violet Maxwell. "AI Reads Ancient Granite Chemistry to Reconstruct Earth’s Tectonic Past." Scienmag, 3 October 2026, https://scienmag.com/ai-reads-ancient-granite-chemistry-to-reconstruct-earths-tectonic-past/. Accessed 3 October 2026.
Violet Maxwell. "AI Reads Ancient Granite Chemistry to Reconstruct Earth’s Tectonic Past." Scienmag. October 3, 2026. https://scienmag.com/ai-reads-ancient-granite-chemistry-to-reconstruct-earths-tectonic-past/

