Deep beneath the farmland of São Paulo State, in the quarry walls of the Monte Olimpo mine, lies one of South America’s most intriguing petroleum mysteries: the Permian Irati Formation, a 280-million-year-old succession of bituminous shale and dolomitic limestone that has long been recognized as one of the Paraná Sedimentary Basin’s principal hydrocarbon source rocks. Now, a team of Brazilian and Portuguese researchers has shown that the key to unlocking this ancient rock’s secrets lies not only in chemistry, but in code. By combining two classic multivariate statistical techniques with a rigorously validated Python workflow, they have produced a reproducible framework for turning dense geochemical datasets into clear geological stories.
The study, published in Discover Geoscience, was led by Gabrielle Roveratti and Daniel Marcos Bonotto of São Paulo State University (UNESP), together with Carla Alexandra de Figueiredo Patinha of the University of Aveiro in Portugal. Their target was a problem that plagues petroleum geoscientists everywhere: carbonate-rich shale sequences like the Irati are extraordinarily heterogeneous, and interpreting the floods of elemental data they generate is often treated as a black box, with software defaults quietly shaping the conclusions. The team set out to make every step of the analysis explicit, testable, and repeatable across independent computing environments.
The stakes are considerable. The Paraná Sedimentary Basin sprawls across roughly 1.5 million square kilometers of Brazil, Paraguay, Uruguay, and Argentina, accumulating sediments and volcanic rocks from the late Ordovician to the late Cretaceous that reach thicknesses of up to seven kilometers along the Paraná River axis. Within this vast basin, the Permian Irati Formation stands out. Its bituminous shales of the Assistência Member contain total organic carbon values ranging from 0.1 to 23 percent, averaging around 2 percent, and host Type I kerogen of algal origin, a variety with significant potential for generating liquid hydrocarbons. Interest surged in the late 2000s as the global energy sector turned to oil shales and unconventional reservoirs, and the Irati has since attracted attention as a potential analogue for fractured limestone reservoirs and low-permeability shale plays.
The rocks themselves carry a remarkable history. The Irati was deposited under restricted marine conditions, with limited water circulation between the basin’s interior and the Panthalassa Ocean, producing hypersaline environments in which carbonates and evaporites accumulated in the north while organic-rich bituminous shales blanketed the south. The formation is regionally correlated with the Mangrullo Formation in Uruguay and the Whitehill Formation in South Africa, all sharing abundant Mesosaurid fossils that once helped cement Wegener’s continental drift theory. U–Pb zircon dating of Mangrullo-equivalent strata places these units in the Cisuralian, between about 279 and 276 million years ago. Adding to the complexity, igneous dikes and sills have locally heated the succession, while diagenesis, oxidation–reduction reactions, weathering, and fracture-guided fluid flow have all reshaped the elemental signatures preserved in the rock.
To sample this complexity, the researchers collected 162 specimens, labeled with the prefix Pam19, from five quarry fronts operated by the Amaral Machado Mining Company at Saltinho, including the Monte Olimpo site. The sections consist of interbedded limestones and shales with a basal carbonate layer, plus a transitional impure limestone level known locally as Lajinha and peculiar massive, often spherical carbonate concretions called Cucuruto. Sampling intervals were deliberately uneven because lithological variability and heavily weathered shales in carbonate-rich intervals sometimes made collection impossible. Each sample was pulverized to below roughly 74 micrometers, mixed with Oregon wax binder, and pressed into pellets for X-ray fluorescence analysis of thirteen major oxides and trace components, from SiO₂ and Al₂O₃ to SrO and Cl, on a Bruker S8 TIGER spectrometer. Gamma-ray spectrometry with a sodium iodide detector then quantified the radioelements potassium-40, uranium, and thorium in every sample.
The computational heart of the study was a three-stage workflow: data preprocessing, multivariate statistical analysis, and geological integration. Before analysis, samples with multiple values below detection limits or extreme outliers were excluded, including one specimen with anomalously high P₂O₅ that would have distorted the multivariate structure. The team then ran Principal Component Analysis and Hierarchical Cluster Analysis twice over, once in IBM SPSS Statistics and once in Python 3.9 using pandas, scikit-learn, SciPy, matplotlib, and seaborn. In Python they even tested two preprocessing strategies, median imputation and casewise deletion, to check whether the choice of how to handle missing data would change the outcome. The full Jupyter Notebook code and a synthetic dataset mirroring the original data’s statistical properties were released on GitHub, making the entire pipeline open to scrutiny.
The statistical diagnostics were emphatic. The Kaiser-Meyer-Olkin measure of sampling adequacy reached 0.815, rated very good, and Bartlett’s test of sphericity was significant at p < 0.001, confirming the data were suitable for factor analysis. Principal Component Analysis extracted four components with eigenvalues above one, together explaining 72.1 percent of the total variance. The first and dominant component, accounting for 43.9 percent, revealed a striking binary mixing axis: strong positive loadings from terrigenous elements like Al₂O₃, K₂O, TiO₂, and SiO₂ opposed strong negative loadings from carbonate-associated CaO, MgO, Cl, and SrO. This is the geochemical fingerprint of the rhythmic alternation between siliciclastic shales and dolomitic limestones that defines the Assistência Member. The second component, explaining 13.2 percent, isolated the radioactive trio of uranium, potassium, and thorium, while the third and fourth components, at 7.6 and 7.4 percent, tracked iron and manganese oxides, hinting at secondary diagenetic or provenance controls.
Hierarchical Cluster Analysis, run with the Between-Groups Linkage method and Pearson’s correlation coefficient, independently reproduced the same structure, sorting the variables into a terrigenous suite, a radioactive suite, and a carbonate suite. Applied to the samples themselves, the clustering carved the stratigraphic column into coherent chemofacies that matched, but did not perfectly follow, the visible lithology. Two calcareous zones emerged with particular clarity. Calcareous Zone A, between 7.4 and 11.6 meters depth, showed strongly negative PC1 scores and belonged firmly to the carbonate cluster, consistent with dolomitic limestones of low detrital input. Calcareous Zone B, a narrow interval between 18.5 and 19.2 meters, shared the carbonate signature but stood apart with significantly higher PC2 scores and elevated strontium oxide content, a combination the authors attribute to possible differences in organic matter preservation, clay mineral contribution, or diagenetic overprinting. Intriguingly, some shale samples clustered with the carbonates, suggesting early diagenetic carbonate cementation that overprinted original depositional chemistry, though the team cautions that petrographic and isotopic work would be needed to confirm this.
The cross-platform validation delivered the study’s methodological punchline. The Python replication, using casewise deletion to mirror the SPSS defaults, produced a KMO of 0.810, the same four-component structure explaining over 70 percent of variance, and identical geochemical associations in the loading plots. Minor differences, such as Python splitting the SPSS first component into two correlated components, did not alter the geological interpretation. A sensitivity analysis with median imputation likewise showed that although imputation introduced small shifts in the variance structure, the principal associations and clustering patterns remained stable. The findings, in other words, are not artifacts of one company’s algorithms; they reflect genuine geochemical structure in the rock.
Beyond the Irati, the study offers a transferable template for anyone wrestling with heterogeneous carbonate–siliciclastic systems, from unconventional reservoir characterization to paleoenvironmental reconstruction. By insisting on transparent preprocessing, dual-software validation, and open code, Roveratti and colleagues demonstrate that reproducibility is not bureaucratic overhead but a scientific instrument in its own right, one sharp enough to detect subtle diagenetic signals that cross lithological boundaries. As exploration for oil shales and fractured carbonate reservoirs accelerates worldwide, workflows like this one promise to turn sprawling geochemical spreadsheets into actionable geological insight, one validated principal component at a time.
Subject of Research: Reproducible multivariate statistical and Python-based geochemical characterization of the Permian Irati Formation in the Paraná Sedimentary Basin, Brazil
Article Title: Using multivariate statistical analysis and Python programming for geochemical characterization at the Paraná Sedimentary Basin, Brazil
Article References: Using multivariate statistical analysis and Python programming for geochemical characterization at the Paraná Sedimentary Basin, Brazil. (n.d.). https://doi.org/10.1007/s44288-026-00686-0
Image Credits: AI Generated
DOI: 10.1007/s44288-026-00686-0
Keywords: geochemistry, Irati Formation, Paraná Basin, principal component analysis, hierarchical cluster analysis, Python, multivariate statistics, shale reservoir, chemofacies, X-ray fluorescence, gamma-ray spectrometry, reproducibility
Cite Scienmag News
Violet Maxwell. (October 8, 2026). Python and Statistics Decode the Geochemical Secrets of Brazil’s Oil-Rich Irati Shale. Scienmag. https://scienmag.com/python-and-statistics-decode-the-geochemical-secrets-of-brazils-oil-rich-irati-shale/
Violet Maxwell. "Python and Statistics Decode the Geochemical Secrets of Brazil’s Oil-Rich Irati Shale." Scienmag, 8 October 2026, https://scienmag.com/python-and-statistics-decode-the-geochemical-secrets-of-brazils-oil-rich-irati-shale/. Accessed 8 October 2026.
Violet Maxwell. "Python and Statistics Decode the Geochemical Secrets of Brazil’s Oil-Rich Irati Shale." Scienmag. October 8, 2026. https://scienmag.com/python-and-statistics-decode-the-geochemical-secrets-of-brazils-oil-rich-irati-shale/

