Biomass may look like an ordinary pile of agricultural residue, forestry waste, or discarded plant material, but inside that material is a chemically complex mixture that determines how efficiently it can become fuel, heat, electricity, chemicals, or carbon-rich products. Now, researchers have shown how machine learning can extract that hidden chemical information from derivative thermogravimetric data, potentially making biomass analysis faster, cheaper, and more adaptable to the enormous variety of feedstocks used in modern bioenergy systems. The study, published in Scientific Reports, examines whether patterns in thermal-decomposition curves can be used to predict the chemical composition of different biomass materials without relying entirely on lengthy laboratory procedures.
The work by S. Pattanayak, D. Saha, C. Loha and colleagues addresses a central challenge in biomass conversion. No two feedstocks are chemically identical. Rice husks, wood chips, straw, grasses, shells, leaves, and other residues contain different proportions of cellulose, hemicellulose, lignin, extractives, moisture, and mineral matter. Those differences strongly influence how a material behaves during combustion, pyrolysis, gasification, and other thermochemical processes. A feedstock rich in cellulose may decompose differently from one dominated by lignin, while high ash content can interfere with conversion equipment and alter the quality of the resulting products. Knowing the composition before processing is therefore essential, but conventional chemical analysis can be time-consuming, labor-intensive, and expensive.
The researchers focus on derivative thermogravimetry, commonly abbreviated as DTG. In a thermogravimetric experiment, a small biomass sample is heated under a controlled atmosphere while its mass is continuously measured. As the temperature rises, water evaporates, volatile compounds are released, and the major structural polymers in the plant material begin to break down. The instrument records the change in mass as a function of temperature. Derivative thermogravimetry takes this one step further by calculating the rate of mass loss, producing a curve that highlights the temperatures at which decomposition is fastest. Instead of treating the curve as a simple visual trace, the study uses it as a thermal fingerprint of the sample’s chemical structure.
That fingerprint is closely linked to the behavior of the three principal lignocellulosic polymers. Hemicellulose generally begins decomposing at lower temperatures than cellulose and produces a comparatively broad thermal response. Cellulose often generates a sharper and more intense decomposition peak, reflecting the breakdown of its ordered carbohydrate chains. Lignin is more structurally irregular and decomposes across a wider temperature range, with a substantial fraction of its mass persisting to higher temperatures. In real biomass, these processes overlap, and the resulting DTG curve is not a direct chemical inventory. It is a blended signal shaped by particle size, heating rate, atmosphere, moisture, ash, and interactions among the components. Machine learning is valuable precisely because it can identify relationships in that overlapping signal that are difficult to isolate by inspection alone.
The study’s central idea is to connect measurable features of DTG profiles with laboratory estimates of biomass composition. A machine-learning model can be trained using samples for which both thermal curves and chemical composition are known. During training, the algorithm searches for mathematical relationships between variables such as peak temperature, peak height, peak width, mass-loss regions, and the overall shape of the derivative curve. Once trained, the model can process the DTG data from a new feedstock and generate predictions for its compositional characteristics. This approach does not eliminate the need for reliable reference measurements during model development, but it could reduce the amount of repeated wet-chemical analysis required when large numbers of samples must be screened.
The potential impact extends far beyond the laboratory. Industrial biomass plants routinely receive materials whose properties vary with season, geography, cultivation practices, storage conditions, and harvesting methods. A rapid compositional estimate could help operators blend feedstocks more intelligently, adjust reactor temperature, control residence time, and anticipate changes in product yield. In pyrolysis, for example, the balance of cellulose, hemicellulose, and lignin affects the distribution of bio-oil, gases, and solid char. In combustion, moisture and ash influence ignition, flame stability, emissions, and slagging. In gasification, composition affects syngas production and the formation of unwanted tars. A prediction system based on DTG data could provide a practical quality-control tool before material enters a conversion process.
The research also illustrates why artificial intelligence is becoming increasingly important in thermal analysis. Traditional interpretation often depends on simplified assumptions, such as assigning individual temperature ranges to particular biomass components. Those assumptions are useful, but real samples produce overlapping reactions and are influenced by experimental conditions. Machine-learning algorithms can handle multiple variables simultaneously and detect nonlinear patterns that conventional regression methods may overlook. However, the usefulness of any prediction depends on the quality and diversity of the training data. If a model is trained mainly on a narrow group of feedstocks, it may perform well on familiar samples but struggle with an unfamiliar crop residue, contaminated waste stream, or material with unusually high mineral content.
That limitation makes validation especially important. A model must be tested on data that were not used during training, and its predictions should be compared with established analytical methods. Researchers and industrial users also need to know not only the predicted value but how uncertain that prediction may be. Differences in heating rate, atmosphere, sample preparation, instrument calibration, and data processing can shift DTG peaks and change the apparent relationship between thermal behavior and composition. For machine learning to become a dependable analytical technology, models will need transparent reporting, carefully standardized experiments, and broad datasets representing the real diversity of biomass encountered outside the laboratory.
The study points toward a future in which thermal instruments and computational models operate as a rapid screening platform for renewable carbon resources. A small sample could be heated, its DTG profile captured, and its likely chemical composition estimated within a workflow considerably faster than a full suite of conventional analyses. Such systems could support decentralized biomass processing, where facilities need to make quick decisions about locally available materials. They could also contribute to better resource management by helping identify feedstocks suited to particular products rather than treating all biomass as interchangeable. The broader message is not that machine learning replaces chemistry, but that it can transform a complex thermal signal into actionable information when it is built on sound measurements and carefully validated data.
As countries search for lower-carbon energy and materials, technologies that improve the efficiency of biomass conversion are attracting growing attention. Yet the sustainability of bioenergy depends on more than simply using plant-derived material; it also depends on matching the right feedstock to the right process and avoiding waste, emissions, and inefficient operation. By linking derivative thermogravimetric analysis with machine learning, Pattanayak, Saha, Loha and their colleagues offer a route toward faster chemical characterization of diverse biomass resources. The approach could help turn an unpredictable raw material into a more precisely managed industrial input, bringing data-driven control to one of the most variable resources in the renewable-energy landscape.
Subject of Research: Machine-learning prediction of biomass chemical composition using derivative thermogravimetric data.
Article Title: Machine-learning prediction of biomass chemical composition using derivative thermogravimetric data of different biomass feedstocks.
Article References: Pattanayak, S., Saha, D., Loha, C. et al. “Machine-learning prediction of biomass chemical composition using derivative thermogravimetric data of different biomass feedstocks.” Scientific Reports 16, 25132 (2026). https://doi.org/10.1038/s41598-026-65442-3
Image Credits: AI Generated
DOI: https://doi.org/10.1038/s41598-026-65442-3
Keywords: biomass, machine learning, derivative thermogravimetry, thermogravimetric analysis, cellulose, hemicellulose, lignin, bioenergy, pyrolysis, biomass characterization

