For decades, measuring how efficiently a crop photosynthesizes has meant clamping leaves into portable gas-exchange devices one at a time, a slow and painstaking process that limits how many plants a breeder can realistically evaluate. Now a team of researchers in China has built a deep learning framework that can estimate more than a dozen photosynthetic traits simultaneously from hyperspectral images, pixel by pixel, with accuracies that rival direct measurement. The work, published in the journal Artificial Intelligence in Agriculture, could reshape how plant scientists phenotype crops and how breeders select varieties capable of sustaining yields under a changing climate.
The research, led by Mengqi Zhang and colleagues including Xinguang Zhu and Minjuan Wang, was conducted at the Songjiang Experimental Station of the Center of Excellence in Molecular Plant Sciences at the Chinese Academy of Sciences. The team worked with two very different crops: dwarf ‘MicroTom’ tomato plants grown in a greenhouse under three light intensities, and 33 rice cultivars raised in outdoor plots under high and low nitrogen treatments. From these plants they collected 18 distinct photosynthetic phenotypic parameters, ranging from net photosynthetic rate and stomatal conductance to chlorophyll concentrations, leaf nitrogen content, water-use efficiency, and light-response traits such as maximum photosynthetic rate and the light compensation point.
The central obstacle the researchers set out to overcome is a familiar one in machine learning: data hunger. Traditional hyperspectral studies typically match an averaged reflectance spectrum from a region of leaf tissue with a single averaged physiological measurement, yielding datasets of only 100 to 200 samples. That is far too little to train a deep neural network reliably, and models built on such data often generalize poorly across cultivars, environments, and scales. Reported coefficients of determination in the existing literature typically range from 0.55 to 0.90, with considerable instability.
The team’s solution was a pixel-level data construction strategy. Instead of averaging reflectance across a leaf region, they treated every individual pixel of a hyperspectral image as an independent training sample, pairing each pixel’s 273-band reflectance vector with the leaf-level physiological measurements taken for that leaf. Because a single region of interest contains thousands of pixels, this approach expanded the dataset dramatically: 20,144 pixel samples for tomato and 12,648 for rice, compared with the 129 and 165 leaves that were physically measured. The authors describe this as a weakly supervised framework, since the ground truth exists only at the leaf level, and they were careful to partition the data at the leaf level rather than the pixel level to prevent information leakage between training and test sets.
At the heart of the framework sits the Multi-task Transformer Inversion Network, or MTI-Net. The architecture takes the 273-dimensional reflectance spectrum of each pixel, projects it into a 128-dimensional latent space, and passes it through a Transformer encoder that uses self-attention to capture long-range dependencies across the spectral range. A shared decoder then produces simultaneous predictions for all target parameters at once. Because different physiological traits operate on wildly different numerical scales, the model employs a dynamic loss-weighting scheme in which the weights themselves are learnable parameters, updated by a dedicated optimizer and constrained by a regularization penalty to prevent extreme values. The entire network contains only about 110,000 parameters, making it remarkably compact for a deep learning model.
The performance gains were striking. On the tomato dataset, MTI-Net achieved an average coefficient of determination of 0.9120 across 13 simultaneously predicted parameters, with a residual prediction deviation of 3.48, well above the threshold of 2 that is conventionally considered excellent in spectroscopic modeling. On rice, the model reached an R-squared of 0.9616 across 8 parameters. It outperformed every baseline tested, including partial least squares regression, support vector regression, random forest, XGBoost, and both single-task and multi-output deep neural networks. Crucially, when the pixel-trained model was applied to averaged spectra, it still maintained an R-squared of 0.84 on tomato and 0.98 on rice, demonstrating that the model learns genuine physiological relationships rather than merely exploiting spectral redundancy.
The study also delivered a practical answer to a question that has long hindered the field: how much data is enough? By systematically subsampling their datasets, the researchers found that accuracy climbed steadily with training-set size and only stabilized once the dataset exceeded roughly 5,000 pixel samples, at which point R-squared values exceeded 0.78. Below about 3,000 samples, models struggled to generalize from pixel-level training to averaged spectra. These findings give the community, for the first time, a quantitative benchmark for planning hyperspectral phenotyping campaigns.
A second innovation, the Differentiable Ranking and Selection Network, or DRS-Net, tackles the curse of dimensionality. Hyperspectral sensors capture hundreds of adjacent, highly collinear bands, many of which carry redundant or noisy information. DRS-Net learns to rank bands through a differentiable approximation of top-N selection, allowing the network to be trained end-to-end while converging on a minimal subset of informative wavelengths. The results showed that just 60 selected bands could nearly match the performance of the full 273-band spectrum, and 40 bands maintained an R-squared above 0.75. When the researchers mapped the selected bands onto a typical vegetation reflectance curve, they found the network had independently converged on biologically meaningful regions, clustering around the red-edge position between 650 and 750 nanometers and the green peak between 510 and 600 nanometers, areas known to be sensitive to chlorophyll content and leaf structure.
Beyond the methodology, the framework produced genuine biological insights. In tomato, the inferred parameters revealed that most photosynthetic traits peaked around 32 days after seedling emergence before declining, and that the photosynthetic dominance of the canopy shifted over time: middle-canopy leaves drove assimilation early in development, while upper-canopy leaves took over as they expanded and intercepted more light, particularly under high irradiance. In rice, the retrieved parameters confirmed strong responses of photosynthesis, leaf nitrogen, and chlorophyll to nitrogen fertilization, with the magnitude of those responses varying across the 33 genotypes, evidence that genetic background modulates physiological sensitivity to nutrient stress. The authors note that such model-derived patterns should be interpreted as high-probability trends rather than direct empirical measurements, given the weakly supervised labeling scheme.
The researchers are candid about the limitations. The experiments were conducted under controlled illumination in a darkroom and a greenhouse, so the framework has not yet faced the atmospheric interference, variable solar angles, and shadow effects of open-field conditions. Cross-species transfer also proved difficult: when the team froze the tomato-trained encoder and fine-tuned only the decoder on rice data, performance dropped sharply, a domain shift the authors attribute to differences in leaf morphology, trichome density, cuticle structure, and canopy architecture between the two species. Still, the study stands as a proof of concept that Transformer-based architectures combined with pixel-level data expansion can push hyperspectral phenotyping into a regime of accuracy and throughput that traditional statistical methods cannot reach, and that intelligent band selection could one day enable cheap, task-specific multispectral sensors for breeders and growers alike.
Subject of Research: Pixel-level multi-task deep learning for high-throughput estimation of crop photosynthetic phenotyping parameters from hyperspectral imagery
Article Title: Pixel-level high-throughput estimation of crop photosynthetic phenotyping parameters using a multi-task deep learning framework
Article References: Zhang, M., Zhou, J., Song, Q., Zhang, M., Zhu, X., & Wang, M. (2026). Pixel-level high-throughput estimation of crop photosynthetic phenotyping parameters using a multi-task deep learning framework. Artificial Intelligence in Agriculture. https://doi.org/10.1016/j.aiia.2026.08.014
Image Credits: AI Generated
DOI: 10.1016/j.aiia.2026.08.014
Keywords: hyperspectral imaging, deep learning, photosynthesis, plant phenotyping, multi-task learning, Transformer, precision agriculture, tomato, rice, band selection, remote sensing, crop breeding
Cite Scienmag News
Alan Morgan. (October 2, 2026). AI Reads Crop Photosynthesis Pixel by Pixel in Deep Learning Leap for Farming. Scienmag. https://scienmag.com/ai-reads-crop-photosynthesis-pixel-by-pixel-in-deep-learning-leap-for-farming/
Alan Morgan. "AI Reads Crop Photosynthesis Pixel by Pixel in Deep Learning Leap for Farming." Scienmag, 2 October 2026, https://scienmag.com/ai-reads-crop-photosynthesis-pixel-by-pixel-in-deep-learning-leap-for-farming/. Accessed 2 October 2026.
Alan Morgan. "AI Reads Crop Photosynthesis Pixel by Pixel in Deep Learning Leap for Farming." Scienmag. October 2, 2026. https://scienmag.com/ai-reads-crop-photosynthesis-pixel-by-pixel-in-deep-learning-leap-for-farming/

