Quantum machine learning has promised much and delivered, so far, mostly small demonstrations. Now a team of Spanish researchers has taken an unusually sober step: instead of claiming that quantum computers can outpredict classical models on crop yields, they built a meticulous benchmark to ask whether quantum machine learning is even reproducible across the software platforms scientists actually use. The answer is a qualified yes, and the details are more interesting than the headline.
The study, published in Smart Agricultural Technology by Fouad Ailabouni and colleagues at the University of Salamanca’s International Chair on Trustworthy Artificial Intelligence and Demographic Challenge, sits within Spain’s national response to rural depopulation, the so-called reto demográfico. That agenda calls for AI tools that can sustain agricultural productivity in dryland regions such as Salamanca, Zamora and Teruel, where warming and declining rainfall are steadily shifting the climate that yield models were trained on. A crop-yield predictor deployed for a five-year decision cycle may therefore face conditions materially different from its training data, and the operators who rely on it must be able to audit exactly how the software behaves.
The researchers chose a Quantum Support Vector Machine, or QSVM, as their test vehicle. This quantum primitive embeds classical data points into the Hilbert space of a multi-qubit register using a parametrised feature map, then computes a kernel matrix measuring the overlap between every pair of embedded points. That matrix feeds a standard classical support vector classifier. Because the kernel is a deterministic function of the software stack that simulates it, the QSVM is an ideal probe for cross-platform reproducibility: any numerical discrepancy between frameworks must trace back to the stack itself, not to training randomness. The authors are explicit that this is a benchmark of reproducibility and robustness, not a claim of quantum advantage.
The synthetic agronomic dataset was built from seven variables calibrated to Spanish dryland ranges: nitrogen, phosphorus, potassium, temperature, humidity, soil pH and rainfall, with a smooth latent yield score dominated by temperature and rainfall. The classifier saw only a two-dimensional projection of temperature and rainfall, standardised and rescaled using statistics frozen on the training split. A four-qubit ZZ-style feature map, drawn from the established quantum-kernel literature, was implemented identically in Qiskit, Cirq and PennyLane behind a common interface, and the resulting 60 by 60 kernel matrices were compared at machine precision across 40 independent data-generation seeds.
The cross-platform result is striking. Qiskit and PennyLane agreed on the kernel matrix to a mean relative divergence of 1.91 times ten to the minus sixteen, the floor of IEEE-754 double-precision arithmetic. Pairs involving Cirq showed a larger divergence of roughly 1.8 times ten to the minus seven, but a dedicated ablation traced this entirely to Cirq’s default simulator using single-precision complex64 numbers rather than any difference in gate implementation. When the researchers forced Cirq to double precision, the gap collapsed by eight orders of magnitude. Crucially, none of this mattered for the classifier: across all 160 seed-and-regime comparisons per platform pair, the three SDKs produced identical accuracy, F1 score and area under the ROC curve at every single cell. The smallest margin slack in the trained classifiers was far larger than the floating-point perturbation, so no support-vector decision was ever flipped.
The second axis of the benchmark simulated climate shift. Three perturbation regimes were applied to the test set only, with the classifier never retrained: a drought regime adding three degrees Celsius and halving rainfall, a wet regime dropping two degrees and multiplying rainfall by 1.6, and an extreme-heat regime adding five degrees. The magnitudes are author-designed stress tests calibrated to be broadly consistent with the direction of Iberian climate change projected under the CMIP6 ensemble, and all sit within or below documented Iberian extremes, including the 2003 heatwave and the 2004/05 drought. Under a strict covariate-shift protocol, the label threshold, scaler and classifier were all frozen on training data, so any accuracy change reflects genuine distributional drift rather than an artefact of label generation.
The robustness findings subvert simple narratives. Neither the QSVM nor most classical baselines showed a statistically significant accuracy change under drought or extreme heat across 40 paired seeds with Holm-corrected Wilcoxon tests. Under the wet regime, the QSVM and three non-linear classical baselines, the RBF and polynomial SVCs and a small multilayer perceptron, all gained accuracy by similar amounts, with balanced accuracy and AUROC rising alongside. A majority-class baseline, which contains no learned information at all, moved sharply in the opposite direction, dropping as test-set prevalence shifted by up to 19 percentage points. That opposing movement rules out prevalence drift alone as the explanation: the wet perturbation appears to change the geometry of the samples relative to the learned decision boundaries, an effect shared across model families rather than specific to the quantum method.
Against the seven classical baselines, the QSVM told a familiar story. It beat the linear models by a large, uniformly significant margin at every regime, matched the RBF, polynomial, random forest and gradient boosting models with no significant difference in any of 16 comparison cells, and was significantly outperformed by the multilayer perceptron at every regime. Even under full grid-search tuning of every model, the tuned QSVM was beaten by at least one classical model in 130 of 160 individual seed-and-regime cells. The most damning result for quantum-advantage hopes, however, is analytic: because every gate in the feature map after the initial Hadamard layer is diagonal in the computational basis, the team derived an exact, SDK-free closed form for the kernel as a 16-term finite sum. This classical evaluation reproduced the QSVM’s predictions perfectly at roughly 40 times lower wall-clock cost, a positive demonstration that the four-qubit circuit carries no computational quantum advantage at this scale.
The team then stress-tested their own conclusions with an unusually thorough battery of sensitivity analyses. A feature-map ablation showed the primary conclusions are insensitive to encoding choice. A shot-noise study found that finite-sampled kernels converge toward the statevector reference but always require positive-semidefinite correction, a structural property any hardware deployment should heed. A sample-size replication at 150 training points over 20 seeds confirmed every central claim. Three real-data anchors, the UCI Wine dataset, a sensor-measured crop recommendation dataset sharing the same seven-variable schema, and a FAOSTAT-derived maize yield task, reproduced the same cross-platform parity pattern, with a single isolated prediction disagreement on the measured-yield task. A generator-sensitivity check showed the cross-SDK parity is fully data-independent, while the wet-regime gain weakened under correlated feature sampling, an honest caveat the authors flag prominently.
The practical upshot is a methodology rather than a breakthrough. The complete benchmark, including instance generators, runners, raw logs and pinned environment files, is publicly archived and runs end-to-end on a single workstation in minutes, with no quantum hardware required. For practitioners, the authors recommend pinning SDK versions, verifying kernel divergence below ten to the minus five against a reference implementation, explicitly documenting numerical precision, calibrating climate-shift regimes to local projections, and reporting balanced accuracy or AUC alongside raw accuracy before drawing any robustness conclusion. The findings are bounded to noise-free statevector simulators, a fixed four-qubit map and a synthetic benchmark, and the authors are careful not to claim general climate robustness for any classifier family. But as quantum machine learning inches toward real-world decision support in agriculture and beyond, this study sets a template: audit the software stack, test under distribution shift, and let a majority-class baseline keep everyone honest.
Subject of Research: Cross-SDK numerical reproducibility and climate-shift robustness of quantum kernel classifiers on agronomic data
Article Title: A controlled cross-SDK numerical reproducibility benchmark for quantum kernels under synthetic agronomic distribution shifts
Article References: A controlled cross-SDK numerical reproducibility benchmark for quantum kernels under synthetic agronomic distribution shifts. (n.d.). Original publication
Image Credits: AI Generated
DOI: Not provided
Keywords: quantum machine learning, QSVM, reproducibility, Qiskit, Cirq, PennyLane, distribution shift, crop yield prediction, climate robustness, quantum kernels, agricultural AI, NISQ simulation
Cite Scienmag News
Katie Riggs. (October 10, 2026). Quantum Kernels Prove Identical Across Qiskit, Cirq and PennyLane in New Agronomic Benchmark. Scienmag. https://scienmag.com/quantum-kernels-prove-identical-across-qiskit-cirq-and-pennylane-in-new-agronomic-benchmark/
Katie Riggs. "Quantum Kernels Prove Identical Across Qiskit, Cirq and PennyLane in New Agronomic Benchmark." Scienmag, 10 October 2026, https://scienmag.com/quantum-kernels-prove-identical-across-qiskit-cirq-and-pennylane-in-new-agronomic-benchmark/. Accessed 10 October 2026.
Katie Riggs. "Quantum Kernels Prove Identical Across Qiskit, Cirq and PennyLane in New Agronomic Benchmark." Scienmag. October 10, 2026. https://scienmag.com/quantum-kernels-prove-identical-across-qiskit-cirq-and-pennylane-in-new-agronomic-benchmark/

