<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>benchmark &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/benchmark/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 09 Oct 2026 06:13:03 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>benchmark &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>AI Models Flunk Engineering Simulation Test in Massive New Benchmark</title>
		<link>https://scienmag.com/ai-models-flunk-engineering-simulation-test-in-massive-new-benchmark/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Fri, 09 Oct 2026 06:13:03 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI evaluation]]></category>
		<category><![CDATA[AI in engineering analysis]]></category>
		<category><![CDATA[AI limitations in technical visualization]]></category>
		<category><![CDATA[AI model evaluation]]></category>
		<category><![CDATA[benchmark]]></category>
		<category><![CDATA[Carnegie Mellon University]]></category>
		<category><![CDATA[Communications Engineering]]></category>
		<category><![CDATA[engineering simulation interpretation]]></category>
		<category><![CDATA[engineering simulations]]></category>
		<category><![CDATA[evaluation of AI reasoning in simulations]]></category>
		<category><![CDATA[finite element analysis]]></category>
		<category><![CDATA[impact of dataset scale on AI assessment]]></category>
		<category><![CDATA[large-scale simulation benchmark]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[mechanical engineering]]></category>
		<category><![CDATA[mechanical engineering AI testing]]></category>
		<category><![CDATA[OpenSeeSimE]]></category>
		<category><![CDATA[OpenSeeSimE dataset]]></category>
		<category><![CDATA[question answering]]></category>
		<category><![CDATA[simulation output comprehension]]></category>
		<category><![CDATA[statistical robustness in AI benchmarking]]></category>
		<category><![CDATA[vision-language model performance]]></category>
		<category><![CDATA[vision-language models]]></category>
		<category><![CDATA[visual reasoning]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=252249</guid>

					<description><![CDATA[A new 200,000-question benchmark from Carnegie Mellon University reveals that leading vision-language models perform at random chance when interpreting engineering simulation outputs, despite excelling at general visual reasoning tasks.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence systems that can ace medical imaging quizzes and describe photographs in astonishing detail may be far less capable than they appear when confronted with the colorful contour plots and meshed geometries of engineering simulations. That is the sobering conclusion of a new study published in Communications Engineering, in which researchers at Carnegie Mellon University built one of the largest benchmarks ever assembled for evaluating how well vision-language models can interpret engineering simulation outputs. The result: ten state-of-the-art models, many of them celebrated for their general visual reasoning skills, performed at or near random chance when asked questions about simulation results, with effect sizes so small the researchers describe them as negligible.</p>
<p>The benchmark, called OpenSeeSimE, consists of more than 200,000 question-answer pairs spanning 10,000 parametrically varied simulations. According to the authors, Jessica Ezemba, Jason Pohl, Conrad Tucker, and Christopher McComb of Carnegie Mellon&#8217;s Department of Mechanical Engineering, the dataset represents an 850-fold scale increase over what had previously been possible in this niche, enabling statistically robust evaluation across diverse simulation configurations and question types. That scale matters enormously. Small benchmarks can produce results that look impressive but dissolve under statistical scrutiny; with hundreds of thousands of question-answer pairs, the researchers could measure model performance with enough statistical power to detect even incremental progress, and to show confidently when there was none.</p>
<p>The motivation behind the work stems from a genuine bottleneck in engineering design cycles. Interpreting simulation outputs, whether from structural analyses, fluid dynamics studies, or thermal models, requires expensive domain expertise. Engineers must validate complex outputs to ensure safety and performance, and that validation step can slow product development considerably. Large language models have been proposed as assistants in this interpretation task, but they face a fundamental scalability limitation: even modest simulations produce data that exceed the context windows of the best-in-class LLMs. A simulation that generates millions of data points simply cannot be fed into a model that processes text sequentially, no matter how large its context window has grown.</p>
<p>Vision-language models offer a promising alternative, and the reasoning behind that promise is elegant. Rather than ingesting raw simulation data, a VLM can process a visualization of the results, a contour plot, a deformed geometry, a field map, as a compressed representation of the underlying information. VLMs have already demonstrated success across technical visual reasoning domains, from medical imaging to materials characterization. If a model can identify anomalies in radiographs or classify microstructures in electron micrographs, why not read a stress distribution from a finite element plot? The problem, the Carnegie Mellon team found, was that nobody had rigorously tested this assumption, largely because large-scale evaluation frameworks did not exist and expert annotation of simulation outputs was prohibitively expensive.</p>
<p>OpenSeeSimE addresses both obstacles through parametric generation. By systematically varying simulation parameters across 10,000 simulations, the researchers could produce an enormous volume of question-answer pairs without relying on scarce expert annotators to label each example individually. The parametric approach also ensures diversity: the benchmark covers a wide range of simulation configurations and question types, so a model cannot succeed by memorizing a handful of visual patterns. The team built the simulation pipeline using PyAnsys, and the authors acknowledge support and guidance from Raunak Borker, Sanjay Ranganayakulu, Chris Hawkins, and the broader Ansys team in utilizing and understanding the software.</p>
<p>The evaluation itself covered ten state-of-the-art vision-language models, and the findings were striking in their consistency. Models that demonstrate strong performance on general visual reasoning benchmarks, the very tasks that have fueled excitement about multimodal AI, scored between 29 and 47 percent on the engineering simulation questions. For context, random chance on multiple-choice questions typically falls in that range, meaning the models were effectively guessing. The effect sizes were negligible, indicating that the differences between models were not meaningful and that none of them had developed even a rudimentary grasp of simulation interpretation.</p>
<p>This disconnect between general visual competence and domain-specific failure carries a broader lesson about the current generation of AI systems. The pattern recognition abilities that allow a VLM to caption a photograph or answer questions about natural images do not automatically transfer to the stylized, information-dense visual language of engineering analysis. Contour plots encode quantitative information through color scales that must be read precisely; meshed geometries convey structural information through conventions that differ sharply from natural scenes. A model trained predominantly on internet images and text has little exposure to these conventions, and the benchmark results suggest that exposure to general visual data provides essentially no foundation for simulation literacy.</p>
<p>The practical implication, according to the authors, is that deploying vision-language models for simulation interpretation will require domain-specific training rather than reliance on general-purpose models. Off-the-shelf multimodal assistants, however impressive their demo performances, cannot be trusted to read a simulation output today. But the benchmark also provides the infrastructure needed to change that. Because OpenSeeSimE is large enough to measure incremental progress with statistical confidence, researchers developing domain-adapted models can use it to track whether fine-tuning and specialized training data are actually working, rather than relying on small evaluation sets that produce noisy, unreliable signals.</p>
<p>The study also establishes critical baselines for the field. Knowing that current models perform at chance levels gives future researchers a clear starting point against which to measure improvement. The benchmark&#8217;s reusable framework, with its 850-fold scale increase over prior evaluation efforts, is designed to serve as a lasting resource for the community. As AI systems are increasingly proposed for safety-critical engineering tasks, from validating structural designs to checking thermal performance, rigorous evaluation of their actual capabilities becomes not just scientifically interesting but essential for responsible deployment.</p>
<p>Published open access in Communications Engineering on 23 September 2026, the paper arrives at a moment when the gap between AI hype and AI capability in specialized domains is under intense scrutiny. The Carnegie Mellon work adds engineering simulation to the growing list of technical areas, alongside medicine, law, and scientific research, where general-purpose models stumble without targeted training. For engineers hoping that AI might soon ease the burden of interpreting complex simulation outputs, the message is clear: the tools are not ready yet, but for the first time, there is a rigorous, large-scale yardstick to measure how close they are getting.</p>
<p><strong>Subject of Research:</strong> Benchmark evaluation of vision-language models for engineering simulation question answering</p>
<p><strong>Article Title:</strong> A large-scale benchmark to assess vision-language model question answering capabilities in engineering simulations</p>
<p><strong>Article References:</strong> Ezemba, J., Pohl, J., Tucker, C., &amp; McComb, C. (2026). A large-scale benchmark to assess vision-language model question answering capabilities in engineering simulations. <em>Communications Engineering</em>. <a href="https://doi.org/10.1038/s44172-026-00786-2" rel="noopener noreferrer">https://doi.org/10.1038/s44172-026-00786-2</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s44172-026-00786-2" rel="noopener noreferrer">10.1038/s44172-026-00786-2</a></p>
<p><strong>Keywords:</strong> vision-language models, engineering simulations, benchmark, OpenSeeSimE, question answering, machine learning, mechanical engineering, Carnegie Mellon University, visual reasoning, finite element analysis, AI evaluation, Communications Engineering</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">252249</post-id>	</item>
		<item>
		<title>New Benchmark Puts AI Models to the Test on Real Math Problems</title>
		<link>https://scienmag.com/new-benchmark-puts-ai-models-to-the-test-on-real-math-problems/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 21:45:58 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[advancements in AI mathematical reasoning]]></category>
		<category><![CDATA[AI mathematical reasoning benchmark]]></category>
		<category><![CDATA[answer grading]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[benchmark]]></category>
		<category><![CDATA[comparison of AI math performance]]></category>
		<category><![CDATA[comprehensive machine learning evaluation]]></category>
		<category><![CDATA[cross-lingual math problem datasets]]></category>
		<category><![CDATA[data contamination]]></category>
		<category><![CDATA[DeepSeek]]></category>
		<category><![CDATA[Education]]></category>
		<category><![CDATA[evaluation metrics]]></category>
		<category><![CDATA[evaluation of large language models on math tasks]]></category>
		<category><![CDATA[Gaokao]]></category>
		<category><![CDATA[GPT-4]]></category>
		<category><![CDATA[improving consistency in AI math assessments]]></category>
		<category><![CDATA[integration of diverse math datasets]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[mathematical reasoning]]></category>
		<category><![CDATA[MathEval]]></category>
		<category><![CDATA[MathEval benchmark for AI models]]></category>
		<category><![CDATA[multi-step problem solving in AI]]></category>
		<category><![CDATA[standardization of math AI testing]]></category>
		<category><![CDATA[testing higher mathematics comprehension in AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=229195</guid>

					<description><![CDATA[A new benchmark called MathEval unifies 22 mathematical datasets and uses annually refreshed Gaokao exam problems to measure large language models' true reasoning abilities.]]></description>
										<content:encoded><![CDATA[<p>Mathematical reasoning has long been treated as a kind of litmus test for machine intelligence. A system that can parse a word problem, plan a multi-step solution, and carry out exact arithmetic without slipping is doing something qualitatively different from predicting the next word in a sentence. Yet as large language models have grown more capable, the research community has struggled to agree on just how good they actually are at mathematics. Different papers use different datasets, different prompting strategies, and different grading rules, producing scores that are difficult to compare and sometimes wildly inconsistent. A new benchmark called MathEval, described in the journal Frontiers of Digital Education, aims to end that confusion with a single, comprehensive, and continuously refreshed testing ground for the mathematical minds of machines.</p>
<p>The benchmark, developed by Tianqiao Liu, Zui Chen, Zhensheng Fang, Weiqi Luo, Mi Tian, and Zitao Liu of the Guangdong Institute of Smart Education at Jinan University and TAL Education Group, consolidates 22 distinct datasets into one evaluation framework. The collection spans an enormous range of mathematical territory: elementary arithmetic and word problems, competition-level mathematics, and higher mathematics at university and beyond. It covers problems written in both English and Chinese, and it deliberately includes material at every difficulty level from primary school exercises to problems that challenge even strong human mathematicians. By drawing on well-known resources such as GSM8K, MATH, MathQA, MAWPS, Ape210K, OlympiadBench, and AGIEval, the benchmark situates itself within the existing landscape of evaluation research while unifying those scattered efforts under one roof.</p>
<p>The motivation for such a unification is straightforward. Previous assessments of language models on mathematics have been, in the words of the research team, inconsistent and incomplete. One study might report that a model solves 80 percent of grade-school word problems, while another finds the same model failing on similar material because the answer format differed, the prompt was phrased differently, or the grading script treated equivalent expressions as mismatched. Because mathematical outputs can take many valid forms — a fraction versus a decimal, a simplified radical versus an approximation, a solution set written in different notation — automatically comparing a model&#8217;s answer to a reference answer is far harder than it might appear. Small differences in extraction and comparison logic can shift reported accuracy by several percentage points, enough to change the ranking of competing models.</p>
<p>MathEval tackles this grading problem with an unusual and technically interesting solution: it uses GPT-4 itself as an automated pipeline for answer extraction and comparison. Rather than relying on brittle regular expressions or rigid string matching, the benchmark prompts GPT-4 to read a model&#8217;s full response, identify the final answer, and judge whether it is mathematically equivalent to the reference solution. This approach adapts to diverse models, diverse output styles, and diverse prompt formats, which is essential when evaluating dozens of systems that each express their conclusions differently. The team validated the reliability of this automated judging process, drawing on established statistical measures of agreement to confirm that the pipeline&#8217;s judgments are consistent and trustworthy.</p>
<p>There is, however, an obvious practical drawback to using GPT-4 as a grader: not every research group has reliable access to it, and running every comparison through a commercial frontier model is expensive. To solve this, the researchers trained a publicly available model, built on the DeepSeek-LLM-7B-Base architecture, using GPT-4&#8217;s comparison results as training data. The result is a compact answer-validation model that reproduces GPT-4&#8217;s grading behavior without requiring any access to GPT-4 itself. This distillation strategy means that any laboratory, anywhere, can run the full MathEval evaluation pipeline on local hardware, democratizing access to a high-quality assessment tool and lowering the barrier to rigorous, reproducible measurement of mathematical reasoning.</p>
<p>Perhaps the most forward-looking feature of MathEval is its defense against data contamination. A persistent worry in language model evaluation is that test problems leak into training data: web-scraped corpora inevitably contain popular benchmark datasets, so a model may appear to solve problems it has effectively memorized rather than reasoned through. This concern inflates scores and makes genuine progress impossible to measure. MathEval addresses it by incorporating an annually refreshed set of problems drawn from the most recent Chinese National College Entrance Examination, the Gaokao, including the 2023 and 2024 editions. Because these exam questions are newly written each year and administered under strict security, they cannot have contaminated any model&#8217;s training data at evaluation time. Performance on the Gaokao subset therefore offers a cleaner signal of true problem-solving ability, and the benchmark is designed to keep pace with each year&#8217;s exam, providing a moving target that models cannot simply memorize their way past.</p>
<p>The choice of the Gaokao is also scientifically apt. The examination is famous for its demanding mathematics section, which requires not only computation but also careful reading, multi-step planning, and the synthesis of ideas from algebra, geometry, trigonometry, and calculus. It is bilingual in practice, since the benchmark as a whole tests both English and Chinese, and cross-lingual evaluation matters: a model that excels on English problems may falter on equivalent Chinese ones, and vice versa. By embedding these fresh exam problems within a broader multilingual, multi-domain suite, MathEval can reveal whether a model&#8217;s mathematical competence is a genuine, transferable skill or a narrow, language-specific and dataset-specific trick.</p>
<p>The benchmark&#8217;s design also reflects a broader trend in artificial intelligence research toward holistic evaluation. Earlier efforts such as BIG-Bench, HELM, LongBench, and CMMLU demonstrated that single-number scores on isolated tasks tell an incomplete story about what language models can and cannot do. MathEval extends that philosophy to mathematics specifically, evaluating models across multiple dimensions simultaneously: mathematical discipline, language, problem category, and difficulty level. This multidimensional structure allows researchers to localize failures with precision. A model might handle arithmetic flawlessly yet collapse on olympiad geometry, or perform well in English while degrading in Chinese, or succeed on familiar textbook formats while stumbling on novel competition problems. Each of these failure patterns suggests a different underlying weakness and points toward different remedies in training data or model design.</p>
<p>The stakes of getting this measurement right extend well beyond leaderboard rivalry. Mathematical reasoning underpins applications that society increasingly wants to hand to AI systems: tutoring students, verifying scientific calculations, assisting engineers, and generating reliable quantitative analyses. The research team&#8217;s own related work on math reasoning in language models, on step-level reward models, and on chain-of-thought decoding shows how much active effort is being invested in improving these capabilities. But improvement cannot be demonstrated without measurement that is fair, consistent, and resistant to gaming. A benchmark that changes its test set annually, grades answers with validated automated judges, and spans the full spectrum of mathematical difficulty provides exactly the kind of stable yardstick that the field has lacked.</p>
<p>MathEval is already gaining traction in the research community, with the published paper accumulating citations and thousands of accesses since its release in May 2025. Its open availability, combined with the distilled answer-checking model that removes the GPT-4 dependency, makes it a practical standard that other groups can adopt immediately. As each new generation of language models arrives claiming better reasoning, the benchmark&#8217;s refreshed Gaokao problems will be waiting — unseen, unmemorized, and unforgiving. In a field where the gap between claimed and actual capability can be obscured by contaminated test sets and inconsistent grading, that kind of honest examination may prove to be the most valuable result of all.</p>
<p><strong>Subject of Research:</strong> Benchmarking the mathematical reasoning capabilities of large language models</p>
<p><strong>Article Title:</strong> MathEval: A Comprehensive Benchmark for Evaluating Large Language Models on Mathematical Reasoning Capabilities</p>
<p><strong>Article References:</strong> Liu, T., Chen, Z., Fang, Z., Luo, W., Tian, M., &amp; Liu, Z. (2025). MathEval: A Comprehensive Benchmark for Evaluating Large Language Models on Mathematical Reasoning Capabilities. <em>Frontiers of Digital Education, 2</em>(2), Article 16. <a href="https://doi.org/10.1007/s44366-025-0053-z" rel="noopener noreferrer">https://doi.org/10.1007/s44366-025-0053-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44366-025-0053-z" rel="noopener noreferrer">10.1007/s44366-025-0053-z</a></p>
<p><strong>Keywords:</strong> MathEval, large language models, mathematical reasoning, benchmark, GPT-4, DeepSeek, Gaokao, answer grading, data contamination, evaluation metrics, artificial intelligence, education</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">229195</post-id>	</item>
		<item>
		<title>Cold-Start Benchmark Exposes Limits of Drug Synergy Models</title>
		<link>https://scienmag.com/cold-start-benchmark-exposes-limits-of-drug-synergy-models/</link>
		
		<dc:creator><![CDATA[Louis Brooks]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 10:53:09 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[AI in pharmacology]]></category>
		<category><![CDATA[benchmark]]></category>
		<category><![CDATA[challenges in clinical translation of AI]]></category>
		<category><![CDATA[cold start]]></category>
		<category><![CDATA[computational biology]]></category>
		<category><![CDATA[data leakage]]></category>
		<category><![CDATA[data leakage in bioinformatics]]></category>
		<category><![CDATA[drug combination]]></category>
		<category><![CDATA[drug combination benchmarking]]></category>
		<category><![CDATA[drug discovery]]></category>
		<category><![CDATA[drug discovery computational methods]]></category>
		<category><![CDATA[drug synergy]]></category>
		<category><![CDATA[drug synergy prediction]]></category>
		<category><![CDATA[DrugComb dataset analysis]]></category>
		<category><![CDATA[evaluating drug synergy models]]></category>
		<category><![CDATA[generalization in drug synergy models]]></category>
		<category><![CDATA[leakage-controlled]]></category>
		<category><![CDATA[LightGBM]]></category>
		<category><![CDATA[limitations of current drug synergy benchmarks]]></category>
		<category><![CDATA[Machine learning]]></category>
		<category><![CDATA[machine learning drug interaction models]]></category>
		<category><![CDATA[model overfitting in pharmacology]]></category>
		<category><![CDATA[pharmacology]]></category>
		<category><![CDATA[ZIP score]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=227303</guid>

					<description><![CDATA[A new benchmark reveals that drug synergy prediction models often rely on memorizing historical data rather than learning generalizable pharmacological principles, significantly impacting their utility in real-world drug discovery.]]></description>
										<content:encoded><![CDATA[<p>Computational models designed to predict whether two drugs will work better together than alone are often evaluated under conditions that inadvertently allow them to memorize answers rather than learn underlying biological principles. A new study published in BMC Bioinformatics reveals that when machine learning algorithms are tested on standard datasets, their high accuracy scores may largely reflect a simple ability to recall historical outcomes for specific cell lines and drug pairs, rather than a genuine understanding of pharmacological interactions. This finding has significant implications for the development of artificial intelligence tools in drug discovery, suggesting that current benchmarks may be overstating the readiness of these models for real-world clinical applications where new combinations must be evaluated without prior experimental data.</p>
<p>The research, led by Xulin Pan from The Pennsylvania State University, introduces a leakage-controlled benchmark specifically designed to test how well synergy classification models perform when they are forced to generalize to entirely new entities. The study utilizes the DrugComb v1.5 dataset, a comprehensive repository containing 739,964 experiments across 26 studies, 17 tissues, 288 cell lines, and 4,268 compounds. By strictly excluding any measured dose-response readouts and relying solely on pre-treatment identity and context descriptors such as compound identities, cell line types, tissue origins, and clinical development phases, the researchers created a controlled environment to isolate the predictive power of metadata alone.</p>
<p>The core of the investigation focuses on the Zero Interaction Potency (ZIP) score, a standard metric used to quantify drug synergy. The researchers derived a three-class endpoint from this score using conventional thresholds of plus or minus 10, categorizing interactions as synergistic, additive, or antagonistic. Two primary models were evaluated under this framework: a logistic regression baseline and a class-weighted gradient-boosting model using the LightGBM algorithm. The evaluation was conducted under both predefined grouped splits and specific cold-start splits, which simulate scenarios where the model encounters drugs, cell lines, or entire studies it has never seen during training.</p>
<p>On the standard test set, the LightGBM model demonstrated strong performance, achieving an accuracy of 0.782 and a balanced accuracy of 0.805. It significantly outperformed the logistic regression baseline, a result confirmed by paired McNemar testing with a p-value less than 0.001. The model also achieved a macro-F1 score of 0.703 and a macro one-vs-rest area under the curve of 0.922. Precision-recall analysis indicated that the model provided roughly six-fold enrichment for the rare antagonistic and synergistic classes over their base rates, suggesting that it could effectively triage potential combinations for further experimental screening.</p>
<p>However, the study’s most critical findings emerge when the models are subjected to distribution shifts that mimic real-world deployment challenges. When the model was tested on unseen drugs, its balanced accuracy dropped to 0.694. The decline was more severe for unseen cell lines, where balanced accuracy fell to 0.482. Most strikingly, when the model was evaluated on entirely unseen studies, its performance collapsed to the random floor, indicating that it could not generalize beyond the specific experimental contexts it had learned. An ablation study confirmed that this failure was not driven by the study identifier itself, but rather by the lack of generalizable features across different experimental setups.</p>
<p>Permutation importance analysis identified cell-line identity as the dominant predictor in the model’s decision-making process. This finding led the researchers to test a simple baseline: a lookup table that retrieved base rates for each specific cell line and drug pair. Surprisingly, this non-parametric approach recovered a balanced accuracy of 0.778, which was within just 0.027 of the full LightGBM model’s performance. This proximity suggests that the majority of the model’s apparent success on standard splits is attributable to conditional memorization of entity-specific base rates rather than the learning of complex pharmacological relationships.</p>
<p>The implications of these results are profound for the field of computational pharmacology. The study argues that standard random or grouped splits can materially overstate the performance of synergy prediction models when they are deployed in cold-start scenarios. By providing a leakage-controlled lower bound, the benchmark offers a rigorous standard against which richer molecular and omics-based models can be measured. The authors emphasize that future evaluations should routinely include cold-start splits alongside conventional metrics to provide a more accurate picture of a model’s true generalization capabilities.</p>
<p>This work highlights a broader issue in machine learning for biology: the gap between in-distribution performance and out-of-distribution robustness. While high accuracy scores on standard benchmarks are often cited as evidence of a model’s utility, they may mask fundamental limitations in its ability to predict outcomes for novel combinations. As the pharmaceutical industry increasingly turns to AI to reduce the cost and time of drug discovery, the need for rigorous, leakage-controlled benchmarks becomes essential to avoid investing in models that cannot reliably predict the behavior of new drug pairs.</p>
<p>The study does not dismiss the value of machine learning in drug synergy prediction but rather refines the expectations for what these models can achieve. It demonstrates that useful synergy triage is achievable for known entities without the need for complex molecular features, but it warns that this capability is heavily dependent on the specific context of the training data. For new drugs or cell lines, the predictive power diminishes sharply, underscoring the need for models that can integrate deeper mechanistic insights or molecular descriptors that are invariant across different experimental conditions.</p>
<p>By establishing this benchmark, the researchers provide a critical tool for the scientific community to assess the true generalization potential of synergy prediction algorithms. The findings call for a shift in how models are evaluated, moving beyond simple accuracy metrics to include rigorous cold-start tests that reflect the realities of drug development. This approach will help ensure that computational tools are not only accurate in retrospective analyses but also reliable in prospective applications, ultimately contributing to more efficient and effective drug discovery processes.</p>
<p><strong>Subject of Research:</strong> Evaluation of machine learning generalization in drug-combination synergy prediction</p>
<p><strong>Article Title:</strong> A leakage-controlled cold-start benchmark for drug-combination synergy classification</p>
<p><strong>Article References:</strong> Pan, X. (2026). A leakage-controlled cold-start benchmark for drug-combination synergy classification. <em>BMC Bioinformatics</em>. <a href="https://doi.org/10.1186/s12859-026-06665-z" rel="noopener noreferrer">https://doi.org/10.1186/s12859-026-06665-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s12859-026-06665-z" rel="noopener noreferrer">10.1186/s12859-026-06665-z</a></p>
<p><strong>Keywords:</strong> drug synergy, machine learning, benchmark, cold-start, pharmacology, LightGBM, data leakage, drug discovery, computational biology, ZIP score, leakage-controlled, drug-combination</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">227303</post-id>	</item>
		<item>
		<title>New Benchmark Puts Spatial Transcriptomics Batch Correction Methods to the Test</title>
		<link>https://scienmag.com/new-benchmark-puts-spatial-transcriptomics-batch-correction-methods-to-the-test/</link>
		
		<dc:creator><![CDATA[Brooke Gardner]]></dc:creator>
		<pubDate>Sun, 20 Sep 2026 19:34:55 +0000</pubDate>
				<category><![CDATA[Biology]]></category>
		<category><![CDATA[batch effect]]></category>
		<category><![CDATA[batch effect correction]]></category>
		<category><![CDATA[benchmark]]></category>
		<category><![CDATA[benchmarking]]></category>
		<category><![CDATA[biological signal preservation]]></category>
		<category><![CDATA[cross-platform integration]]></category>
		<category><![CDATA[data integration]]></category>
		<category><![CDATA[gene expression mapping in tissues]]></category>
		<category><![CDATA[Genome Biology]]></category>
		<category><![CDATA[multi-platform spatial omics integration]]></category>
		<category><![CDATA[reproducibility]]></category>
		<category><![CDATA[reproducibility in spatial biology]]></category>
		<category><![CDATA[semi-synthetic simulation]]></category>
		<category><![CDATA[SpaBEAT]]></category>
		<category><![CDATA[SpaBEAT method for batch correction]]></category>
		<category><![CDATA[spatial batch effect assessment]]></category>
		<category><![CDATA[spatial omics data analysis]]></category>
		<category><![CDATA[Spatial transcriptomics]]></category>
		<category><![CDATA[spatial transcriptomics data benchmarking]]></category>
		<category><![CDATA[spatial transcriptomics technical challenges]]></category>
		<category><![CDATA[systematic]]></category>
		<category><![CDATA[tissue sample processing variability]]></category>
		<category><![CDATA[tissue slice processing artifacts]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=201739</guid>

					<description><![CDATA[A new benchmarking framework called SpaBEAT defines four types of batch effects in spatial transcriptomics and shows that no correction method is universally optimal.]]></description>
										<content:encoded><![CDATA[<p>Spatial transcriptomics has rapidly become one of the most transformative technologies in modern biology, allowing researchers to map gene expression across intact tissue slices at increasingly high resolution. By revealing not only which genes are active but precisely where they are active, the technology has reshaped studies of tumor architecture, brain organization, and development. Yet a stubborn technical problem has shadowed this progress: batch effects, the systematic distortions introduced when samples are processed in different runs, on different platforms, or with different protocols. These artifacts can masquerade as biological signals, corrupt integrated analyses, and undermine reproducibility. A major new study published in Genome Biology now tackles the problem head-on with a systematic benchmark designed to define, measure, and correct batch effects in spatial omics data.</p>
<p>The research, led by Minghui Zhao, Yingxin Zhang, Ming Jing, Na Zhou, and colleagues at Shandong University together with collaborators at Shandong Women&#8217;s University, introduces SpaBEAT, short for Spatial Batch Effect Assessment and Testing. Rather than treating batch effects as a single vague nuisance, the team breaks them down into four distinct categories that capture the real-world ways spatial datasets diverge from one another. Inter-slice effects arise between individual tissue slices processed separately, even from the same sample. Inter-sample effects reflect differences between distinct biological specimens that may be entangled with genuine biological variation. Cross-protocol and cross-platform effects emerge when data generated with different chemistries, instruments, or assay designs must be combined. Finally, intra-slice effects occur within a single tissue slice, for example when different regions are imaged or sequenced under subtly different conditions.</p>
<p>This taxonomy matters because each type of batch effect poses a different analytical challenge. Correcting differences between two slides of the same tumor is fundamentally easier than integrating data from a spot-based platform with data from a high-resolution imaging platform, where technical differences can be far larger than the biological differences researchers hope to detect. By explicitly separating these scenarios, SpaBEAT gives the field a common language for discussing what exactly a correction method is supposed to accomplish, and it exposes the uncomfortable truth that a tool excelling in one scenario may fail badly in another.</p>
<p>Using this framework, the authors benchmarked ten spatial integration methods across a strikingly diverse collection of spatial transcriptomics modalities. Their test bed included spot-based datasets, which profile gene expression at discrete spatial locations; high-resolution platforms that approach single-cell or sub-cellular resolution; image-based targeted assays that measure carefully chosen panels of genes; and cross-platform datasets that require harmonizing fundamentally different measurement technologies. This breadth is unusual among published benchmarks, which have often relied on a handful of convenient datasets. By spanning modalities, the study provides a far more realistic picture of how correction methods behave in the heterogeneous settings where working biologists actually deploy them.</p>
<p>A central methodological innovation of the study lies in its use of controlled and semi-synthetic simulations. Evaluating batch correction on real data alone is plagued by a circularity problem: researchers rarely know with certainty which expression differences are technical and which are biological, so they cannot tell whether a method has removed noise or destroyed signal. The Shandong team sidestepped this trap by constructing simulations in which technical variation was deliberately injected while predefined biological differences were preserved and known in advance. This allowed them to disentangle the two sources of variation with unprecedented clarity, quantifying precisely how much of any observed improvement reflected genuine batch removal rather than opportunistic erasure of biology. The simulations also enabled systematic stress tests of robustness, probing how sensitive each method was to preprocessing choices, to the degree of overlap in targeted gene panels, and to different cell-segmentation strategies used to convert images into cell-level expression matrices.</p>
<p>Performance was scored with a rigorous panel of metrics capturing the two pillars of any integration task: batch-effect removal and biological signal preservation. The first asks whether the method successfully merges data from different batches so that technical artifacts no longer dominate the structure of the data. The second asks whether known biological features, such as tissue domains, cell-type identities, and spatially patterned gene programs, survive the correction intact. The authors combined these metrics into a hierarchical ranking scheme that also accounted for task coverage, how many scenarios a method could handle at all, and computational efficiency, an increasingly practical concern as spatial datasets swell to millions of measured locations.</p>
<p>The headline finding is sobering but clarifying: no single method is universally optimal. Performance proved strongly context-dependent, with every tool exhibiting distinct trade-offs between the aggressiveness with which it removes batch effects and its fidelity in preserving genuine biological structure. A method that aggressively harmonizes batches may blur the very boundaries between tissue domains that researchers care about, while a conservative method may leave technical gradients that contaminate downstream clustering, trajectory inference, or differential expression analyses. Which trade-off is acceptable depends on the tissue, the platform, the batch-effect scenario, and the biological question at hand. The benchmark makes clear that choosing a correction method is not a matter of finding a champion but of matching tool characteristics to task requirements.</p>
<p>The practical implications for the spatial omics community are substantial. Researchers planning multi-slice or multi-platform studies can now consult empirically grounded guidance on which integration approaches perform best under their specific conditions, rather than relying on intuition, convenience, or outdated single-cell benchmarks that may not transfer to spatial data. The study also highlights the importance of robustness checks: because method performance shifted with preprocessing decisions and segmentation strategies, the authors caution that integration results should be validated across reasonable alternative pipelines before being trusted. For the growing number of atlas-scale projects that aggregate data from many laboratories, the findings suggest that careful attention to batch design should begin at the experimental stage, not after the data have been collected.</p>
<p>Beyond its immediate results, the study delivers a durable piece of community infrastructure. The authors have released benchmark datasets, simulation frameworks, reproducible workflows, and evaluation resources, effectively transforming what was previously scattered tacit knowledge into a standardized, testable framework. This open approach mirrors successful benchmarking efforts in single-cell genomics and machine learning, where shared evaluation standards have accelerated method development by making strengths and weaknesses transparent and comparable. New spatial integration methods can now be assessed against the same yardsticks, on the same data, under the same definitions of success, raising the bar for claims of improved performance.</p>
<p>As spatial transcriptomics moves from specialized core facilities toward routine clinical and translational use, the stakes of proper batch correction will only rise. Multi-center studies of tumors, organoids, and diseased tissues will depend on confidently distinguishing true spatial biology from technical artifacts, and flawed integration could distort biomarker discovery or mislead therapeutic interpretations. By defining the problem precisely, quantifying the trade-offs honestly, and providing the tools for others to test and improve upon its conclusions, the SpaBEAT framework marks an important step toward spatial genomics results that are not only beautiful maps of biology but also reliable, reproducible ones. For a field racing to map the molecular architecture of life in place, knowing exactly where the technical noise lies may prove as important as knowing where the genes are.</p>
<p><strong>Subject of Research:</strong> Systematic benchmarking of batch effect correction methods for spatial transcriptomics</p>
<p><strong>Article Title:</strong> A systematic benchmark of batch effect correction methods for spatial transcriptomics</p>
<p><strong>Article References:</strong> Zhao, M., Zhang, Y., Jing, M., Zhou, N., Liu, X., Wang, X., Liu, R., Yuan, G., Xue, F., &amp; Hou, Q. (2026). A systematic benchmark of batch effect correction methods for spatial transcriptomics. <em>Genome Biology</em>. <a href="https://doi.org/10.1186/s13059-026-04281-x" rel="noopener noreferrer">https://doi.org/10.1186/s13059-026-04281-x</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1186/s13059-026-04281-x" rel="noopener noreferrer">10.1186/s13059-026-04281-x</a></p>
<p><strong>Keywords:</strong> spatial transcriptomics, batch effect, data integration, benchmarking, SpaBEAT, cross-platform integration, semi-synthetic simulation, biological signal preservation, Genome Biology, reproducibility, systematic, benchmark</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">201739</post-id>	</item>
	</channel>
</rss>
