<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>evaluation of large language models on math tasks &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/evaluation-of-large-language-models-on-math-tasks/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Fri, 02 Oct 2026 21:45:58 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>evaluation of large language models on math tasks &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Benchmark Puts AI Models to the Test on Real Math Problems</title>
		<link>https://scienmag.com/new-benchmark-puts-ai-models-to-the-test-on-real-math-problems/</link>
		
		<dc:creator><![CDATA[Courtney Benton]]></dc:creator>
		<pubDate>Fri, 02 Oct 2026 21:45:58 +0000</pubDate>
				<category><![CDATA[Social Science]]></category>
		<category><![CDATA[advancements in AI mathematical reasoning]]></category>
		<category><![CDATA[AI mathematical reasoning benchmark]]></category>
		<category><![CDATA[answer grading]]></category>
		<category><![CDATA[Artificial Intelligence]]></category>
		<category><![CDATA[benchmark]]></category>
		<category><![CDATA[comparison of AI math performance]]></category>
		<category><![CDATA[comprehensive machine learning evaluation]]></category>
		<category><![CDATA[cross-lingual math problem datasets]]></category>
		<category><![CDATA[data contamination]]></category>
		<category><![CDATA[DeepSeek]]></category>
		<category><![CDATA[Education]]></category>
		<category><![CDATA[evaluation metrics]]></category>
		<category><![CDATA[evaluation of large language models on math tasks]]></category>
		<category><![CDATA[Gaokao]]></category>
		<category><![CDATA[GPT-4]]></category>
		<category><![CDATA[improving consistency in AI math assessments]]></category>
		<category><![CDATA[integration of diverse math datasets]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[mathematical reasoning]]></category>
		<category><![CDATA[MathEval]]></category>
		<category><![CDATA[MathEval benchmark for AI models]]></category>
		<category><![CDATA[multi-step problem solving in AI]]></category>
		<category><![CDATA[standardization of math AI testing]]></category>
		<category><![CDATA[testing higher mathematics comprehension in AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=229195</guid>

					<description><![CDATA[A new benchmark called MathEval unifies 22 mathematical datasets and uses annually refreshed Gaokao exam problems to measure large language models' true reasoning abilities.]]></description>
										<content:encoded><![CDATA[<p>Mathematical reasoning has long been treated as a kind of litmus test for machine intelligence. A system that can parse a word problem, plan a multi-step solution, and carry out exact arithmetic without slipping is doing something qualitatively different from predicting the next word in a sentence. Yet as large language models have grown more capable, the research community has struggled to agree on just how good they actually are at mathematics. Different papers use different datasets, different prompting strategies, and different grading rules, producing scores that are difficult to compare and sometimes wildly inconsistent. A new benchmark called MathEval, described in the journal Frontiers of Digital Education, aims to end that confusion with a single, comprehensive, and continuously refreshed testing ground for the mathematical minds of machines.</p>
<p>The benchmark, developed by Tianqiao Liu, Zui Chen, Zhensheng Fang, Weiqi Luo, Mi Tian, and Zitao Liu of the Guangdong Institute of Smart Education at Jinan University and TAL Education Group, consolidates 22 distinct datasets into one evaluation framework. The collection spans an enormous range of mathematical territory: elementary arithmetic and word problems, competition-level mathematics, and higher mathematics at university and beyond. It covers problems written in both English and Chinese, and it deliberately includes material at every difficulty level from primary school exercises to problems that challenge even strong human mathematicians. By drawing on well-known resources such as GSM8K, MATH, MathQA, MAWPS, Ape210K, OlympiadBench, and AGIEval, the benchmark situates itself within the existing landscape of evaluation research while unifying those scattered efforts under one roof.</p>
<p>The motivation for such a unification is straightforward. Previous assessments of language models on mathematics have been, in the words of the research team, inconsistent and incomplete. One study might report that a model solves 80 percent of grade-school word problems, while another finds the same model failing on similar material because the answer format differed, the prompt was phrased differently, or the grading script treated equivalent expressions as mismatched. Because mathematical outputs can take many valid forms — a fraction versus a decimal, a simplified radical versus an approximation, a solution set written in different notation — automatically comparing a model&#8217;s answer to a reference answer is far harder than it might appear. Small differences in extraction and comparison logic can shift reported accuracy by several percentage points, enough to change the ranking of competing models.</p>
<p>MathEval tackles this grading problem with an unusual and technically interesting solution: it uses GPT-4 itself as an automated pipeline for answer extraction and comparison. Rather than relying on brittle regular expressions or rigid string matching, the benchmark prompts GPT-4 to read a model&#8217;s full response, identify the final answer, and judge whether it is mathematically equivalent to the reference solution. This approach adapts to diverse models, diverse output styles, and diverse prompt formats, which is essential when evaluating dozens of systems that each express their conclusions differently. The team validated the reliability of this automated judging process, drawing on established statistical measures of agreement to confirm that the pipeline&#8217;s judgments are consistent and trustworthy.</p>
<p>There is, however, an obvious practical drawback to using GPT-4 as a grader: not every research group has reliable access to it, and running every comparison through a commercial frontier model is expensive. To solve this, the researchers trained a publicly available model, built on the DeepSeek-LLM-7B-Base architecture, using GPT-4&#8217;s comparison results as training data. The result is a compact answer-validation model that reproduces GPT-4&#8217;s grading behavior without requiring any access to GPT-4 itself. This distillation strategy means that any laboratory, anywhere, can run the full MathEval evaluation pipeline on local hardware, democratizing access to a high-quality assessment tool and lowering the barrier to rigorous, reproducible measurement of mathematical reasoning.</p>
<p>Perhaps the most forward-looking feature of MathEval is its defense against data contamination. A persistent worry in language model evaluation is that test problems leak into training data: web-scraped corpora inevitably contain popular benchmark datasets, so a model may appear to solve problems it has effectively memorized rather than reasoned through. This concern inflates scores and makes genuine progress impossible to measure. MathEval addresses it by incorporating an annually refreshed set of problems drawn from the most recent Chinese National College Entrance Examination, the Gaokao, including the 2023 and 2024 editions. Because these exam questions are newly written each year and administered under strict security, they cannot have contaminated any model&#8217;s training data at evaluation time. Performance on the Gaokao subset therefore offers a cleaner signal of true problem-solving ability, and the benchmark is designed to keep pace with each year&#8217;s exam, providing a moving target that models cannot simply memorize their way past.</p>
<p>The choice of the Gaokao is also scientifically apt. The examination is famous for its demanding mathematics section, which requires not only computation but also careful reading, multi-step planning, and the synthesis of ideas from algebra, geometry, trigonometry, and calculus. It is bilingual in practice, since the benchmark as a whole tests both English and Chinese, and cross-lingual evaluation matters: a model that excels on English problems may falter on equivalent Chinese ones, and vice versa. By embedding these fresh exam problems within a broader multilingual, multi-domain suite, MathEval can reveal whether a model&#8217;s mathematical competence is a genuine, transferable skill or a narrow, language-specific and dataset-specific trick.</p>
<p>The benchmark&#8217;s design also reflects a broader trend in artificial intelligence research toward holistic evaluation. Earlier efforts such as BIG-Bench, HELM, LongBench, and CMMLU demonstrated that single-number scores on isolated tasks tell an incomplete story about what language models can and cannot do. MathEval extends that philosophy to mathematics specifically, evaluating models across multiple dimensions simultaneously: mathematical discipline, language, problem category, and difficulty level. This multidimensional structure allows researchers to localize failures with precision. A model might handle arithmetic flawlessly yet collapse on olympiad geometry, or perform well in English while degrading in Chinese, or succeed on familiar textbook formats while stumbling on novel competition problems. Each of these failure patterns suggests a different underlying weakness and points toward different remedies in training data or model design.</p>
<p>The stakes of getting this measurement right extend well beyond leaderboard rivalry. Mathematical reasoning underpins applications that society increasingly wants to hand to AI systems: tutoring students, verifying scientific calculations, assisting engineers, and generating reliable quantitative analyses. The research team&#8217;s own related work on math reasoning in language models, on step-level reward models, and on chain-of-thought decoding shows how much active effort is being invested in improving these capabilities. But improvement cannot be demonstrated without measurement that is fair, consistent, and resistant to gaming. A benchmark that changes its test set annually, grades answers with validated automated judges, and spans the full spectrum of mathematical difficulty provides exactly the kind of stable yardstick that the field has lacked.</p>
<p>MathEval is already gaining traction in the research community, with the published paper accumulating citations and thousands of accesses since its release in May 2025. Its open availability, combined with the distilled answer-checking model that removes the GPT-4 dependency, makes it a practical standard that other groups can adopt immediately. As each new generation of language models arrives claiming better reasoning, the benchmark&#8217;s refreshed Gaokao problems will be waiting — unseen, unmemorized, and unforgiving. In a field where the gap between claimed and actual capability can be obscured by contaminated test sets and inconsistent grading, that kind of honest examination may prove to be the most valuable result of all.</p>
<p><strong>Subject of Research:</strong> Benchmarking the mathematical reasoning capabilities of large language models</p>
<p><strong>Article Title:</strong> MathEval: A Comprehensive Benchmark for Evaluating Large Language Models on Mathematical Reasoning Capabilities</p>
<p><strong>Article References:</strong> Liu, T., Chen, Z., Fang, Z., Luo, W., Tian, M., &amp; Liu, Z. (2025). MathEval: A Comprehensive Benchmark for Evaluating Large Language Models on Mathematical Reasoning Capabilities. <em>Frontiers of Digital Education, 2</em>(2), Article 16. <a href="https://doi.org/10.1007/s44366-025-0053-z" rel="noopener noreferrer">https://doi.org/10.1007/s44366-025-0053-z</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44366-025-0053-z" rel="noopener noreferrer">10.1007/s44366-025-0053-z</a></p>
<p><strong>Keywords:</strong> MathEval, large language models, mathematical reasoning, benchmark, GPT-4, DeepSeek, Gaokao, answer grading, data contamination, evaluation metrics, artificial intelligence, education</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">229195</post-id>	</item>
	</channel>
</rss>
