Friday, October 9, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Mathematics

AI System Cracks Decades-Old Math Problems Using Machine-Verified Proofs

October 9, 2026
in Mathematics
Reid Dalton
By Reid Dalton Scienmag Editorial Profile - Applied Mathematics
Reading Time: 4 mins read
0
AI System Cracks Decades-Old Math Problems Using Machine-Verified Proofs

AI System Cracks Decades-Old Math Problems Using Machine-Verified Proofs

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Mathematicians have long dreamed of a computational collaborator that could not only calculate but genuinely reason — one that could explore the far reaches of open problems and return with arguments that hold up under the strictest scrutiny. That dream moved measurably closer to reality this week, as researchers unveiled AlphaProof Nexus, a large language model-based artificial intelligence framework capable of autonomously searching for formal mathematical proofs and verifying their logical soundness without human intervention. In a study published in Science, the team reports that the system solved dozens of previously open problems spanning several branches of mathematics, including two Erdős problems that had resisted resolution for more than half a century.

The central obstacle the researchers set out to overcome is one that anyone who has worked with modern AI systems will recognize: large language models are remarkably capable but chronically unreliable. They can sketch plausible solutions to difficult problems, yet they are prone to subtle logical errors and outright hallucinations — confident assertions that simply are not true. In everyday applications, such mistakes are annoying. In research mathematics, where a single flawed step invalidates an entire proof, they are disqualifying. Without extensive review by human experts, LLM-generated mathematics has remained untrustworthy at the level of rigor that the discipline demands.

The solution adopted by the team, led by George Tsoukalas, is elegantly pragmatic: let the AI write its mathematics in a formal language where correctness is not a matter of opinion. AlphaProof Nexus generates proofs in Lean, a formal proof programming language in which every logical step is checked automatically by a compiler. If a step does not follow from what came before, the compiler rejects it. This architecture creates a closed loop in which the language model proposes, the verifier disposes, and only arguments that survive machine-level scrutiny can be counted as genuine proofs. Errors cannot slip through unnoticed, because the verification environment is unforgiving by design.

Formal verification of this kind is not new in itself. Proof assistants such as Lean have been successfully applied to competition mathematics — the kind of olympiad-style problems that have become a standard benchmark for AI reasoning systems — and to the human-aided formalization of natural-language mathematical arguments, in which mathematicians translate existing proofs into machine-checkable form. What remained genuinely unknown was whether this approach could scale to open research-level problems: the messy, uncharted territory where no solution exists to be checked against, and where even the path toward a proof is unclear.

To close that gap, the researchers built AlphaProof Nexus as a multi-agent framework. Rather than relying on a single model attempting each problem in isolation, the system deploys multiple AI agents that search the space of possible proofs, with each attempt generating feedback from the Lean compiler. Failed attempts are not wasted; the compiler’s responses tell the agents where an argument breaks down, guiding subsequent exploration. This iterative dialogue between generation and verification allows the system to progressively refine its strategies, much as a human mathematician learns from a proof that collapses halfway through.

At the top of this architecture sits a more advanced, full-featured agent that coordinates the subagents through an evolutionary algorithm, treating AlphaProof itself as a specialized proof tool. In evolutionary terms, candidate proof strategies compete, the most promising ones are selected and recombined, and successive generations of approaches grow more effective. The result is a system that does not merely answer questions but organizes an entire search process — allocating effort, pruning dead ends, and directing specialized resources toward the hardest parts of a problem.

The benchmark results are striking. Out of 353 attempted Erdős problems — questions posed by the legendary mathematician Paul Erdős that collectively form a barometer of difficulty across combinatorics and number theory — AlphaProof Nexus solved nine, including two that had remained open for more than 50 years. The system also resolved 44 of 492 open conjectures drawn from the On-Line Encyclopedia of Integer Sequences, a vast community-maintained repository where patterns in integer sequences often point toward deep unsolved questions. Beyond these collections, the framework tackled and solved several other research-level problems in fields as diverse as algebraic geometry, optimization, quantum optics, and graph theory.

Perhaps the most intriguing implication, however, concerns the problems the system did not solve. In a related Perspective accompanying the research, mathematicians Jeremy Avigad and Matthew Ballard highlight the authors’ observation that even unsuccessful proof attempts by the AI could help researchers understand the problems better and make progress toward solving them. A failed machine-generated proof is not merely a null result; it maps out territory where an argument cannot easily go, exposes hidden structure in a problem, and can suggest new angles for human investigators. In this view, the value of AI in mathematics extends beyond the binary of solved and unsolved to something closer to genuine mathematical insight.

Avigad and Ballard frame this as a guiding principle for the field: an essential goal of developing AI for mathematics is to support mathematicians in the search for knowledge and understanding that lead to further advances. That framing matters. It positions systems like AlphaProof Nexus not as replacements for human mathematicians but as instruments — powerful, tireless, and rigorously honest about what they can actually prove — that extend the reach of human inquiry. The machine guarantees logical validity; the human supplies meaning, context, and the judgment about which questions are worth asking in the first place.

The work arrives at a moment of intense debate about the role of AI in fundamental research, and it offers a concrete answer to skeptics who have dismissed language models as sophisticated pattern matchers incapable of real reasoning. By anchoring every claim in machine-verified formal proof, AlphaProof Nexus sidesteps the trust problem that has plagued AI-generated mathematics, and its record against genuinely open problems suggests that automated mathematical discovery has moved from thought experiment to working reality. For a discipline whose practitioners have spent centuries perfecting the art of certainty, the arrival of a tireless collaborator that never asserts anything it cannot prove may prove to be one of the most consequential tools mathematics has ever gained.

Subject of Research: AI-driven formal mathematical proof discovery using large language models and automated verification

Article Title: Introducing AlphaProof Nexus: An AI tool for formal mathematical proof discovery

Article References: Introducing AlphaProof Nexus: An AI tool for formal mathematical proof discovery. (n.d.). Original publication

Image Credits: AI Generated

DOI: Not provided

Keywords: AlphaProof Nexus, artificial intelligence, large language models, formal proof verification, Lean programming language, Erdős problems, On-Line Encyclopedia of Integer Sequences, automated mathematical discovery, algebraic geometry, graph theory, quantum optics, Science journal

Cite Scienmag News

Reid Dalton. (October 9, 2026). AI System Cracks Decades-Old Math Problems Using Machine-Verified Proofs. Scienmag. https://scienmag.com/ai-system-cracks-decades-old-math-problems-using-machine-verified-proofs/

Reid Dalton. "AI System Cracks Decades-Old Math Problems Using Machine-Verified Proofs." Scienmag, 9 October 2026, https://scienmag.com/ai-system-cracks-decades-old-math-problems-using-machine-verified-proofs/. Accessed 9 October 2026.

Reid Dalton. "AI System Cracks Decades-Old Math Problems Using Machine-Verified Proofs." Scienmag. October 9, 2026. https://scienmag.com/ai-system-cracks-decades-old-math-problems-using-machine-verified-proofs/

Tags: AI-assisted mathematical discoveryAI-driven mathematical proof verificationalgebraic geometryAlphaProof NexusAlphaProof Nexus AI systemapplication of AI to Erdős problemsArtificial Intelligenceautomated mathematical discoveryautomation in mathematical researchautonomous problem-solving in mathematicsErdős problemsformal proof exploration using artificial intelligenceformal proof verificationgraph theorylarge language modelslarge language models in mathematicsLean programming languagelogical soundness in AI-generated proofsmachine-verified formal proofsOn-Line Encyclopedia of Integer Sequencesovercoming reliability issues in AI mathematicsquantum opticsresolving long-standing mathematical open problemsScience journal
Share26Tweet16
Previous Post

Black Holes Hide Supercritical Phase Boundaries in the Complex Plane, Study Finds

Next Post

How Tumors Survive Radiation: Plastic Cells and Tolerant Niches Drive Recurrence

Related Posts

Agentic AI Promises Autonomy, But Hallucinations and Scaling Woes Stand in the Way
Mathematics

Agentic AI Promises Autonomy, But Hallucinations and Scaling Woes Stand in the Way

October 9, 2026
New Statistical Test Puts Climate Models Under a Sharper Microscope
Climate

New Statistical Test Puts Climate Models Under a Sharper Microscope

October 9, 2026
How Few Ensemble Members Does a Weather-Style Forecast Filter Really Need? Chaos Sets the Limit
Earth Science

How Few Ensemble Members Does a Weather-Style Forecast Filter Really Need? Chaos Sets the Limit

October 9, 2026
AI-Powered Statistics Project Aims to Decode the Genetic Switches Behind Eye Disease
Mathematics

AI-Powered Statistics Project Aims to Decode the Genetic Switches Behind Eye Disease

October 9, 2026
Smarter Fuel Breaks: New Optimization Model Fights Wildfire While Protecting Caribou Routes
Mathematics

Smarter Fuel Breaks: New Optimization Model Fights Wildfire While Protecting Caribou Routes

October 9, 2026
New Statistical Study Reveals How Climate Change Reshapes Both Light and Extreme Rainfall
Climate

New Statistical Study Reveals How Climate Change Reshapes Both Light and Extreme Rainfall

October 9, 2026
Next Post
How Tumors Survive Radiation: Plastic Cells and Tolerant Niches Drive Recurrence

How Tumors Survive Radiation: Plastic Cells and Tolerant Niches Drive Recurrence

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • How Tumors Survive Radiation: Plastic Cells and Tolerant Niches Drive Recurrence
  • AI System Cracks Decades-Old Math Problems Using Machine-Verified Proofs
  • Black Holes Hide Supercritical Phase Boundaries in the Complex Plane, Study Finds
  • Wood Smoke Tops the List as Belgrade’s Air Pollution Fingerprint Revealed in Year-Long Study

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Science News
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,150 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading