Fields Medalists Warn AI Benchmark Race Is Misaligning Mathematics With the Values of the Discipline

Image: Universiteitleiden
Main Takeaway
Twenty-five Fields Medalists have warned that AI companies’ race to solve mathematical problems as benchmarks threatens research norms around understanding, credit, disclosure, and accountability.
Jump to Key PointsSummary
Why mathematicians issued the declaration
Twenty-five Fields Medalists have signed a declaration arguing that AI companies’ race to solve mathematical problems as benchmarks is damaging mathematics and widening a conflict over what the discipline is for. The statement emerged from discussions during the week before its release and was published with an invitation for additional signatures.
The signatories’ concern follows rapid advances in large language models and related systems, which have begun tackling research-level problems across several mathematical fields. Their objection centers on the incentives surrounding those demonstrations: benchmark performance rewards an answer, while mathematicians value understanding, explanatory insight, reliable attribution, and contributions that can be checked and built upon.
The benchmark dispute
The immediate trigger was OpenAI’s announcement that an AI system had solved the longstanding unit-distance problem associated with Paul Erdős, a result that drew both excitement and alarm. The problem asks about the minimum number of distinct distances determined by points in the plane, and its history reflects the kind of open-ended inquiry that mathematicians distinguish from answering a standardized test.
The declaration argues that turning such problems into public AI benchmarks changes their role. A difficult theorem becomes a score in a competition among companies, while the surrounding reasoning, novelty, and human contribution receive less attention. Scientific American described the response as an effort by mathematicians, computer scientists, and historians of mathematics to establish guardrails before AI systems become dominant in research practice.
What counts as mathematical progress
The debate reaches beyond whether AI can produce correct proofs. Mathematics develops through explanations that reveal structures, connect areas, expose useful methods, and allow other researchers to verify and extend the work. An output that reaches a valid conclusion without a comprehensible route raises questions about whether it advances knowledge in the same way as a human-readable proof.
A related essay by Terence Tao frames the issue around the purposes and values of mathematical research rather than the timing of AI capability improvements. It treats research mathematics as a practice concerned with understanding shapes, numbers, and natural phenomena, and asks how those goals should guide the use of systems that perform research-level tasks. The declaration also places the dispute within broader conflicts between AI development incentives and the values of scientific and creative professions.
Credit, errors, and disclosure
The proposed response focuses on transparency. Researchers using AI should disclose that use, identify the systems involved, and make clear which parts of a result were generated, checked, or developed by people. These details matter when a proof is difficult to understand, when a model draws on prior work without clear attribution, or when responsibility for an error is disputed.
The Leiden Declaration describes AI-generated proofs as raising practical questions about authorship, accountability, originality, and plagiarism. It calls for a community response to symbolic and neural systems that generate or formalize mathematics, while recognizing that technology has repeatedly changed mathematical practice. The International Mathematical Union has endorsed that declaration, according to the declaration’s publication page, giving the discussion an institutional dimension beyond individual researchers.
A wider split over AI’s role
The declarations do not reject every use of AI in mathematics. Their focus is the direction set by commercial incentives and the decision to treat difficult mathematical problems as competitive demonstrations. Tools that help formalize arguments, search literature, test conjectures, or check calculations occupy a different position from systems presented as autonomous researchers, although the boundary will be contested as capabilities improve.
That distinction explains the urgency in the Fields Medalists’ statement and the more consultative tone of the Leiden initiative. One group released its declaration quickly because it judged the situation pressing; the other presents a broader community process and lists institutional endorsement. Together, they show a field trying to establish norms while the technology is still changing faster than professional rules.
What happens next
The next phase will center on adoption rather than rhetoric. Mathematical journals, conferences, universities, and research funders will need policies for AI disclosure, proof verification, authorship, attribution, and the evaluation of machine-assisted results. Formal verification can help establish whether a proof is valid, but it doesn't by itself settle questions about explanation, originality, or intellectual credit.
The declarations also put pressure on AI companies to define success more carefully. A benchmark result can demonstrate capability, yet it can't represent the full value of mathematical research. The growing list of signatures, the International Mathematical Union’s endorsement of the Leiden Declaration, and continuing attention from major news outlets indicate that AI in mathematics has moved from a technical curiosity to a governance dispute inside a core scientific discipline.
Key Points
Twenty-five Fields Medalists warned that AI math benchmarks conflict with mathematics’ values of understanding and accountability.
OpenAI’s reported solution to the Erdős unit-distance problem intensified debate over AI-generated mathematical discoveries.
The Leiden Declaration calls for disclosure, attribution, accountability, and scrutiny of AI-assisted mathematical research.
Terence Tao’s essay centers mathematics’ goals and values rather than predictions about AI capability timelines.
Formal verification can test proof validity while leaving originality, explanation, and intellectual credit unresolved.
Questions Answered
The 25 Fields Medalists criticized AI companies because benchmark contests reward solutions while mathematics also values understanding, explanation, originality, and accountability. Their declaration says commercial incentives are severely misaligned with the goals of mathematical research.
OpenAI announced that an AI system solved the famous Erdős unit-distance problem. The announcement prompted excitement about AI’s mathematical ability and concern that major research problems were becoming competitive company benchmarks.
The Leiden Declaration calls for greater transparency about AI use in mathematical research. It addresses disclosure, authorship, attribution, responsibility for errors, and the evaluation of AI-generated or AI-formalized proofs.
Terence Tao examines the goals and values of mathematical research while assuming that systems capable of research-level tasks will arrive. His approach focuses on what mathematics should preserve, rather than predicting when AI capabilities will emerge.
Formal verification can establish whether a proof satisfies a formal logical standard, but it doesn't settle questions about explanation, originality, attribution, or intellectual credit. Those issues require research and publishing norms alongside technical checks.
Source Reliability
57% of sources are highly trusted · Avg reliability: 73
Go deeper with Organic Intel
Simple AI systems for your life, work, and business. Each one includes copyable prompts, guides, and downloadable resources.
Explore Systems