Mathematicians Want AI Labs to Fund the Understanding of Their Proofs
A community document asks AI labs to publish prompts, cost and failures with each proof, and to pay for humans to learn what the machines found.
On September 29, 2026, a group of mathematicians published recommendations for how AI labs should release mathematical results, drawing on more than 600 replies from the mathematical community. One line stands out: the authors ask frontier labs to stop testing advanced mathematical problems on proprietary models that the wider scientific community cannot access. They acknowledge the labs are doing it anyway, and wrote the rest of the document with that in mind.
The Signal: Understanding Has Become the Scarce Resource
The document is a symptom of a shift. For centuries, mathematics ran on a simple norm: the authors of a paper understand the argument, have checked it themselves, and answer for it. The recommendations open by observing that AI can now produce mathematical arguments that the person who prompted it cannot understand, verify, or take responsibility for.
That breaks the old supply chain. Generating a result used to be the hard part, and understanding came along for free. Now generation may be cheap and comprehension expensive. The document's survey was prompted by a specific episode: OpenAI announced the existence of many results without giving details. Read the text as a bid to set terms before that pattern becomes routine, and as a preview of what other fields may ask of labs once models start producing work their users cannot audit.
How the Proposed Release Process Works
The recommendations split results into two kinds. If a mathematician fully understands a paper and takes responsibility, the traditional path applies: post a preprint, submit to a journal, give talks. The harder case is output that nobody yet understands, where the authors say the lab itself should do the cleanup work rather than leaving it to mathematicians afterward.
- Attribution: search the literature and cite the papers that first introduced related ideas, even if the model rediscovered them independently.
- Readable writeups: produce proofs in conventional paper style, with friendly introductions and precise statements, not wordy reasoning in non-standard terminology.
- Neutral repositories: deposit results somewhere not controlled by any lab, with persistent citable identifiers, recorded modifications, and ideally comments.
- Provenance: publish the model name, prompts, a summarized chain of thought, time taken, and estimated compute cost.
- Formalization: where possible, provide a machine-checkable version of the proof. Formalization means translating an argument into code that a proof-checking program can verify step by step. Where that causes unacceptable delay, state the status plainly, for example formalized modulo standard accepted results.
- Denominators: document how AI came to be used on each problem. For batch releases, say how many comparable problems the models tried and failed, and how problems were chosen.
The last item matters most. A list of solved problems without the list of failures is a selection effect, and the authors also urge labs not to use result releases as marketing vehicles.
Who Pays, and Who Decides
The second half of the document is about money. Labs that release output without immediate human understanding must ensure understanding follows, including through significant funding. The suggested uses are conferences or summer schools, workshops and working groups, funded postdocs and students, and expert-written books or long expository articles.
The guardrails are deliberate. Funding decisions should go through existing nonprofit institutions with established processes, and the development of understanding should stay community-led, not directed by the labs, even for their own results. The authors add that accepting this support would not confer legitimacy on lab practices. They also warn that proprietary internal models risk a two-tier system, and that unequal access to public models could worsen existing inequalities.
Questions You Should Be Asking
- When a vendor announces AI-generated results, can anyone outside the company reproduce them, and has the vendor published the prompts, cost, and chain of thought?
- How many comparable problems did the system attempt and fail? Without that denominator, what does the success rate actually mean?
- Who inside our organization is accountable for an AI-produced analysis or proof that nobody on the team can verify?
- Is there a machine-checkable verification step, and if not, is that stated clearly or buried?
- If AI output outruns human comprehension, who funds the work of catching up, and who controls how the money is spent?
What To Watch Next
The signal is whether any lab publicly adopts even part of this: a failure denominator alongside a batch release, or funding routed through an independent nonprofit. The document is a set of requests, not a binding rule, so adoption would be voluntary. If labs ignore it, expect the mathematical community's next move to be about access and credit instead of etiquette.
- 1Request that AI labs release mathematical proofs with detailed explanations and open access to enable independent verification by the broader research community.
- 2Establish collaborative frameworks between AI labs and mathematicians to validate AI-generated proofs before publication to maintain scientific rigor standards.
- 3Advocate for your institution to require transparency reports from AI labs on their mathematical testing methodologies and results accessibility policies.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
