Skip to content

OpenAI releases 372 maths results at once

A still-secret OpenAI model produced 722 mathematical manuscripts, posted online without going through journals. What they contain, how they are checked, and why mathematicians are divided.

Advertisement

Picture a mathematics lecture hall in the evening, after everyone has left. In the morning, the professors find the blackboards covered in proofs. Not one or two: hundreds. Nobody yet knows which are correct, which are new, and which quietly copy someone else's work.

That is roughly the situation OpenAI has just put mathematicians around the world in.

722 manuscripts, posted at once

This week OpenAI published, on the GitHub platform, a set of 722 mathematical manuscripts grouped into 372 families of results. They were produced by an internal model the company has not released.

The topics span algebra, number theory, theoretical computer science, logic and topology. Among the highlighted results: a proof, in certain cases, of a formula linked to the Birch and Swinnerton-Dyer conjecture, one of the seven great Millennium Problems, and a result on spin glasses, materials whose disorder fascinates physicists.

"We are entering a new era of discovery now," said Sam Altman, OpenAI's chief executive.

How the machine worked

According to OpenAI, the model was given about 4,000 problems. For most published results, a single prompt to a single agent was enough, though some needed several attempts. Each result used on average about three hours of computing, the equivalent of ChatGPT Pro's thinking mode.

In other words, nobody guided the machine step by step. It was given the statement and searched on its own: try a lead, drop it, try another, go back. Just like a researcher pulling an all-nighter at the blackboard, except it pulled thousands of them in parallel.

Three hours for a result researchers had not found is dizzying. But proportions matter: out of 4,000 problems attempted, 372 families of results were kept. The machine also failed a lot, and those failures are not in the shop window.

The automatic proofreader of maths

The real problem with 722 manuscripts is not writing them. It is reading them. Checking a single difficult proof can take a specialist weeks.

So OpenAI relies on Lean, a proof-checking software. Think of it as a spell checker, but for logic: you translate the proof into its language, and it checks, step by step, that each deduction really follows from the previous one. If a single step is off, it refuses.

Many of the published proofs went through Lean. That is a real sign of seriousness. But Lean checks that reasoning is correct, not that it is new, nor that it matters. And not every proof has been reviewed by human experts yet.

A precedent in August 📐
In August, OpenAI had already announced ten results from an internal model. Mathematicians then accused the company of presenting already-published ideas as new, without properly crediting their authors. One of them, Stephen Miller of Yeshiva University, spoke in Scientific American of research misconduct. OpenAI promised corrections. The question of fair attribution remains open for the 722 new manuscripts.

Why mathematicians are divided

Some are enthusiastic. A machine able to clear hundreds of problems could save research years, and free researchers for the deepest questions.

Others are worried. According to the specialist site The Decoder, twenty-five Fields Medal winners, the highest honour in mathematics, signed an open letter denouncing a deep disconnect between the goals of the AI industry and those of mathematics. Their fear: that mass-producing true statements ends up taking over from understanding.

The choice to publish directly on GitHub, without going through journals and their peer review, also sends a message. In essence, it says the usual pace of science is too slow for this volume of results. Journals have not changed their pace.

The takeaway

What this release shows is less a machine that has "solved maths" than a shift in balance. Producing a result is becoming fast and cheap. Checking it, placing it among what already exists, and understanding what it means, remains slow, and deeply human.

The blackboards are full. The question is who will have time to read them. That is the subject of our reflection tonight.

Advertisement