OpenAI’s math solutions aren’t meeting the field’s standards yet
OpenAI released hundreds of claimed solutions to difficult math problems this week, but according to reports, the outputs fell short of standards set by an elite mathematics advisory group.
OpenAI released hundreds of claimed solutions to difficult math problems this week, but according to reports, the outputs fell short of standards set by an elite mathematics advisory group. The Advisory Group on Mathematics and Artificial Intelligence, hosted by Princeton University’s Institute for Advanced Studies, issued guidelines in late September asking labs to stop testing advanced problems on proprietary models and to ensure human understanding of results.
OpenAI evaluated its proprietary models using open research problems. While the lab followed some guidelines by releasing results quickly and sharing how models reached conclusions, it missed others. Only ten of the 719 manuscripts included the model's chain of thought, and 58% of the proofs were formalized. A new paper from University of Cambridge and King's College London researchers highlighted gaps between natural language explanations and formally expressed code solutions in Lean, a programming language used to confirm proof accuracy.
The paper found two discrepancies in OpenAI's solution to a fluid behavior problem. Mathematicians emphasize that human researchers typically take responsibility for results through talks and peer review, while model prompts often leave behind solutions without adequate human comprehension.
WireUnWired turns the supplied report into a clearer brief, preserves the original publisher and author details, and adds relevant context without hiding where the information came from.
Relevant WireUnWired coverage is connected so one story can lead into the larger technology context.
Original publication: 8 October 2026 23:40
