OpenAI published 722 math manuscripts on GitHub from an unreleased internal model that tackled roughly 4,000 problems, averaging about three hours of ChatGPT Pro thinking compute per result. Only 162 of the 722 papers, roughly 22%, have a computer-checked main result in the Lean formal verification language, and OpenAI acknowledges that some unformalized results could have issues. Mathematicians remain skeptical: MIT’s Andrew Sutherland demands tangible proof before accepting the claim that a single AI agent, prompted with one prompt, solved these problems. The Institute for Advanced Study recommended disclosing the model name, prompts, and compute costs—information OpenAI has not yet shared.
Source: Read the original article

