Skip to content
research

OpenAI’s Secret Math Model Is a Bigger Deal Than GPT-7

A viral label is racing ahead of the evidence, while mathematicians face a more consequential question: how should AI-generated discoveries be checked? The answer could reshape research long before anyone settles on a model name.

Aki Tanaka
OpenAI’s Secret Math Model Is a Bigger Deal Than GPT-7

Forget GPT-7: What OpenAI Actually Released

OpenAI did not announce a new named model like "GPT-7." Instead, the company released a substantial collection of 722 mathematical manuscripts, organized into 372 families of related results. This release, produced by an unreleased internal model, directly addresses roughly 4,000 open research problems.

The widespread "GPT-7" speculation originates from viral guesses, not official OpenAI confirmation. OpenAI consistently refers to the system as an "unreleased internal model," emphasizing its current operational status within the company rather than as a commercial product or a named successor to existing GPT series.

Scale of this work matters because it moves beyond standard benchmarks. The release includes not just claims or striking chatbot answers, but also supporting artifacts, such as formal proofs verified with the Lean theorem prover. This allows external researchers to independently validate the mathematical claims, marking a significant shift towards verifiable, machine-assisted discovery.

This collection covers a broad spectrum of mathematics, including:

  • Number theory
  • Geometry
  • Probability
  • Theoretical computer science
  • Mathematical physics

The average compute expenditure per result was approximately three hours of ChatGPT Pro thinking time, primarily via single prompts, a notable contrast to earlier, more resource-intensive approaches involving massive agent swarms.

These Were Research Questions, Not Homework

Unlike typical homework problems with known solutions, OpenAI’s model tackled open problems: questions without established answers, even among experts. The model engaged with approximately 4,000 such problems, encompassing diverse mathematical fields.

These research questions spanned:

  • Numbers
  • Geometry
  • Probability
  • Computer science
  • Mathematical physics

OpenAI states its list of 372 result families is not a ranking of importance, reflecting the breadth of its internal model’s explorations.

Reported outcomes varied significantly. Some of the 722 manuscripts claim full solutions, while others detail partial progress on complex problems. Crucially, some papers present counterexamples, disproving long-held conjectures and demonstrating a nuanced understanding of mathematical boundaries.

One particularly noteworthy claim, Family 003, concerns the quasi-Riemann hypothesis. This establishes a zero-free region for the Riemann zeta function, specifically $\zeta(s) \neq 0$ for $\text{Re}(s) > 7/8$. It is vital to understand this is a weaker claim, not a proof of the full Riemann hypothesis, which remains unsolved. Researchers have independently reproduced and verified this specific result using the Lean theorem prover.

A Proof Is Stronger Than a Confident Answer

Mathematical proof demands a valid argument covering every possible case, not simply a pattern holding across numerous examples. Testing a rule on thousands of instances only reveals a trend; it does not prove the rule's universal truth. A single counterexample can invalidate an entire conjecture, underscoring the rigor required.

OpenAI utilized Lean, an interactive theorem prover, to formalize many proof steps. Lean can check these steps against stated rules and assumptions, offering a level of certainty beyond persuasive prose. This computer-checked artifact provides stronger evidence, as an automated kernel verifies the logical consistency of the argument.

Not all results boast computer-checkable proofs, and even verified work requires scrutiny to ensure it genuinely proves the claim in the manuscript. The Advisory Group on Mathematics and Artificial Intelligence (AGMAI) consulted on open-science standards, but its involvement does not imply verification of every result. For more details on the initiative, see Sharing AI Progress in Mathematics - OpenAI.

For instance, one notable claim, the quasi-Riemann hypothesis (Family 003), posits a zero-free region for the Riemann zeta function. This weaker version of the famous hypothesis has been independently reproduced and verified by researchers using the Lean compiler, highlighting the power of formal verification in advancing mathematical understanding.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

The Real Shift Is How AI Research Gets Tested

OpenAI reported that each successful result averaged about three hours of ChatGPT Pro thinking time from its unreleased model. This figure quantifies the computational effort for that specific internal system to generate a result, not a guarantee that any consumer-grade ChatGPT Pro subscription can solve open research problems by running for three hours. The number also excludes the extensive human and automated verification steps.

Contrast this with OpenAI’s earlier Navier–Stokes project, which involved roughly 10,000 AI agents working for 88 hours. These two figures measure different aspects of computational load across distinct problems and methodologies; they do not establish a direct speedup or a sudden leap in general model capabilities. Different tasks require different computational strategies.

Mathematics and code offer unusually verifiable testing grounds for AI reasoning because proofs can be checked algorithmically, often with tools like Lean. This allows for unprecedented rigor in evaluating AI-generated solutions. However, success in these structured domains does not automatically prove general intelligence or reliable performance across less formal fields like ethics, social science, or creative writing.

The productive question for the scientific community remains whether independent mathematicians can reproduce, contextualize, and improve these results. This emphasis on verifiable, open science stands in stark contrast to the speculative "GPT-7" headlines that often dominate the news cycle, shifting the focus from hype to tangible, testable progress.

Frequently Asked Questions

Did OpenAI announce a model called GPT-7?

No. The release describes an unreleased internal model; GPT-7 is speculation, not an official name.

What did the model produce?

OpenAI shared 722 mathematical manuscripts grouped into 372 result families, based on work across roughly 4,000 research problems.

Did the model solve the Riemann hypothesis?

No. A manuscript claims a weaker, quasi-Riemann result; the full Riemann hypothesis remains unsolved.

Are all the mathematical results verified?

No. Some proofs have computer-checkable Lean artifacts, while others still require expert review and may contain errors.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$199 · AI tools & software only

For builders

This page is doing a job for someone else’s tool.

AI agents read it. Buyers land on it. It answers in eight languages and over MCP. Your tool can have one like it — live in 24 hours.