Google's Cogentic AI Agents Crack Five Open Math Problems

Seven researchers at Google Research submitted a paper to arXiv on September 30, introducing a multi-agent system called Cogentic built to find research-level mathematical proofs. Running on Gemini as its base model, the system advanced five open problems across online learning, auction theory, and mechanism design, according to the paper. Every proof was independently verified by domain experts afterward and expanded into full papers co-authored with those experts.

The author list includes Yang Cai, who also holds an appointment at Yale University, and Vineet Gupta, who is also affiliated with Google DeepMind, along with Aranyak Mehta, Christopher Liaw, Di Wang, and others. The paper is classified under cs.AI and cs.GT (game theory).

One Generation Pass Isn't Enough, So It Became a Pipeline

The paper's starting premise is simple: language models can already produce decent mathematical ideas, but problems that require testing several conjectures at once and grinding away for days overwhelm a single generation pass.

Cogentic splits the search for a proof across different roles:

  • An orchestrator holds the global state and decides how many provers to assign to which direction;
  • Provers write candidate proofs in parallel;
  • Verifiers hunt for errors from complementary angles, in an adversarial role;
  • A literature retriever supplies background material;
  • A ledger records only intermediate conclusions that have passed verification, carrying them across rounds so later proofs can cite them directly;
  • An additional "advisor" role watches the whole process for patterns and adjusts parameters on the fly.

Prove, verify, prove again — the cycle repeats. The ledger addresses a familiar failure mode in long-horizon tasks: a lemma the model worked out in one round often gets forgotten or restated differently in the next. Now, once it clears verification, it's locked in place.

How Far Each of the Five Problems Got

Per the paper's results:

  1. Online inverse linear optimization: the first efficient O(d) regret bound, independent of the time horizon T, with O(d²) computation per round;
  2. Competitive complexity of two-sided markets: proof that adding just 2 more sellers to the smaller side lets trading revenue match the optimal allocation;
  3. Anytime regret for n experts: an anytime algorithm whose constant term matches that of the fixed-horizon version;
  4. Simple mechanisms versus optimal revenue: the approximation ratio for a single additive buyer's revenue improved from 5.2 to 3.52;
  5. Price of anarchy for automated bidding: 1.5 — the optimum — for 2 bidders, and 2 − 1/(4n+1) for n bidders.

On cost, the paper says most problems took on the order of a hundred calls to Gemini, and the hardest took on the order of a thousand; it doesn't specify which version of Gemini was used.

Set Beside OpenAI's Proof Efforts

AI-for-math news has been thick on the ground in the second half of this year, and set side by side, the difference is in how verification is done.

In August, OpenAI's Astra produced 10 results in mathematics and theoretical computer science, backed by a 249-page paper and Lean formalization certificates. The Navier–Stokes proof it announced in September ran roughly ten thousand agents for 88 hours, also with an accompanying Lean file. Machine-checkable formal certificates are the selling point OpenAI keeps emphasizing on that front.

Cogentic's paper puts its weight elsewhere: inside the system, adversarial verifiers screen the work once; after the system is done, people who know the field read it line by line, then co-write it into a formal paper with experts. The scope is narrower too — all five problems come from the authors' own research areas, problems where insiders are already watching and peers can judge the results.

That choice makes the results easier to vouch for, at the cost of generalizability. The paper itself admits the problems were "selected from areas familiar to the authors," and doesn't address how the approach would fare on directions the authors don't know well.

Readers Can't Keep Up With the Output

One line in the paper echoes almost exactly what Terence Tao's August essay worried about:

"A system like this can produce candidate results faster than they can be read."

Cogentic's five problems each had an expert checking the work, so the paper can call them "verified." But once this kind of orchestration opens up to more people and more problems, the bottleneck shifts to the reviewers. The paper doesn't say how long expert verification took per proof — and that's exactly the number that would tell you how far this can scale.

Sources: arXiv paper 2609.40324, CocoLoop, Google Research; all bounds, approximation ratios, and call-count orders of magnitude for the five results follow the paper's own text.