Meta published six math papers on its research blog on October 2, all produced through collaboration between mathematicians and the thinking modes of Muse Spark 1.1 and 1.2. Five of the six resolve questions that had been publicly posed before but remained unanswered.
The blog post is restrained in how it describes the model's role. In its own words:
Proof strategies were developed and revised with help from Muse Spark
What each of the six papers solved
The six papers span six different fields: probability theory, differential equations, group theory, optimization, arithmetic physics, and non-associative algebra:
| Field | Paper | What it did |
|---|---|---|
| Probability theory | Sharp threshold for Gaussian ellipsoid fitting | Established the sharp threshold at which high-dimensional Gaussian points can be fit by an ellipsoid |
| Differential equations | Finite-time blow-up of radial negative-energy solutions | Proved that solutions to the mass-critical biharmonic nonlinear Schrödinger equation blow up in finite time |
| Group theory | Semi-abelian groups need not be monomial groups | Used a group of order 384 to disprove a conjecture M. Kida proposed in 2024 |
| Optimization | Ring-based relaxations are tight | Answered a question Del Pia and Khajavirad posed earlier this year |
| Arithmetic physics | String two-point functions equal height functions on curves | Connected p-adic string theory to number theory, extending an idea Manin proposed in the 1980s |
| Non-associative algebra | Solvable evolution algebras | Used a three-dimensional example to disprove a classification conjecture by García-Martínez and Pérez-Rodríguez |
The credited mathematicians include Aykut Arslan (two papers), Leonard Dinh, Joseph Phillip Brennan and Milana Golich, and Andres Barei. The arithmetic physics paper was completed by a team of five led by Anindya Dey.
How the process was set
Meta laid out four practices: a team of mathematicians sets the research direction; a separate group of mathematicians reviews independently; each paragraph of a paper is labeled as drafted by the researcher or by the AI; and all prior work is credited in full.
That last practice applied directly to the probability paper. The Gaussian-ellipsoid-fitting threshold problem was independently solved by three teams at almost the same time this past August, and Meta's paper acknowledges the concurrent work. That suggests the problem was already within reach of human mathematicians, and the model's contribution looked more like acceleration than opening a path no one else could see.
The blog post also spelled out its intent: "Our goal here wasn't to mass-produce papers, but to empower researchers."
Set against Google's paper from the same week
The same week, Google Research published Cogentic, which uses multiple agents working together to explore mathematical proofs. Placed side by side, the two approaches diverge clearly.
Google's problems cluster around online learning, auction theory, and mechanism design. Most of its five problems called Gemini roughly a hundred times each, with the hardest one calling it about a thousand times; the approximation ratio for single additive buyer revenue was brought down from 5.2 to 3.52, and every proof was checked by a domain expert. The emphasis there is on the system — how multiple agents divide labor and check one another.
Meta's problems, by contrast, are scattered across six unrelated branches. It didn't disclose call counts or describe any agent architecture; the emphasis there is on people — who led, who reviewed, who drafted which paragraph. The two narratives map to two different ways of verifying the work: Google's lets you estimate cost from call volume, while Meta's lets you trace accountability through the byline.
Both materials are missing the same thing: failure cases. Neither company has said how many problems the model worked on but failed to crack. Without that denominator, the five successful papers can only tell us so much, and outsiders have no real way to gauge Muse Spark's actual hit rate in mathematics.
Sources: Meta AI Research blog, Google Research blog, CocoLoop; titles, authors, counterexample scale, and model versions for all six papers were verified against Meta's blog post.