Skip to content

Cursor: agent swarm discovers novel solution to Problem Six of the First Proof challenge

Cursor CEO Michael Truell announces that Cursor’s agent swarm — the same long-running system that previously “built a full browser from scratch unsupervised” per follow-on reporting — solved Problem Six (a spectral-graph-theory problem) of the First Proof challenge, a 2026-02 math research benchmark created by a group of Fields Medalists and MacArthur Fellows to test “research-grade” rather than “competition-style” mathematical reasoning. The claim, per Truell, is that the agent produced a solution yielding stronger results than the official human-written solution, after four days of zero human nudging. Truell frames this as evidence that Cursor’s agent-coordination technique “might generalise beyond coding.” No paper, write-up, or proof artifact is linked from the tweet; this page captures the product claim as it circulated.

  • Cursor’s agent swarm produced a novel solution to Problem Six of the First Proof challenge that the team believes yields stronger results than the human-written reference solution [tweet body].
  • The First Proof challenge is a research-math benchmark released 2026-02, designed by prominent mathematicians (including Fields Medalists and MacArthur Fellows) explicitly to move past pattern-matchable competition-style problems [external context from etn.’s summary thread, not in the original tweet].
  • The agent run was ~4 days with zero human nudging, on a spectral-graph-theory problem; same agent infrastructure previously reported to have “built a full browser from scratch unsupervised” [external context from the AI Investor’s commentary thread, not in the original tweet].
  • Truell frames the result as evidence that the agent-coordination technique may generalize beyond coding [tweet body + follow-on commentary].

Not disclosed. The tweet is a product/result announcement. No paper, proof artifact, transcript, or technical write-up is linked. From follow-on commentary (which is third-party, not from Cursor) the setup is characterized as a long-running agent swarm operating without human intervention over multiple days on a single math problem; no architecture, training recipe, tool stack, or compute budget is given.

  • One headline claim: novel solution to one problem (Problem Six, spectral graph theory) on one benchmark (First Proof challenge), reported to be stronger than the human reference solution.
  • No quantitative axes (no benchmark score, no comparison against other models on the same benchmark, no ablation, no inter-run variance).
  • No third-party verification of the math is referenced in the tweet.

A real datapoint for the “general-purpose agent” framing that Introducing GPT-5.3-Codex articulates at the closed-frontier — that a coding agent is “becoming a general-purpose agent that can reason, build, and execute” — but pushed harder: this is the same agent stack moving from SWE-Bench-style coding to research-grade math without (per Truell) a new training run. That makes it the cleanest filed instance so far of an Agentic Software Engineering system claiming out-of-distribution transfer to a non-coding research benchmark, complementing FARS: Fully Automated Research System (FARS) and OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis which target the research-agent stack natively. It also fits the multi-day, multi-agent Inference-Time Scaling regime — the same direction as Kimi K2.5’s PARL (Reasoning RL) and MiroThinker’s interaction-depth scaling (MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling) — but at far longer wall-clock (4 days vs minutes-to-hours) on a single problem with no online human feedback. Worth tracking, but with the standard caveats for tweet-as-primary claims about math: no proof artifact, no third-party verification, no methodological detail. Swayam’s accompanying note (“maybe its worth it to just lock up a codex 5.3 with a h100 node and see if it can do some research by itself”) captures the obvious follow-up — whether the result reproduces with a different agent stack at a fraction of Cursor’s compute.