Skip to content

Sophon — AI Research Catalog of Evals, Tools, and Labs

Sophon (sophon.at) is a browsable catalog of AI evaluations, the tools and methods that improve scores on them, and the labs producing both. It indexes leaderboards (LMArena Text / Vision / Webdev / Text-to-Image, etc.), tags each benchmark with the models scoring on it, ranks tools by “known eval-lift evidence,” and groups benchmarks by capability and by how close their top scores sit to ceiling. Useful as a quick external dashboard for “what’s saturated, what isn’t, and which intervention moved which benchmark.”

  • Surfaces leaderboards across modalities — current standings include Arena - Text / Text Style Control / Vision / Vision Style Control / Webdev / Text-to-Image, all sourced from LMArena [Leaderboards section].
  • Highlights benchmarks “closest to saturation” with explicit top-score percentages — MBPP at 100.0%, MATH-500 at 99.2%, τ²-bench at 99.1%, AIME 2025 at 98.7%, AIME 2024 at 96.7%, GPQA Diamond at 94.1%, GPQA full at 93.2%, LiveBench-Math at 93.2% [Closest to saturation section].
  • Organizes capability-tagged tool rankings (“what lifts scores most”) so a reader can search for tools by capability rather than by paper title [What lifts scores most / Browse by capability sections].
  • Supports keyboard search (⌘K) over the entire eval/tool/lab catalog [Start here section].

The site is structured as four cross-linked indices: Leaderboards (live rating-system standings, currently dominated by LMArena), Recommender / “What lifts scores most” (tools ranked by accumulated eval-lift evidence), Capabilities (capability-keyed views of which tools train toward each capability), and Evals (every indexed benchmark, sortable by saturation). Individual leaderboard entries display the #1 model and its rating (e.g., Claude Opus 4.6 at 1550 on Arena Text Style Control; Gemini 3.1 Pro at 1536 on Arena Text; gpt-image-2 at 1360 on Arena Text-to-Image). Coverage on each benchmark page reports the number of models scored — GPQA Diamond at 448, τ²-bench at 319, AIME 2025 at 207, MATH-500 at 178, AIME 2024 at 172, LiveBench-Math at 51, MBPP at 15, GPQA full at 10 — implying the catalog is mostly autopopulated from leaderboard submissions rather than hand-curated. Authorship and methodology are not disclosed on the landing page.

Not a research artifact — no benchmark numbers to report. The site itself is the artifact. As a data point on the broader landscape: seven of the eight “closest to saturation” benchmarks Sophon highlights are math/coding/PhD-QA benchmarks, with the top score above 93% on all eight and at 100% on MBPP. That is consistent with the trend already filed in the wiki that frontier-LM evaluations on math and code are running out of ceiling, and that “the next benchmark” is increasingly a video / agentic / interleaved-multimodal one.

Sophon is a third-party companion to the wiki’s VLM-as-Evaluator concept — instead of evaluating one paper at a time, it lets a researcher browse which benchmarks the field has converged on and which are saturated. Complements the in-wiki benchmark-drift discussion (e.g., GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation documenting silent drift in GenEval, and ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation on sample-efficient evaluation): Sophon is the dashboard, those papers are the methodology critique. Worth bookmarking when scoping a new project — “is this benchmark still informative or already saturated?” is one click away.