Sophon — AI Research Catalog of Evals, Tools, and Labs
Sophon (sophon.at) is a browsable catalog of AI evaluations, the tools and methods that improve scores on them, and the labs producing both. It indexes leaderboards (LMArena Text / Vision / Webdev / Text-to-Image, etc.), tags each benchmark with the models scoring on it, ranks tools by “known eval-lift evidence,” and groups benchmarks by capability and by how close their top scores sit to ceiling. Useful as a quick external dashboard for “what’s saturated, what isn’t, and which intervention moved which benchmark.”
Key claims
Section titled “Key claims”- Surfaces leaderboards across modalities — current standings include Arena - Text / Text Style Control / Vision / Vision Style Control / Webdev / Text-to-Image, all sourced from LMArena [Leaderboards section].
- Highlights benchmarks “closest to saturation” with explicit top-score percentages — MBPP at 100.0%, MATH-500 at 99.2%, τ²-bench at 99.1%, AIME 2025 at 98.7%, AIME 2024 at 96.7%, GPQA Diamond at 94.1%, GPQA full at 93.2%, LiveBench-Math at 93.2% [Closest to saturation section].
- Organizes capability-tagged tool rankings (“what lifts scores most”) so a reader can search for tools by capability rather than by paper title [What lifts scores most / Browse by capability sections].
- Supports keyboard search (
⌘K) over the entire eval/tool/lab catalog [Start here section].
Method
Section titled “Method”The site is structured as four cross-linked indices: Leaderboards (live rating-system standings, currently dominated by LMArena), Recommender / “What lifts scores most” (tools ranked by accumulated eval-lift evidence), Capabilities (capability-keyed views of which tools train toward each capability), and Evals (every indexed benchmark, sortable by saturation). Individual leaderboard entries display the #1 model and its rating (e.g., Claude Opus 4.6 at 1550 on Arena Text Style Control; Gemini 3.1 Pro at 1536 on Arena Text; gpt-image-2 at 1360 on Arena Text-to-Image). Coverage on each benchmark page reports the number of models scored — GPQA Diamond at 448, τ²-bench at 319, AIME 2025 at 207, MATH-500 at 178, AIME 2024 at 172, LiveBench-Math at 51, MBPP at 15, GPQA full at 10 — implying the catalog is mostly autopopulated from leaderboard submissions rather than hand-curated. Authorship and methodology are not disclosed on the landing page.
Results
Section titled “Results”Not a research artifact — no benchmark numbers to report. The site itself is the artifact. As a data point on the broader landscape: seven of the eight “closest to saturation” benchmarks Sophon highlights are math/coding/PhD-QA benchmarks, with the top score above 93% on all eight and at 100% on MBPP. That is consistent with the trend already filed in the wiki that frontier-LM evaluations on math and code are running out of ceiling, and that “the next benchmark” is increasingly a video / agentic / interleaved-multimodal one.
Why it’s interesting
Section titled “Why it’s interesting”Sophon is a third-party companion to the wiki’s VLM-as-Evaluator concept — instead of evaluating one paper at a time, it lets a researcher browse which benchmarks the field has converged on and which are saturated. Complements the in-wiki benchmark-drift discussion (e.g., GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation documenting silent drift in GenEval, and ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation on sample-efficient evaluation): Sophon is the dashboard, those papers are the methodology critique. Worth bookmarking when scoping a new project — “is this benchmark still informative or already saturated?” is one click away.
See also
Section titled “See also”- VLM-as-Evaluator — Sophon is essentially an external dashboard view of the same evaluator-aggregation problem this concept tracks.
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation — argues benchmarks silently drift as models improve; Sophon’s saturation view operationalizes that observation.
- ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation — methodology for efficient evaluation and failure discovery, complementary to Sophon’s catalog approach.
- awesome-diffusion-papers.vercel.app — time-sorted, searchable browser for awesome-diffusion-categorized — similar third-party catalog tool, but for diffusion papers rather than evaluations.