Skip to content

ResearchClawBench: Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery

ResearchClawBench is an agent-evaluation benchmark from InternScience for automated research, spanning a “Re-Discovery to New-Discovery” spectrum. Scores are normalized so that 50 = matches the original paper and 100 = surpasses it; a leaderboard tracks the best score per task across all agents. As of filing the leaderboard is empty (“No scored runs yet”), so this is a benchmark-in-launch rather than a set of measured results, but its scoring semantics directly operationalize the “AI-for-AI research” evaluation gap flagged elsewhere in the wiki.

  • The benchmark evaluates AI agents on automated research spanning re-discovery of prior results through to new discoveries [landing page].
  • Scoring is normalized so that 50 = matches the original paper’s result and 100 = surpasses it [landing page frontier section].
  • The leaderboard tracks the best score per task across all agents; at filing time no scored runs are posted [landing page leaderboard section].

The public landing page describes the framing (Re-Discovery → New-Discovery evaluation) and the normalized 50/100 scoring anchor against the original paper’s result, but does not publish task lists, agent-harness specifications, or judging methodology at filing time. The leaderboard shell is present but empty, suggesting the benchmark is at launch/preview stage.

No scored runs yet [landing page]. There are no numeric agent scores, per-task breakdowns, or comparisons to prior benchmarks to report from the current landing page.

The “50 = matches paper, 100 = surpasses” scoring anchor is the concrete missing piece in the AI-for-AI Research cluster: filed AI-scientist systems (AlphaGo Moment for Model Architecture Discovery, FARS: Fully Automated Research System, Kosmos: An AI Scientist for Autonomous Discovery) each report their own headline scaling axis (SOTA-hits-per-compute, papers-per-day, expert-months-per-cycle), which makes cross-system comparison hard — a normalized per-task score against a known ground-truth paper is what would let those systems be ranked head-to-head. Complements Agent² RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training? (agent² RL-Bench), which evaluates agents on a different rung — interactive RL post-training — with its own per-task scoring; ResearchClawBench is aiming at the paper-reproduction/paper-extension rung above that. The empty leaderboard is worth flagging: this is a launch pointer, not a result yet.