Introducing Claude Opus 4.7
Claude Opus 4.7 is Anthropic’s GA frontier closed-source model, positioned as a direct upgrade over Opus 4.6 with the headline gains in advanced software engineering: longer-horizon autonomy, stricter instruction-following, and self-verification before reporting back. The release introduces an xhigh effort level between high and max for finer inference-compute control, raises image input resolution to 2,576px long edge (~3.75 MP, >3× prior Claude), and reports state-of-the-art on Finance Agent and GDPval-AA among partner-supplied numbers. A less-capable sibling of an undisclosed “Mythos Preview” model held back under Project Glasswing cyber-safeguards; Opus 4.7 is the first model to ship with automatic cyber-misuse detection and a Cyber Verification Program for legitimate security users.
Key claims
Section titled “Key claims”- Opus 4.7 introduces a new
xhigheffort level betweenhighandmax, presented as a finer-grained inference-compute knob; Claude Code defaults toxhighfor all plans at launch [§Also launching today]. - Maximum image input is now 2,576 pixels on the long edge (~3.75 MP), more than 3× the prior Claude limit, applied as a model-level change rather than an API parameter [§Testing Claude Opus 4.7, fn 1].
- Opus 4.7 is described as the first model trained with explicit differential reduction of cyber capabilities relative to a more capable internal model (“Mythos Preview”), and the first to ship with automatic detection/blocking of prohibited or high-risk cybersecurity requests; security professionals can apply to a Cyber Verification Program for vulnerability research, penetration testing, and red-teaming [§Introducing Claude Opus 4.7].
- On a 93-task internal coding benchmark cited by Anthropic from a partner (Vercel), Opus 4.7 lifts resolution by 13% over Opus 4.6, including four tasks neither Opus 4.6 nor Sonnet 4.6 could solve [§Testing Claude Opus 4.7].
- Partner-reported coding gains: CursorBench 70% vs Opus 4.6’s 58% (Cursor); Rakuten-SWE-Bench 3× more production tasks resolved vs Opus 4.6 (Rakuten); BigLaw Bench 90.9% substantive accuracy at high effort (Harvey); Hex internal finance-analyst eval reports low-effort Opus 4.7 ≈ medium-effort Opus 4.6 [§Testing Claude Opus 4.7].
- XBOW’s visual-acuity benchmark for autonomous penetration testing jumps from 54.5% (Opus 4.6) to 98.5% (Opus 4.7), attributed to the higher input image resolution [§Testing Claude Opus 4.7].
- Opus 4.7 is claimed state-of-the-art on the third-party GDPval-AA economically-valuable knowledge-work evaluation, and on Anthropic’s internal Finance Agent evaluation [§Testing Claude Opus 4.7].
- Pricing unchanged from Opus 4.6: 25/M output tokens. API id
claude-opus-4-7. Available on Claude products, API, Amazon Bedrock, Vertex AI, Microsoft Foundry [§Introducing Claude Opus 4.7]. - An updated tokenizer increases input-token count by roughly 1.0–1.35× depending on content type; output tokens also increase at higher effort levels, particularly in later agentic turns [§Migrating from Opus 4.6 to Opus 4.7].
- Safety: similar overall profile to Opus 4.6 with improvements in honesty and prompt-injection resistance and a modest regression on overly-detailed controlled-substance harm-reduction advice; alignment assessment summarized as “largely well-aligned and trustworthy, though not fully ideal” — Mythos Preview is internally claimed to be the best-aligned Anthropic model trained to date [§Safety and alignment].
- Claude Code adds a
/ultrareviewslash command (dedicated review session for change sets) and extends Auto Mode permissions to Max users; API adds task budgets in public beta [§Also launching today].
Method
Section titled “Method”This is a closed-source model release announcement, not a technical paper — no architecture, training data, or post-training recipe is disclosed beyond a few framing points. What is described:
- A continuation of the Opus 4.x line, with the differentiator framed as longer-horizon autonomy and self-verification rather than a step change in architecture.
- An updated tokenizer, with a 1.0–1.35× input-token expansion factor depending on content type — an unusually concrete disclosure for a closed release, motivated by migration planning.
- A new inference-time effort tier (
xhigh) inserted betweenhighandmax. The post does not say whetherxhighis a separately-tuned reasoning budget or a runtime hyperparameter, only that it “thinks more” thanhighand that Anthropic recommends starting athighorxhighfor coding and agentic use. - A novel deployment-side safety primitive: automatic per-request detection of prohibited cyber uses, applied at the model layer with a verification path (Cyber Verification Program) for legitimate security work. Framed as a precondition to eventually releasing the more capable Mythos-class model.
- Roughly two dozen partner testimonials with internal benchmark numbers (CursorBench, Rakuten-SWE-Bench, BigLaw Bench, XBOW visual-acuity, Hex finance, Databricks OfficeQA Pro, CodeRabbit precision/recall, Notion Agent tool-error rate, Factory Droids task success, etc.), spanning coding (Cursor, Replit, Vercel, Bolt, Warp, Devin, Codeium-style platforms), code review (CodeRabbit, Qodo), agent platforms (Notion, Hebbia, Genspark, Factory, Ramp), domain work (Harvey/legal, Hex/finance, Solve Intelligence/life sciences patents, Databricks/document reasoning), and security (XBOW).
Results
Section titled “Results”Headline partner-supplied numbers (all comparisons vs Opus 4.6 unless noted, no shared external benchmark across partners):
- Vercel 93-task internal coding bench: +13% task resolution, 4 tasks newly solved that neither Opus 4.6 nor Sonnet 4.6 could.
- Cursor CursorBench: 70% vs 58%.
- Rakuten-SWE-Bench: 3× more production tasks resolved, double-digit lifts on Code Quality and Test Quality.
- Harvey BigLaw Bench: 90.9% substantive accuracy at high effort.
- XBOW visual-acuity benchmark for autonomous pen-testing: 98.5% vs 54.5% (attributed to the new ~3.75 MP image input limit).
- Hex finance-analyst eval: low-effort 4.7 ≈ medium-effort 4.6.
- Notion Agent multi-step workflows: +14% over Opus 4.6, ~⅓ the tool errors.
- Databricks OfficeQA Pro: -21% errors when reasoning over source documents.
- CodeRabbit code review: +10% recall with stable precision; faster than GPT-5.4 xhigh on their harness.
- Anthropic internal research-agent benchmark: tied for top across six modules at 0.715; General Finance 0.813 vs 0.767.
- Claimed SOTA on third-party GDPval-AA (knowledge work across finance, legal, other domains) and on Anthropic’s internal Finance Agent benchmark.
Cost/perf notes: pricing flat at 25 per M input/output tokens; total token consumption increases via (a) new tokenizer (1.0–1.35× input) and (b) more output thinking at higher effort levels.
Safety: alignment assessment per the system card concludes “largely well-aligned and trustworthy, though not fully ideal.” Improvements in honesty and prompt-injection resistance; modest regression in controlled-substance harm-reduction verbosity. Cyber capability is the first frontier where Anthropic ships a deployed-time safeguard (automatic detection + verification program) in addition to training-time differential reduction.
Why it’s interesting
Section titled “Why it’s interesting”The closed-source partner-benchmark wedge widens further. Opus 4.7 follows the same release shape as Introducing GPT-5.3-Codex — a closed-frontier model whose benchmark story is told mostly through partner testimonials and a few internal evals rather than a system card with disclosed methodology — which makes it useful as the next closed-source reference point on the Agentic Software Engineering Pareto, even though direct numerical comparison to the open MoE cluster (Qwen3-Coder-Next Technical Report, Step 3.5 Flash) is impossible without a shared benchmark axis. The XBOW visual-acuity jump from 54.5 → 98.5% on the same axis is the cleanest single datapoint that input resolution, not just reasoning, is now a load-bearing axis for agentic computer-use — relevant to the OSWorld / computer-use gap flagged on the Agentic Software Engineering page. The xhigh effort level is a small but concrete addition to the Inference-Time Scaling taxonomy: a closed-source frontier model now exposes four effort tiers as a user-facing knob, formalizing the “spend more tokens to think more” axis as a product surface rather than just a research artifact. And the cyber-safeguard story — differential capability reduction during training plus runtime misuse detection plus a verification program — is a new deployment-time safety pattern that didn’t appear in earlier filed releases.
See also
Section titled “See also”- Agentic Software Engineering — closed-frontier reference point alongside GPT-5.3-Codex; partner-benchmark testimony pattern for coding-agent capability
- Inference-Time Scaling —
xhigheffort tier formalizes spend-more-think-more as a user-facing knob - Introducing GPT-5.3-Codex — sibling closed-frontier release with the same partner-testimonial benchmark format
- GPT-5.4 Thinking and GPT-5.4 Pro rolling out in ChatGPT (product announcement) — GPT-5.4 launch tweet, similarly minimal disclosure; tracker for the next closed-frontier datapoint
- Step 3.5 Flash — open-MoE counterpart on the agentic-SWE Pareto; explicit Claude-Code compatibility
- Qwen3-Coder-Next Technical Report — open agentic-SWE recipe; comparison anchor for what the closed-source release implicitly omits