Overview
Anthropic has released Claude Opus 4.6, its most advanced model to date, featuring significant improvements in coding, reasoning, and long-context performance. It is available immediately on claude.ai, the API, and major cloud platforms.
Key Improvements & Features
- Enhanced Coding & Agentic Skills: Improved planning, longer sustained agentic tasks, better performance in large codebases, and superior code review/debugging.
- 1M Token Context Window (Beta): First Opus-class model with a 1M token context window, available on the Claude Developer Platform.
- State-of-the-Art Performance: Leads on multiple benchmarks:
- Terminal-Bench 2.0: Highest score in agentic coding.
- Humanity's Last Exam: Leads all frontier models in complex multidisciplinary reasoning.
- GDPval-AA: Outperforms GPT-5.2 by ~144 Elo points and its predecessor (Opus 4.5) by 190 points on economically valuable knowledge work.
- BrowseComp: Best performance in locating hard-to-find online information.
- Everyday Work Capabilities: Excels at financial analyses, research, and creating documents, spreadsheets, and presentations. Within Cowork, it can multitask autonomously.
- Safety Profile: Shows an overall safety profile as good as or better than any other frontier model, with low rates of misaligned behavior and the lowest rate of over-refusals of any recent Claude model.
Early Access Partner Feedback
Partners highlighted its autonomous operation, ability to handle complex multi-step tasks, and superior reasoning:
"Claude Opus 4.6 is the strongest model Anthropic has shipped. It takes complicated requests and actually follows through, breaking them into concrete steps, executing, and producing polished work even when the task is ambitious." – Notion
"Claude Opus 4.6 is a huge leap for agentic planning. It breaks complex tasks into independent subtasks, runs tools and subagents in parallel, and identifies blockers with real precision." – Devin
"Claude Opus 4.6 handled a multi-million-line codebase migration like a senior engineer. It planned up front, adapted its strategy as it learned, and finished in half the time." – Vercel
Technical & API Updates
- Adaptive Thinking: Model can now decide when deeper reasoning is needed, replacing the previous binary extended thinking toggle.
- Effort Controls: Four levels (
low,medium,high[default],max) to control intelligence, speed, and cost. - Context Compaction (Beta): Automatically summarizes older context to enable longer-running tasks.
- 128k Output Tokens: Supports larger outputs without breaking tasks into multiple requests.
- US-Only Inference: Available at 1.1× token pricing for workloads requiring US-based processing.
- Pricing: Remains at $5/$25 per million input/output tokens. Premium pricing applies for prompts exceeding 200k tokens ($10/$37.50 per million tokens).
Product Updates
- Claude Code: Now supports agent teams (research preview) to work on tasks in parallel.
- Claude in Excel: Substantial upgrades for handling long-running tasks, planning, and ingesting unstructured data.
- Claude in PowerPoint (Research Preview): New capability to create presentations from data or descriptions, respecting brand layouts.
Safety & Evaluations
- Comprehensive Testing: The most extensive safety evaluations to date, including new tests for user wellbeing and complex refusal scenarios.
- Cybersecurity: Enhanced abilities led to six new cybersecurity probes to track misuse. Anthropic is accelerating cyberdefensive uses, like finding and patching open-source vulnerabilities.
- Misaligned Behavior: Low rates of deception, sycophancy, and cooperation with misuse, matching the safety profile of its predecessor.
Benchmark Highlights (from system card)
- Long-Context Retrieval: On the 8-needle 1M variant of MRCR v2, Opus 4.6 scores 76% vs. Sonnet 4.5's 18.5%.
- SWE-bench Verified: Averaged 81.42% over 25 trials.
- BigLaw Bench: Achieved 90.2%, with 40% perfect scores.
- CyberGym & OpenRCA: Demonstrated superior performance in cybersecurity investigations and root cause analysis.
Footnotes & Methodology
- The 1M token context is beta and available only on the Claude Developer Platform.
- Evaluations like GDPval-AA were run independently by Artificial Analysis.
- For benchmarks like Humanity's Last Exam, models were run with tools (web search, code execution, etc.) and max reasoning effort.
- A minor score update for HLE with tools (53.1% to 53.0%) was made after improved cheating detection.
来源
暂无来源