Overview
- Model: Claude Opus 4.7, generally available as of April 16, 2026.
- Positioning: A significant upgrade to Opus 4.6, particularly for advanced software engineering and complex tasks. It is less broadly capable than the unreleased Claude Mythos Preview but outperforms Opus 4.6 across benchmarks.
- Key Theme: Enhanced autonomy, reliability, and precision for long-running, difficult tasks.
Key Improvements
- Software Engineering: Major gains on the most difficult coding tasks. Users report confidence in handing off complex work with less supervision.
- Capabilities:
- Vision: Substantially better; can process images up to 2,576 pixels on the long edge (~3.75 megapixels), over 3x prior models.
- Creativity & Taste: Produces higher-quality professional outputs (interfaces, slides, documents).
- Instruction Following: Substantially improved, taking instructions more literally. Prompts from older models may need re-tuning.
- Memory: Better at using file system-based memory across long, multi-session work.
- Cybersecurity: Cyber capabilities are intentionally less advanced than Mythos Preview. Includes new safeguards to detect and block prohibited/high-risk cybersecurity uses. A Cyber Verification Program is available for security professionals.
Availability & Pricing
- Access: Available on all Claude products, API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.
- API Model Name:
claude-opus-4-7 - Pricing: Same as Opus 4.6: $5 per million input tokens and $25 per million output tokens.
Testing Highlights & Quotes
Early-access testers report significant improvements:
-
Coding & Autonomy:
"It catches its own logical faults during the planning phase and accelerates execution, far beyond previous Claude models." (FinTech platform) "It's the first model to pass our implicit-need tests, and it keeps executing through tool failures that used to stop Opus cold." (Notion) "Claude Opus 4.7 autonomously built a complete Rust text-to-speech engine from scratch... Months of senior engineering, delivered autonomously." (Cognition)
-
Performance Gains:
- CursorBench: 70% vs. Opus 4.6's 58%.
- Rakuten-SWE-Bench: Resolves 3x more production tasks than Opus 4.6.
- Internal Coding Benchmark: 13% higher resolution than Opus 4.6.
- Finance Agent Evaluation: State-of-the-art score (0.813 vs. 0.767 for Opus 4.6).
-
Reliability & Quality:
"It correctly reports when data is missing instead of providing plausible-but-incorrect fallbacks." (Hex) "It's a more intelligent, more efficient Opus 4.6: low-effort Opus 4.7 is roughly equivalent to medium-effort Opus 4.6." (Hex) "It's the cleanest jump we've seen since the move from Sonnet 3.7 to the Claude 4 series." (Cognition)
Safety & Alignment
- Profile: Similar to Opus 4.6, with low rates of concerning behavior (deception, sycophancy).
- Improvements: Better honesty and resistance to prompt injection attacks.
- Weaknesses: Modestly weaker in some areas (e.g., overly detailed harm-reduction advice on controlled substances).
- Assessment: "Largely well-aligned and trustworthy, though not fully ideal in its behavior." Mythos Preview remains the best-aligned model.
New Features & Launches
- Effort Control: New
xhigheffort level betweenhighandmaxfor finer control over reasoning/latency tradeoffs. Default in Claude Code is nowxhigh. - Task Budgets (API Beta): Allows developers to guide Claude's token spend for long-running tasks.
- Claude Code Updates:
- New
/ultrareviewcommand for dedicated code review sessions. - Auto Mode extended to Max users, allowing Claude to make decisions on your behalf for longer, less interrupted tasks.
- New
Migration from Opus 4.6
- Tokenizer Update: Same input may map to 1.0–1.35x more tokens.
- Increased Output: Higher effort levels, especially in agentic settings, produce more output tokens for improved reliability.
- Recommendation: Measure impact on real traffic. Use effort parameters, task budgets, or concise prompting to control usage. A migration guide is available.
Footnotes & Methodology
- Benchmarks compared against best available API versions of GPT-5.4 and Gemini 3.1 Pro.
- Various evaluation harnesses and methodologies were used (e.g., Terminus-2 for Terminal-Bench, internal implementations for SWE-bench Multimodal).
- Scores for some benchmarks (e.g., MCP-Atlas, CyberGym) were updated from initial reports.
来源
暂无来源