原始文档 文章 Summary: Claude Opus 4.7 Release

Summary: Claude Opus 4.7 Release

文章 7 min read · 未标注

Overview

  • Model: Claude Opus 4.7, generally available as of April 16, 2026.
  • Positioning: A significant upgrade to Opus 4.6, particularly for advanced software engineering and complex tasks. It is less broadly capable than the unreleased Claude Mythos Preview but outperforms Opus 4.6 across benchmarks.
  • Key Theme: Enhanced autonomy, reliability, and precision for long-running, difficult tasks.

Key Improvements

  • Software Engineering: Major gains on the most difficult coding tasks. Users report confidence in handing off complex work with less supervision.
  • Capabilities:
    • Vision: Substantially better; can process images up to 2,576 pixels on the long edge (~3.75 megapixels), over 3x prior models.
    • Creativity & Taste: Produces higher-quality professional outputs (interfaces, slides, documents).
    • Instruction Following: Substantially improved, taking instructions more literally. Prompts from older models may need re-tuning.
    • Memory: Better at using file system-based memory across long, multi-session work.
  • Cybersecurity: Cyber capabilities are intentionally less advanced than Mythos Preview. Includes new safeguards to detect and block prohibited/high-risk cybersecurity uses. A Cyber Verification Program is available for security professionals.

Availability & Pricing

  • Access: Available on all Claude products, API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.
  • API Model Name: claude-opus-4-7
  • Pricing: Same as Opus 4.6: $5 per million input tokens and $25 per million output tokens.

Testing Highlights & Quotes

Early-access testers report significant improvements:

  • Coding & Autonomy:

    "It catches its own logical faults during the planning phase and accelerates execution, far beyond previous Claude models." (FinTech platform) "It's the first model to pass our implicit-need tests, and it keeps executing through tool failures that used to stop Opus cold." (Notion) "Claude Opus 4.7 autonomously built a complete Rust text-to-speech engine from scratch... Months of senior engineering, delivered autonomously." (Cognition)

  • Performance Gains:

    • CursorBench: 70% vs. Opus 4.6's 58%.
    • Rakuten-SWE-Bench: Resolves 3x more production tasks than Opus 4.6.
    • Internal Coding Benchmark: 13% higher resolution than Opus 4.6.
    • Finance Agent Evaluation: State-of-the-art score (0.813 vs. 0.767 for Opus 4.6).
  • Reliability & Quality:

    "It correctly reports when data is missing instead of providing plausible-but-incorrect fallbacks." (Hex) "It's a more intelligent, more efficient Opus 4.6: low-effort Opus 4.7 is roughly equivalent to medium-effort Opus 4.6." (Hex) "It's the cleanest jump we've seen since the move from Sonnet 3.7 to the Claude 4 series." (Cognition)

Safety & Alignment

  • Profile: Similar to Opus 4.6, with low rates of concerning behavior (deception, sycophancy).
  • Improvements: Better honesty and resistance to prompt injection attacks.
  • Weaknesses: Modestly weaker in some areas (e.g., overly detailed harm-reduction advice on controlled substances).
  • Assessment: "Largely well-aligned and trustworthy, though not fully ideal in its behavior." Mythos Preview remains the best-aligned model.

New Features & Launches

  • Effort Control: New xhigh effort level between high and max for finer control over reasoning/latency tradeoffs. Default in Claude Code is now xhigh.
  • Task Budgets (API Beta): Allows developers to guide Claude's token spend for long-running tasks.
  • Claude Code Updates:
    • New /ultrareview command for dedicated code review sessions.
    • Auto Mode extended to Max users, allowing Claude to make decisions on your behalf for longer, less interrupted tasks.

Migration from Opus 4.6

  • Tokenizer Update: Same input may map to 1.0–1.35x more tokens.
  • Increased Output: Higher effort levels, especially in agentic settings, produce more output tokens for improved reliability.
  • Recommendation: Measure impact on real traffic. Use effort parameters, task budgets, or concise prompting to control usage. A migration guide is available.

Footnotes & Methodology

  • Benchmarks compared against best available API versions of GPT-5.4 and Gemini 3.1 Pro.
  • Various evaluation harnesses and methodologies were used (e.g., Terminus-2 for Terminal-Bench, internal implementations for SWE-bench Multimodal).
  • Scores for some benchmarks (e.g., MCP-Atlas, CyberGym) were updated from initial reports.

来源

暂无来源