原始文档 文章 Introducing Claude Sonnet 4.5

Introducing Claude Sonnet 4.5

文章 5 min read · 未标注

Source: Anthropic | Date: Sep 29, 2025

Core Announcement

Claude Sonnet 4.5 is Anthropic's latest model, positioned as the best coding model in the world and the strongest for building complex agents and computer use. It shows substantial gains in reasoning and math.

Key Product Upgrades & Features

  • Claude Code Enhancements:
    • Checkpoints: Save progress and roll back instantly (highly requested feature).
    • Refreshed terminal interface and a native VS Code extension.
  • Claude API & Agent Capabilities:
    • New context editing feature and memory tool for longer-running, more complex agents.
  • Claude Apps (Web, iOS, Android):
    • Code execution and file creation (spreadsheets, slides, documents) directly in conversations.
  • Claude for Chrome Extension: Now available to Max users from the waitlist.
  • Claude Agent SDK: The infrastructure powering Claude Code is now available for developers to build their own agents.

Performance & Benchmarks

  • Coding: State-of-the-art on SWE-bench Verified (77.2% standard, 82.0% with high compute). Can maintain focus for 30+ hours on complex tasks.
  • Computer Use: Leads on OSWorld benchmark at 61.4% (up from Sonnet 4's 42.2%).
  • Reasoning & Math: Shows improved capabilities across evaluations.
  • Domain Expertise: Experts in finance, law, medicine, and STEM report dramatically better knowledge and reasoning compared to older models, including Opus 4.1.

Customer Testimonials (Key Insights)

  • Cursor: State-of-the-art coding performance, especially on longer-horizon tasks.
  • GitHub Copilot: Significant improvements in multi-step reasoning and code comprehension for agentic experiences.
  • Canva: Handles complex, long-context tasks for 240M+ users, noticeably more intelligent.
  • Figma: Improves prompt iteration and functional prototyping in early testing.
  • Devin: Increased planning performance by 18% and end-to-end eval scores by 12%.
  • Financial Analysis: Delivers investment-grade insights for complex analysis with less human review.

Alignment & Safety

  • Most Aligned Model Yet: Shows large improvements in alignment, reducing concerning behaviors like sycophancy, deception, power-seeking, and encouraging delusions.
  • Prompt Injection Defense: Considerable progress made for agentic and computer use capabilities.
  • Safety Level: Released under AI Safety Level 3 (ASL-3) protections, including classifiers to detect CBRN-related content.
  • False Positives: Reduced by a factor of ten since initial description; users can seamlessly continue interrupted conversations with Sonnet 4.

Availability & Pricing

  • Available everywhere today.
  • API Identifier: claude-sonnet-4-5
  • Pricing: Same as Claude Sonnet 4: $3 / $15 per million tokens (input/output).
  • Recommendation: Anthropic recommends upgrading for all uses as a drop-in replacement.

Bonus: "Imagine with Claude" Research Preview

  • A temporary experiment where Claude generates software on the fly in real time.
  • Available to Max subscribers for five days at claude.ai/imagine.

Technical Methodology (Footnotes)

  • SWE-bench Verified: Results averaged over 10 trials with a 200K thinking budget. "High compute" results use parallel attempts, rejection sampling, and an internal scoring model.

来源

暂无来源