原始文档 ›文章 ›Claude Sonnet 4.6: Summary
Release Date: February 17, 2026
Key Highlights
- Most capable Sonnet model yet, with full upgrades in coding, computer use, long-context reasoning, agent planning, knowledge work, and design.
- 1M token context window available in beta.
- Default model for Free and Pro plans on claude.ai and Claude Cowork.
- Pricing: Same as Sonnet 4.5: $3/$15 per million tokens (input/output).
- Performance: Approaches or matches Opus-class intelligence at a lower cost, preferred by early users over both Sonnet 4.5 and Opus 4.5.
Core Capabilities & Improvements
Computer Use
- Major advancement in using software like a human (mouse/keyboard) without APIs.
- Benchmark Progress: Steady gains on OSWorld over 16 months. Early users report human-level capability in tasks like navigating complex spreadsheets and multi-step web forms.
- Safety: Significant improvement in resistance to prompt injection attacks compared to Sonnet 4.5, performing similarly to Opus 4.6.
Coding & Agentic Work
- User Preference: In Claude Code, users preferred Sonnet 4.6 over Sonnet 4.5 ~70% of the time and over Opus 4.5 59% of the time.
- Key Strengths: Better context reading, logic consolidation, instruction following, fewer hallucinations, and consistent multi-step follow-through.
- Use Cases: Excels at complex code fixes, large codebase searches, bug detection, and deep codebase work previously requiring Opus.
Long-Context & Reasoning
- 1M Token Context: Effectively reasons over entire codebases, lengthy contracts, or dozens of research papers.
- Strategic Planning: Demonstrated superior long-horizon planning in the Vending-Bench Arena evaluation, winning by investing in capacity early and pivoting to profitability.
- Benchmark Performance: Matches Opus 4.6 on OfficeQA (enterprise document comprehension) and shows significant improvements on reasoning benchmarks.
Design & Knowledge Work
- Visual Outputs: Customers report notably more polished frontend code, layouts, animations, and design sensibility.
- Document Analysis: Significant jump in answer match rate for financial services workflows. Outperformed Sonnet 4.5 by 15 percentage points in heavy reasoning Q&A on enterprise documents (Box evaluation).
- Insurance Benchmark: Achieved 94% accuracy, the highest-performing model tested for computer use in mission-critical workflows.
Safety & Character
- Extensive safety evaluations show it is as safe as, or safer than, other recent Claude models.
- Characterized as having a "broadly warm, honest, prosocial, and at times funny character, very strong safety behaviors, and no signs of major concerns around high-stakes forms of misalignment."
Product & Platform Updates
- Thinking Modes: Supports both adaptive thinking and extended thinking.
- Context Compaction (Beta): Automatically summarizes older context as conversations approach limits, increasing effective context length.
- API Tools: Web search and fetch tools now automatically write and execute code to filter results, improving quality and token efficiency. Code execution, memory, programmatic tool calling, and tool use examples are generally available.
- Claude in Excel: Add-in now supports MCP connectors (e.g., S&P Global, LSEG, FactSet) to pull external context into spreadsheets. Available on Pro, Max, Team, and Enterprise plans.
How to Use
- Available now on all Claude plans, Claude Cowork, Claude Code, the API, and major cloud platforms.
- Free tier upgraded to Sonnet 4.6 by default, including file creation, connectors, skills, and compaction.
- API Identifier:
claude-sonnet-4-6
Recommendation
- Sonnet 4.6 is the recommended model for most tasks, offering frontier-level performance at a lower cost.
- Opus 4.6 remains the strongest for tasks demanding the deepest reasoning, such as complex codebase refactoring, multi-agent coordination, and problems where precision is paramount.
来源
暂无来源