原始文档 ›文章 ›Introducing the next generation of Claude
Overview
Anthropic announced the Claude 3 model family on March 4, 2024, setting new industry benchmarks. The family includes three models in ascending order of capability: Haiku, Sonnet, and Opus. They are designed to offer optimal balances of intelligence, speed, and cost.
Availability: Opus and Sonnet are available immediately via the Claude API (now generally available in 159 countries) and on claude.ai. Haiku will be available soon.
Model Details & Specifications
Claude 3 Opus
- Role: Most intelligent model, for highly complex tasks.
- Key Differentiator: "Higher intelligence than any other model available."
- Cost: $15 / $75 per million tokens (Input / Output).
- Context Window: 200K tokens (1M available for specific use cases).
- Ideal For: Task automation, R&D, strategy, advanced analysis.
Claude 3 Sonnet
- Role: Ideal balance of intelligence and speed for enterprise workloads.
- Key Differentiator: "More affordable than other models with similar intelligence; better for scale."
- Cost: $3 / $15 per million tokens (Input / Output).
- Context Window: 200K tokens.
- Ideal For: Data processing (RAG), sales, time-saving tasks like code generation.
Claude 3 Haiku
- Role: Fastest, most compact model for near-instant responsiveness.
- Key Differentiator: "Smarter, faster, and more affordable than other models in its intelligence category."
- Cost: $0.25 / $1.25 per million tokens (Input / Output).
- Context Window: 200K tokens.
- Ideal For: Live customer interactions, content moderation, cost-saving tasks.
Key Features & Improvements
Performance & Intelligence
- Benchmark Leadership: Opus outperforms peers on common evaluations like MMLU (undergraduate knowledge), GPQA (graduate reasoning), and GSM8K (mathematics).
- Speed:
- Haiku: Can read a 10k-token arXiv paper in under three seconds.
- Sonnet: 2x faster than Claude 2/2.1 for most workloads.
- Opus: Similar speed to Claude 2/2.1 but with much higher intelligence.
- Vision: Sophisticated capabilities to process photos, charts, graphs, and technical diagrams.
- Long Context & Recall: 200K context window at launch (1M possible). Opus achieved >99% accuracy on the "Needle In A Haystack" recall test.
Usability & Accuracy
- Fewer Refusals: Models show more nuanced understanding, refusing harmless prompts less often.
- Improved Accuracy: Opus shows a twofold improvement in accuracy on complex factual questions vs. Claude 2.1, with fewer hallucinations. Citations feature coming soon.
- Easier to Use: Better at following complex, multi-step instructions and producing structured output (e.g., JSON).
Safety & Responsibility
- Responsible Design: Dedicated teams track risks (misinformation, CSAM, biological misuse, etc.). Uses Constitutional AI methods.
- Reduced Bias: Shows less bias than previous models on the Bias Benchmark for Question Answering (BBQ).
- Safety Level: Remains at AI Safety Level 2 (ASL-2). Red teaming concludes negligible potential for catastrophic risk at this time.
Availability & Future
- Current Access: Opus and Sonnet are available via the Claude API and claude.ai (Sonnet powers free tier, Opus for Pro subscribers).
- Cloud Partners: Sonnet available on Amazon Bedrock and in private preview on Google Cloud Vertex AI. Opus and Haiku coming soon.
- Future Updates: Frequent updates planned, including features like Tool Use (function calling), interactive coding (REPL), and advanced agentic capabilities.
Anthropic's Stance: "We do not believe that model intelligence is anywhere near its limits... Our hypothesis is that being at the frontier of AI development is the most effective way to steer its trajectory towards positive societal outcomes."
来源
暂无来源