原始文档 文章 Summary: Claude 3.5 Sonnet Launch

Summary: Claude 3.5 Sonnet Launch

文章 4 min read · 未标注

Key Announcement

  • Model: Claude 3.5 Sonnet is the first release in the Claude 3.5 model family.
  • Performance: Outperforms competitor models and Claude 3 Opus on a wide range of evaluations, with the speed and cost of the mid-tier Claude 3 Sonnet.
  • Availability: Now available for free on Claude.ai and the Claude iOS app. Higher rate limits for Claude Pro and Team subscribers. Also available via the Anthropic API, Amazon Bedrock, and Google Cloud's Vertex AI.
  • Pricing: $3 per million input tokens and $15 per million output tokens, with a 200K token context window.

Performance & Capabilities

  • Intelligence Benchmarks: Sets new industry standards for graduate-level reasoning (GPQA), undergraduate-level knowledge (MMLU), and coding proficiency (HumanEval).
  • Speed: Operates at twice the speed of Claude 3 Opus.
  • Coding Proficiency: In an internal agentic coding evaluation, Claude 3.5 Sonnet solved 64% of problems, outperforming Claude 3 Opus (38%). It can independently write, edit, and execute code, and handles code translations effectively for legacy application updates.
  • Vision: The strongest vision model yet, surpassing Claude 3 Opus on standard benchmarks. Excels at visual reasoning (e.g., interpreting charts) and accurately transcribing text from imperfect images.

New Features & Future Roadmap

  • Artifacts: A new feature on Claude.ai that creates a dedicated window for AI-generated content (code, documents, designs), enabling real-time editing and integration into workflows. This marks Claude's evolution into a collaborative work environment.
  • Upcoming Models: Claude 3.5 Haiku and Claude 3.5 Opus will be released later this year to complete the 3.5 family.
  • Future Features: Development of new modalities, enterprise integrations, and Memory (to remember user preferences and interaction history).

Safety & Privacy

  • Safety Level: Remains at ASL-2 despite the intelligence leap.
  • External Testing: Provided to the UK's Artificial Intelligence Safety Institute (UK AISI) for pre-deployment evaluation, with results shared with the US AISI.
  • Expert Feedback: Integrated policy feedback from external experts, including child safety experts at Thorn, to update classifiers and fine-tune models.
  • Privacy Principle: Does not train on user-submitted data without explicit permission.

Actionable Information

  • Access: Use for free on Claude.ai or the iOS app. API access is available through major cloud platforms.
  • Cost-Effective: Ideal for complex, speed-sensitive tasks like customer support and multi-step workflows due to its performance-to-cost ratio.
  • Provide Feedback: Users can submit feedback directly in-product to influence the development roadmap.

来源

暂无来源