原始文档 文章 Anthropic Announcements: New Models & Computer Use

Anthropic Announcements: New Models & Computer Use

文章 5 min read · 未标注

Overview

On October 22, 2024, Anthropic announced three major developments:

  1. An upgraded Claude 3.5 Sonnet model.
  2. A new Claude 3.5 Haiku model.
  3. A groundbreaking computer use capability in public beta.

Update (12/03/2024): Claude 3.5 Haiku pricing revised to $0.80 MTok input / $4 MTok output.


Key Announcements

1. Upgraded Claude 3.5 Sonnet

  • Focus: Significant improvements, especially in coding and agentic tasks.
  • Performance Highlights:
    • SWE-bench Verified (Coding): Score improved from 33.4% to 49.0%, surpassing all publicly available models, including OpenAI o1-preview.
    • TAU-bench (Agentic Tool Use): Retail domain improved from 62.6% to 69.2%; Airline domain from 36.0% to 46.0%.
  • Cost & Speed: Same price and speed as its predecessor.
  • Availability: Immediately available to all users.
  • Safety: Jointly tested by US and UK AI Safety Institutes. Remains classified under ASL-2 Standard.
  • Early Customer Feedback:
    • GitLab: Up to 10% stronger reasoning with no added latency.
    • Cognition: Substantial improvements in coding, planning, and problem-solving.
    • The Browser Company: Outperformed every model they've tested for automating web workflows.

2. New Claude 3.5 Haiku

  • Positioning: Next generation of Anthropic's fastest model.
  • Performance: Matches or exceeds Claude 3 Opus (the previous largest model) on many benchmarks while maintaining similar speed to the previous Haiku.
  • Key Strength: Scores 40.6% on SWE-bench Verified, outperforming many agents using models like the original Claude 3.5 Sonnet and GPT-4o.
  • Use Cases: User-facing products, specialized sub-agent tasks, and generating personalized experiences from large data volumes.
  • Availability: Later in October 2024, initially as a text-only model on Anthropic API, Amazon Bedrock, and Google Cloud's Vertex AI. Image input to follow.

3. Computer Use (Public Beta)

  • Concept: A new capability allowing Claude to use computers like a human—by looking at a screen, moving a cursor, clicking, and typing.
  • Status: Experimental public beta. Described as "at times cumbersome and error-prone."
  • API: Provides a way for Claude to perceive and interact with computer interfaces, translating natural language instructions into computer commands.
  • Performance: On the OSWorld benchmark (screenshot-only category), Claude 3.5 Sonnet scored 14.9%, notably better than the next-best system's 7.8%. With more steps, it scored 22.0%.
  • Current Limitations: Actions like scrolling, dragging, and zooming are challenging. Developers are encouraged to start with low-risk tasks.
  • Safety Measures: Proactive approach with new classifiers to identify computer use and potential harm. Released early for developer feedback.
  • Early Adopters: Asana, Canva, Cognition, DoorDash, Replit, and The Browser Company are exploring it for complex, multi-step tasks.
  • Availability: Today on the Anthropic API, Amazon Bedrock, and Google Cloud's Vertex AI.

Key Excerpts & Quotes

"Claude 3.5 Sonnet is the first frontier AI model to offer computer use in public beta."

"Instead of making specific tools to help Claude complete individual tasks, we're teaching it general computer skills—allowing it to use a wide range of standard tools and software programs designed for people."


Looking Ahead

Anthropic emphasizes that computer use is in its earliest stages. Learning from initial deployments will help understand the potential and implications of increasingly capable AI systems. They welcome developer feedback on the new models and the computer use beta.

来源

暂无来源