原始文档 文章 Progress from Anthropic's Frontier Red Team

Progress from Anthropic's Frontier Red Team

文章 6 min read · 未标注

Source: Anthropic Blog, March 19, 2025 Core Finding: Frontier AI models show "early warning" signs of rapid progress in dual-use capabilities, approaching or exceeding undergraduate-level skills in cybersecurity and expert-level knowledge in some biology areas. However, they currently fall short of thresholds for substantially elevated national security risks.

Key Capability Progress

Cybersecurity

  • Rapid Advancement: 2024 was a "zero to one" moment. Claude improved from a high schooler to an undergraduate level in Capture The Flag (CTF) exercises in one year.
  • Benchmark Performance: On the Cybench public benchmark, Claude 3.7 Sonnet solves ~33% of challenges within five attempts, up from ~5% for the previous frontier model a year prior.
  • Areas of Improvement: Progress is seen across CTF categories: pwn, web, and crypto.
  • Current Limitations: Skills still lag behind expert humans, particularly in reverse engineering binaries and network reconnaissance/exploitation.
  • Realistic Testing: In large (~50 host) cyber range experiments with Carnegie Mellon University, models cannot autonomously succeed. However, when equipped with a researcher-built toolset (e.g., Incalmo), they could replicate a complex, multi-stage attack similar to a known large-scale data theft.

Biosecurity

  • Knowledge Surge: Within a year, Claude went from underperforming to comfortably exceeding world-class virology experts on a lab troubleshooting evaluation (VCT).
  • Uneven Capabilities: Models are approaching/exceeding human expert baselines on tasks like understanding biology protocols and cloning workflows, but remain worse at interpreting scientific figures.
  • Weaponization Risk Assessment:
    • In controlled studies, models provided some uplift to novices compared to those without AI access.
    • However, the highest-scoring AI-assisted plans still contained critical mistakes leading to real-world failure.
    • Expert red-teaming concluded models cannot reliably guide a novice through end-to-end bio-weapon acquisition due to too many critical failures in planning.
  • Mitigation Investment: Due to rapid improvement, Anthropic is investing heavily in monitoring and mitigations (e.g., constitutional classifiers).

Strategic Partnerships & Oversight

  • Government Collaboration: Pre-deployment testing conducted with the US AI Safety Institute and UK AI Security Institute (AISI). Insights informed the AI Safety Level (ASL) determination for Claude 3.7 Sonnet.
  • Nuclear Domain Pilot: Anthropic partnered with the National Nuclear Security Administration (NNSA) for classified red-teaming on nuclear/radiological risk. This demonstrates the feasibility of public-private collaboration in highly sensitive domains.

Looking Ahead & Actionable Insights

  • Evaluation Scaling: Goal is to scale up to more frequent, automated evaluations for earlier risk detection.
  • Imminent Threshold: Based on biology research, models are getting closer to crossing the capabilities threshold requiring AI Safety Level 3 safeguards, prompting additional investment in security measures.
  • Tool Obsolescence: As models improve with extended thinking, current cyber toolkits like Incalmo may become obsolete, with models performing better "out of the box."
  • Core Recommendation: Deeper collaboration between frontier AI labs and governments is essential for improving evaluations and risk mitigations.

Key Quote: "Our assessment is that AI models are displaying 'early warning' signs of rapid progress in key dual-use capabilities... However, present-day models fall short of thresholds at which we consider them to generate substantially elevated risks to national security."

来源

暂无来源