原始文档 ›文章 ›Progress from Anthropic's Frontier Red Team
Source: Anthropic Blog, March 19, 2025 Core Finding: Frontier AI models show "early warning" signs of rapid progress in dual-use capabilities, approaching or exceeding undergraduate-level skills in cybersecurity and expert-level knowledge in some biology areas. However, they currently fall short of thresholds for substantially elevated national security risks.
Key Capability Progress
Cybersecurity
- Rapid Advancement: 2024 was a "zero to one" moment. Claude improved from a high schooler to an undergraduate level in Capture The Flag (CTF) exercises in one year.
- Benchmark Performance: On the Cybench public benchmark, Claude 3.7 Sonnet solves ~33% of challenges within five attempts, up from ~5% for the previous frontier model a year prior.
- Areas of Improvement: Progress is seen across CTF categories:
pwn,web, andcrypto. - Current Limitations: Skills still lag behind expert humans, particularly in reverse engineering binaries and network reconnaissance/exploitation.
- Realistic Testing: In large (~50 host) cyber range experiments with Carnegie Mellon University, models cannot autonomously succeed. However, when equipped with a researcher-built toolset (e.g., Incalmo), they could replicate a complex, multi-stage attack similar to a known large-scale data theft.
Biosecurity
- Knowledge Surge: Within a year, Claude went from underperforming to comfortably exceeding world-class virology experts on a lab troubleshooting evaluation (VCT).
- Uneven Capabilities: Models are approaching/exceeding human expert baselines on tasks like understanding biology protocols and cloning workflows, but remain worse at interpreting scientific figures.
- Weaponization Risk Assessment:
- In controlled studies, models provided some uplift to novices compared to those without AI access.
- However, the highest-scoring AI-assisted plans still contained critical mistakes leading to real-world failure.
- Expert red-teaming concluded models cannot reliably guide a novice through end-to-end bio-weapon acquisition due to too many critical failures in planning.
- Mitigation Investment: Due to rapid improvement, Anthropic is investing heavily in monitoring and mitigations (e.g., constitutional classifiers).
Strategic Partnerships & Oversight
- Government Collaboration: Pre-deployment testing conducted with the US AI Safety Institute and UK AI Security Institute (AISI). Insights informed the AI Safety Level (ASL) determination for Claude 3.7 Sonnet.
- Nuclear Domain Pilot: Anthropic partnered with the National Nuclear Security Administration (NNSA) for classified red-teaming on nuclear/radiological risk. This demonstrates the feasibility of public-private collaboration in highly sensitive domains.
Looking Ahead & Actionable Insights
- Evaluation Scaling: Goal is to scale up to more frequent, automated evaluations for earlier risk detection.
- Imminent Threshold: Based on biology research, models are getting closer to crossing the capabilities threshold requiring AI Safety Level 3 safeguards, prompting additional investment in security measures.
- Tool Obsolescence: As models improve with extended thinking, current cyber toolkits like Incalmo may become obsolete, with models performing better "out of the box."
- Core Recommendation: Deeper collaboration between frontier AI labs and governments is essential for improving evaluations and risk mitigations.
Key Quote: "Our assessment is that AI models are displaying 'early warning' signs of rapid progress in key dual-use capabilities... However, present-day models fall short of thresholds at which we consider them to generate substantially elevated risks to national security."
来源
暂无来源