原始文档 文章 Anthropic's Responsible Scaling Policy (RSP) Version 3.0 - Summary

Anthropic's Responsible Scaling Policy (RSP) Version 3.0 - Summary

文章 5 min read · 未标注

Source: Anthropic, February 24, 2026 Purpose: A voluntary framework to mitigate catastrophic risks from AI systems, updated to reinforce successes, address shortcomings, and increase transparency and accountability.

Core Concept and Theory of Change

The RSP uses conditional ("if-then") commitments tied to AI Safety Levels (ASL). If a model exceeds certain capability thresholds (e.g., in biological science), stricter safeguards are activated.

Original Theory of Change (2023):

  • Internal Forcing Function: Treat safeguards as launch requirements.
  • Race to the Top: Encourage other companies to adopt similar policies.
  • Create Consensus: Use capability thresholds to advocate for multilateral action.
  • Future Coordination: Assume higher ASLs would require government coordination.

Assessment of Previous RSP (2.5 Years Later)

Successes:

  • Incentivized development of stronger safeguards (e.g., ASL-3 classifiers for chemical/biological weapons).
  • Implementation of ASL-3 standard proved feasible (activated May 2025).
  • Encouraged other companies (OpenAI, Google DeepMind) to adopt similar frameworks.
  • Informed early AI policy (e.g., California SB 53, EU AI Act Codes of Practice).

Shortcomings and Structural Challenges:

  • "Zone of Ambiguity": Capability thresholds are more ambiguous than anticipated. Evaluations cannot provide definitive answers, weakening the public case for risk.
  • Unilateral Limits: Higher ASLs (ASL-4+) may require mitigations impossible for one company (e.g., top-tier model weight security).
  • Political Climate: An anti-regulatory environment complicates multilateral action.

Key Updates in RSP v3.0

The policy is restructured to adopt more realistic unilateral commitments while mapping ambitious industry-wide needs.

1. Separating Company Plans from Industry Recommendations

  • Company Commitments: Mitigations Anthropic will pursue unilaterally.
  • Industry Map: An ambitious capabilities-to-mitigations map for the entire AI industry to manage advanced AI risks.

2. Frontier Safety Roadmap

A new requirement to publish concrete, ambitious yet achievable public goals across four areas: Security, Alignment, Safeguards, and Policy. These are "nonbinding but publicly-declared" targets for transparency and accountability.

Example Roadmap Goals:

  • Launch "moonshot R&D" for unprecedented information security.
  • Develop automated red-teaming surpassing collective bug bounty contributions.
  • Implement systematic measures to ensure Claude behaves according to its constitution.
  • Establish centralized, AI-analyzed records of critical development activities to detect insider/security threats.
  • Publish a policy roadmap for a "regulatory ladder" that scales with risk.

3. Risk Reports and External Review

  • Risk Reports: Published every 3-6 months, detailing model safety profiles, threat models, mitigations, and overall risk assessment. They will highlight gaps between Anthropic's measures and its broader industry recommendations.
  • External Review: For certain reports, expert third-party reviewers will have unredacted access and provide public review of Anthropic's reasoning and decisions.

Conclusion

RSP v3.0 is a living document that amplifies past successes, commits to greater transparency, and pragmatically separates what Anthropic can achieve alone from what the industry needs to do collectively. It will continue to evolve with the technology.

来源

暂无来源