原始文档 文章 Summary: Anthropic's Responsible Scaling Policy (RSP)

Summary: Anthropic's Responsible Scaling Policy (RSP)

文章 6 min read · 未标注

Overview

Anthropic has published its Responsible Scaling Policy (RSP), a framework of technical and organizational protocols to manage catastrophic risks from increasingly capable AI systems. The policy focuses on risks where AI directly causes large-scale devastation, either through deliberate misuse (e.g., bioweapons) or autonomous actions contrary to designer intent.

Core Framework: AI Safety Levels (ASL)

The RSP introduces an AI Safety Levels (ASL) system, modeled after the U.S. government's biosafety levels (BSL). It requires safety, security, and operational standards proportional to a model's potential for catastrophic risk.

ASL Definitions & Current Status

  • ASL-1: Systems posing no meaningful catastrophic risk (e.g., a 2018 LLM, a chess-playing AI).
  • ASL-2: Systems showing early signs of dangerous capabilities (e.g., ability to give bioweapon instructions) but where the information is not yet useful due to insufficient reliability or not exceeding non-AI baselines. Current LLMs, including Claude, are ASL-2.
  • ASL-3: Systems that substantially increase catastrophic misuse risk compared to non-AI baselines (e.g., search engines) OR show low-level autonomous capabilities.
  • ASL-4+: Not yet defined, but will likely involve qualitative escalations in misuse potential and autonomy.

Implementation & Key Commitments

  • ASL-2 measures represent Anthropic's current safety/security standards, overlapping with White House commitments.
  • ASL-3 measures include stricter standards requiring intense research/engineering, such as:
    • Unusually strong security requirements.
    • A commitment not to deploy ASL-3 models if they show any meaningful catastrophic misuse risk under adversarial testing by world-class red-teamers.
  • ASL-4 measures are not yet written but will be defined before reaching ASL-3. They may require unsolved research methods (e.g., interpretability to demonstrate a model won't engage in catastrophic behaviors).

Design Philosophy & Business Impact

  • The ASL system balances targeting catastrophic risk with incentivizing beneficial applications and safety progress.
  • It implicitly requires pausing training of more powerful models if AI scaling outstrips the ability to comply with safety procedures.
  • This creates a "race to the top" dynamic where competitive incentives are channeled into solving safety problems.
  • Business perspective: The RSP will not alter current uses of Claude or disrupt product availability. It is analogous to pre-market testing in automotive/aviation industries.

Governance & Iteration

  • The RSP is formally approved by Anthropic's board; changes require board approval following consultations with the Long Term Benefit Trust.
  • Procedural safeguards ensure the integrity of the evaluation process.
  • These commitments are an early iteration; rapid iteration and course correction will be necessary due to the fast pace of AI development.

Acknowledgments

Anthropic thanks ARC Evals for key insights and expertise, particularly regarding evaluations for autonomous capabilities. ARC Evals' broader Responsible Scaling Policy framework inspired Anthropic's approach.

Key Takeaways

  1. ASL Framework: A tiered system (ASL-1 to ASL-4+) that scales safety requirements with model capability.
  2. Current State: Claude is ASL-2; ASL-3 requires stricter security and red-teaming before deployment.
  3. Balanced Incentives: The policy aims to make safety a competitive advantage, not a hindrance.
  4. Governance: Board-approved with a commitment to iterative updates.
  5. Industry Influence: Designed to inspire policymakers, nonprofits, and other companies.

Footnote: Anthropic has consistently found that working with frontier AI models is an essential ingredient in developing new methods to mitigate AI risk.

来源

暂无来源