原始文档 文章 Anthropic's Updated Responsible Scaling Policy (RSP) Summary

Anthropic's Updated Responsible Scaling Policy (RSP) Summary

文章 6 min read · 未标注

Core Purpose & Commitment

  • Risk Governance Framework: Updated RSP mitigates potential catastrophic risks from frontier AI systems.
  • Core Commitment: "We will not train or deploy models unless we have implemented safety and security measures that keep risks below acceptable levels."
  • Principle of Proportional Protection: Safeguards scale with potential risks using AI Safety Level Standards (ASL Standards), inspired by Biosafety Levels.

Key Framework Components

1. Capability Thresholds

Specific AI abilities triggering stronger safeguards:

  • Autonomous AI R&D: If a model can independently conduct complex AI research (potentially accelerating development unpredictably), requires ASL-4+ standards and additional safety assurances.
  • CBRN Weapons: If a model can meaningfully assist someone with basic technical background in creating/deploying chemical, biological, radiological, or nuclear weapons, requires ASL-3 standards.

2. ASL Standards (Graduated Safeguards)

  • ASL-1: Basic capabilities (e.g., chess-playing bots).
  • ASL-2: Current industry best practices (all current Anthropic models operate here).
  • ASL-3: Enhanced security and deployment controls:
    • Security: Internal access controls, robust model weight protection.
    • Deployment: Multi-layered approach including real-time/asynchronous monitoring, rapid response protocols, thorough pre-deployment red teaming.
  • ASL-4+: For highest-risk capabilities (e.g., autonomous AI R&D).

Implementation & Oversight

  • Capability Assessments: Routine evaluations against Capability Thresholds.
  • Safeguard Assessments: Routine evaluation of security/deployment measure effectiveness.
  • Documentation: Processes inspired by safety case methodologies from high-reliability industries.
  • Governance: Internal stress-testing, internal reporting for safety issues, and external expert feedback on methodologies.

Lessons from First Year of Implementation

  • Identified minor procedural shortcomings (e.g., evaluations completed 3 days late, placeholder evaluation clarity).
  • All instances posed minimal risk to model safety.
  • Key Lessons Learned:
    1. Need for more flexibility in policies.
    2. Need for improved compliance tracking processes.

Organizational Changes & Hiring

  • New Responsible Scaling Officer: Co-Founder/Chief Science Officer Jared Kaplan succeeds Sam McCandlish.
  • New Role: Opening position for Head of Responsible Scaling to coordinate cross-company RSP compliance.
  • Teams Contributing to RSP:
    • Frontier Red Team (threat modeling, capability assessments)
    • Trust & Safety (deployment safeguards)
    • Security and Compliance (security safeguards, risk management)
    • Alignment Science (ASL-3+ safety measures, capability evaluations, alignment stress-testing)
    • RSP Team (policy drafting, assurance, execution)

Additional Context

  • Complementary Policies: RSP complements Anthropic's Usage Policy (prohibiting misinformation, violence, fraud, etc.) and broader societal impact research.
  • External Collaboration: Assessment methodology shared with AI Safety Institutes and independent experts for feedback (not endorsement).
  • Resources: Full policy at anthropic.com/rsp, updates at anthropic.com/rsp-updates.

Forward-Looking Statement

"The frontier of AI is advancing rapidly, making it challenging to anticipate what safety measures will be appropriate for future systems. All aspects of our safety program will continue to evolve."

来源

暂无来源