原始文档 文章 Activating AI Safety Level 3 Protections

Activating AI Safety Level 3 Protections

文章 5 min read · 未标注

Anthropic has proactively activated its AI Safety Level 3 (ASL-3) Deployment and Security Standards for the launch of Claude Opus 4. This is a precautionary measure, as the company has not definitively determined whether the model's capabilities require ASL-3, but cannot rule it out. The ASL-3 standards involve enhanced security to prevent model weight theft and targeted deployment measures to limit misuse for chemical, biological, radiological, and nuclear (CBRN) weapons development.

Key Rationale & Background

  • Proactive & Precautionary: Anthropic is implementing ASL-3 before a final capability assessment is complete. This allows them to develop and test protections in advance.

    "We have not yet determined whether Claude Opus 4 has definitively passed the Capabilities Threshold that requires ASL-3 protections. Rather, due to continued improvements in CBRN-related knowledge and capabilities, we have determined that clearly ruling out ASL-3 risks is not possible for Claude Opus 4..."

  • Responsible Scaling Policy (RSP): The RSP framework requires stronger protections as models approach capability thresholds. All previous models operated under the baseline ASL-2 standard.

  • Capability Thresholds: Dangerous capability evaluations are challenging. Proactively enabling a higher standard simplifies releases and allows for iterative learning.

ASL-3 Deployment Measures

Focused narrowly on preventing assistance with end-to-end CBRN weapon workflows, not general queries or single pieces of information.

Three-Part Defense Approach:

  1. Making the system harder to jailbreak: Implemented Constitutional Classifiers—real-time guards trained on synthetic data to block harmful CBRN information, adding moderate compute overhead.
  2. Detecting jailbreaks: Includes a bug bounty program, offline classification, and threat intelligence partnerships.
  3. Iteratively improving defenses: Rapidly remediate discovered jailbreaks using synthetic data to retrain classifiers.
  • Scope: Initially focused exclusively on biological weapons as the primary risk.
  • False Positives: Measures may occasionally affect legitimate queries. Vetted users with dual-use applications can receive targeted exemptions.

ASL-3 Security Measures

Focused on protecting model weights from theft by sophisticated non-state actors. Involves over 100 security controls.

Key Unique Control:

  • Egress Bandwidth Controls: Restricts data flow out of secure environments. Leverages the large size of model weights to detect and block exfiltration attempts via unusual bandwidth usage. This is a forcing function for understanding data flow.

Other Controls Include:

  • Two-party authorization for weight access

  • Enhanced change management protocols

  • Endpoint software controls via binary allowlisting

  • Scope: Targets sophisticated non-state actors. Nation-state threats (using novel attack chains) and sophisticated insider risk are out of scope for ASL-3.

Conclusions & Next Steps

  • The approach is iterative. If Claude Opus 4 is found not to have surpassed the ASL-3 capability threshold, protections may be adjusted or removed.
  • Anthropic will continue to improve measures, collaborate with the industry, government, and civil society, and publish detailed reports to aid others.
  • The practical experience of operating under ASL-3 will help uncover new issues and opportunities.

来源

暂无来源