Anthropic has proactively activated its AI Safety Level 3 (ASL-3) Deployment and Security Standards for the launch of Claude Opus 4. This is a precautionary measure, as the company has not definitively determined whether the model's capabilities require ASL-3, but cannot rule it out. The ASL-3 standards involve enhanced security to prevent model weight theft and targeted deployment measures to limit misuse for chemical, biological, radiological, and nuclear (CBRN) weapons development.
Key Rationale & Background
-
Proactive & Precautionary: Anthropic is implementing ASL-3 before a final capability assessment is complete. This allows them to develop and test protections in advance.
"We have not yet determined whether Claude Opus 4 has definitively passed the Capabilities Threshold that requires ASL-3 protections. Rather, due to continued improvements in CBRN-related knowledge and capabilities, we have determined that clearly ruling out ASL-3 risks is not possible for Claude Opus 4..."
-
Responsible Scaling Policy (RSP): The RSP framework requires stronger protections as models approach capability thresholds. All previous models operated under the baseline ASL-2 standard.
-
Capability Thresholds: Dangerous capability evaluations are challenging. Proactively enabling a higher standard simplifies releases and allows for iterative learning.
ASL-3 Deployment Measures
Focused narrowly on preventing assistance with end-to-end CBRN weapon workflows, not general queries or single pieces of information.
Three-Part Defense Approach:
- Making the system harder to jailbreak: Implemented Constitutional Classifiers—real-time guards trained on synthetic data to block harmful CBRN information, adding moderate compute overhead.
- Detecting jailbreaks: Includes a bug bounty program, offline classification, and threat intelligence partnerships.
- Iteratively improving defenses: Rapidly remediate discovered jailbreaks using synthetic data to retrain classifiers.
- Scope: Initially focused exclusively on biological weapons as the primary risk.
- False Positives: Measures may occasionally affect legitimate queries. Vetted users with dual-use applications can receive targeted exemptions.
ASL-3 Security Measures
Focused on protecting model weights from theft by sophisticated non-state actors. Involves over 100 security controls.
Key Unique Control:
- Egress Bandwidth Controls: Restricts data flow out of secure environments. Leverages the large size of model weights to detect and block exfiltration attempts via unusual bandwidth usage. This is a forcing function for understanding data flow.
Other Controls Include:
-
Two-party authorization for weight access
-
Enhanced change management protocols
-
Endpoint software controls via binary allowlisting
-
Scope: Targets sophisticated non-state actors. Nation-state threats (using novel attack chains) and sophisticated insider risk are out of scope for ASL-3.
Conclusions & Next Steps
- The approach is iterative. If Claude Opus 4 is found not to have surpassed the ASL-3 capability threshold, protections may be adjusted or removed.
- Anthropic will continue to improve measures, collaborate with the industry, government, and civil society, and publish detailed reports to aid others.
- The practical experience of operating under ASL-3 will help uncover new issues and opportunities.
来源
暂无来源