原始文档 ›文章 ›Frontier Model Security - Anthropic Summary
Overview
Anthropic emphasizes that securing frontier AI models is a critical priority due to their rapidly increasing capabilities and strategic importance. The goal is to protect these systems from theft or misuse, with security measures needing to exceed standard commercial practices.
Key Recommendations & Best Practices
1. Two-Party Control (Multi-Party Authorization)
- Core Principle: No single person should have persistent access to production-critical environments.
- Implementation: Requires time-limited access with a business justification, granted by a coworker.
- Rationale: Already used in high-security domains (e.g., vaults, manufacturing, finance) and by major tech companies to defend against advanced threats and insider risk.
- Applicability: Should be applied to all systems involved in developing, training, hosting, and deploying frontier AI models. Even emerging labs can implement these controls.
2. Secure Software Development Framework
- Gold Standards: The NIST Secure Software Development Framework (SSDF) and Supply Chain Levels for Software Artifacts (SLSA).
- Key Benefit: Creates a chain of custody and provenance, allowing a deployed model to be tied back to the developing company.
- Actionable Step: Anthropic encourages extending the SSDF to explicitly encompass model development within NIST's standard-setting process.
- Precedent: U.S. Executive Order 14028 successfully motivated the software industry to adopt higher standards to retain federal contracts.
3. Public-Private Cooperation
- Proposal: Designate frontier AI research as a special sub-sector of the existing critical infrastructure IT sector.
- Purpose: To enable enhanced cooperation and information sharing between labs and government agencies against highly resourced malicious actors.
Implementation & Regulatory Path
- Immediate Action: Anthropic is implementing two-party controls, SSDF, SLSA, and other best practices.
- Regulatory Leverage: Recommends establishing these practices as procurement requirements for AI companies and cloud providers contracting with governments. This can drive broad market adoption in advance of formal regulation.
- Iterative Process: Security measures will need continuous enhancement as model capabilities scale, developed in consultation with government and industry.
Conclusion & Anthropic's Stance
- Security must not be deprioritized, even if it can sometimes interfere with productivity. Creative solutions are needed to limit its impact on research.
- Anthropic acknowledges the dual potential of AI for benefit and risk, stating they take their responsibility to build and deploy Claude safely and securely "seriously."
Related Announcements (from source)
- Expanded partnership with Google and Broadcom for next-generation compute.
- MOU with the Australian government for AI safety and research.
- Launch of the Claude Partner Network with a $100 million investment.
来源
暂无来源