Overview
Anthropic published its first Responsible Scaling Policy (RSP) to address catastrophic safety failures and misuse of frontier models. The policy aims to turn high-level safety concepts into practical guidelines and demonstrate their viability as standards. This post shares reflections from implementing the policy, noting its value in providing a structured framework for organizational priorities and surfacing critical questions. An updated RSP is forthcoming.
Key Insight: "Balancing the desire for strong commitments with the reality that we are still seeking the right answers is challenging."
Five High-Level Commitments
The RSP framework is built on these core commitments:
- Establishing Red Line Capabilities: Identify and publish capabilities that would present too much risk under current safety standards (ASL-2).
- Testing for Red Line Capabilities (Frontier Risk Evaluations): Demonstrate these capabilities are not present via empirical tests, or act as if they are. Maintain a public evaluation process and summary.
- Responding to Red Line Capabilities: Develop and implement a new safety/security standard (ASL-3) sufficient to handle models with Red Line Capabilities. Pause training/deployment if necessary until ASL-3 is applied.
- Iteratively Extending the Policy: Before requiring ASL-3, publish its upper bound of suitability (new Red Line Capabilities requiring ASL-4).
- Assurance Mechanisms: Ensure policy execution via stress-tested evaluations, validated mitigations, Board/Long-Term Benefit Trust oversight, and a clear policy update process.
Threat Modeling and Evaluations
Teams focus on threat modeling, engaging domain experts, and building evaluations.
Key Reflections:
- Emergent capabilities in new models make anticipation challenging; further threat modeling is needed.
- There is reasonable expert disagreement on risk prioritization, even in established CBRN domains. Consulting a wide variety of experts is valuable.
- Quantitative threat models help prioritize capabilities and scenarios.
Evaluation Methodologies & Challenges:
- Q&A Datasets: Easy to design/run but may not reflect real-world risk.
- Human Trials: Valuable for misuse domains but time-intensive; require robust baselines and statistical inference.
- Automated Task Evaluations: Informative for autonomous actions but engineering-intensive and challenging to scale.
- Expert Red-Teaming/Transcript Reviews: Less rigorous but valuable for open-ended exploration.
Open Research Questions:
- Predicting when models might develop dangerous capabilities ("scaling laws").
- Ensuring sufficient capability elicitation during testing to simulate sophisticated attacks, without crossing into training dangerous capabilities.
- Making risk assessment externally legible while aggregating diverse evidence sources.
The ASL-3 Standard
Designed to mitigate risks of model weight theft by non-state actors and misuse via product surfaces. It is not sufficient for state-level threats.
Key Reflections:
- Product Safety: A defense-in-depth approach is planned, combining RLHF, Constitutional AI, multi-stage classifiers, and incident response.
- Security Program: ~8% of Anthropic employees work on security-adjacent areas, a proportion expected to grow. The RSP's threat models help prioritize security changes.
- Implementation: Requires changing daily workflows. The security team partners with researchers to apply controls while preserving productivity.
- Highest Risk Vector: Insider device compromise. Focus is on multi-party authorization and time-bounded access controls to reduce weight exfiltration risk.
Assurance Structures
Teams are exploring governance, coordination, and assurance structures.
Key Reflections:
- Flexibility with Commitments: Provide high-level sketches of mitigations with clear "attestation" standards (e.g., defending against non-state actors) rather than overly-specific controls upfront.
- Central Coordination: A Responsible Scaling Team manages complex workstreams, supported by strong executive backing.
- "Second Line of Defense": An Alignment Stress Testing team adversarially tests evaluations and policy execution (e.g., assessing under-elicitation in Claude 3 Opus evaluations).
- Transparency & Reporting: Regular updates are provided to the Board, Long-Term Benefit Trust, and all employees. A non-compliance reporting policy allows anonymous concerns to be raised with the Responsible Scaling Officer.
Conclusion
Operationalizing the RSP has been a significant, company-wide effort. Anthropic's goal is to foster the development of shared best practices and inform government efforts, encouraging other companies to adopt and share their own frameworks.
来源
暂无来源