原始文档 文章 Reflections on our Responsible Scaling Policy

Reflections on our Responsible Scaling Policy

文章 7 min read · 未标注

Overview

Anthropic published its first Responsible Scaling Policy (RSP) to address catastrophic safety failures and misuse of frontier models. The policy aims to turn high-level safety concepts into practical guidelines and demonstrate their viability as standards. This post shares reflections from implementing the policy, noting its value in providing a structured framework for organizational priorities and surfacing critical questions. An updated RSP is forthcoming.

Key Insight: "Balancing the desire for strong commitments with the reality that we are still seeking the right answers is challenging."

Five High-Level Commitments

The RSP framework is built on these core commitments:

  1. Establishing Red Line Capabilities: Identify and publish capabilities that would present too much risk under current safety standards (ASL-2).
  2. Testing for Red Line Capabilities (Frontier Risk Evaluations): Demonstrate these capabilities are not present via empirical tests, or act as if they are. Maintain a public evaluation process and summary.
  3. Responding to Red Line Capabilities: Develop and implement a new safety/security standard (ASL-3) sufficient to handle models with Red Line Capabilities. Pause training/deployment if necessary until ASL-3 is applied.
  4. Iteratively Extending the Policy: Before requiring ASL-3, publish its upper bound of suitability (new Red Line Capabilities requiring ASL-4).
  5. Assurance Mechanisms: Ensure policy execution via stress-tested evaluations, validated mitigations, Board/Long-Term Benefit Trust oversight, and a clear policy update process.

Threat Modeling and Evaluations

Teams focus on threat modeling, engaging domain experts, and building evaluations.

Key Reflections:

  • Emergent capabilities in new models make anticipation challenging; further threat modeling is needed.
  • There is reasonable expert disagreement on risk prioritization, even in established CBRN domains. Consulting a wide variety of experts is valuable.
  • Quantitative threat models help prioritize capabilities and scenarios.

Evaluation Methodologies & Challenges:

  • Q&A Datasets: Easy to design/run but may not reflect real-world risk.
  • Human Trials: Valuable for misuse domains but time-intensive; require robust baselines and statistical inference.
  • Automated Task Evaluations: Informative for autonomous actions but engineering-intensive and challenging to scale.
  • Expert Red-Teaming/Transcript Reviews: Less rigorous but valuable for open-ended exploration.

Open Research Questions:

  • Predicting when models might develop dangerous capabilities ("scaling laws").
  • Ensuring sufficient capability elicitation during testing to simulate sophisticated attacks, without crossing into training dangerous capabilities.
  • Making risk assessment externally legible while aggregating diverse evidence sources.

The ASL-3 Standard

Designed to mitigate risks of model weight theft by non-state actors and misuse via product surfaces. It is not sufficient for state-level threats.

Key Reflections:

  • Product Safety: A defense-in-depth approach is planned, combining RLHF, Constitutional AI, multi-stage classifiers, and incident response.
  • Security Program: ~8% of Anthropic employees work on security-adjacent areas, a proportion expected to grow. The RSP's threat models help prioritize security changes.
  • Implementation: Requires changing daily workflows. The security team partners with researchers to apply controls while preserving productivity.
  • Highest Risk Vector: Insider device compromise. Focus is on multi-party authorization and time-bounded access controls to reduce weight exfiltration risk.

Assurance Structures

Teams are exploring governance, coordination, and assurance structures.

Key Reflections:

  • Flexibility with Commitments: Provide high-level sketches of mitigations with clear "attestation" standards (e.g., defending against non-state actors) rather than overly-specific controls upfront.
  • Central Coordination: A Responsible Scaling Team manages complex workstreams, supported by strong executive backing.
  • "Second Line of Defense": An Alignment Stress Testing team adversarially tests evaluations and policy execution (e.g., assessing under-elicitation in Claude 3 Opus evaluations).
  • Transparency & Reporting: Regular updates are provided to the Board, Long-Term Benefit Trust, and all employees. A non-compliance reporting policy allows anonymous concerns to be raised with the Responsible Scaling Officer.

Conclusion

Operationalizing the RSP has been a significant, company-wide effort. Anthropic's goal is to foster the development of shared best practices and inform government efforts, encouraging other companies to adopt and share their own frameworks.

来源

暂无来源