原始文档 ›文章 ›Third-Party Testing as a Key Ingredient of AI Policy
Source: Anthropic (March 25, 2024) Core Thesis: Effective third-party testing for frontier AI systems is essential to prevent societal harm and is the best foundation for AI policy.
Policy Overview & Rationale
- Problem: Frontier AI systems (e.g., Claude, Gemini, ChatGPT) are "everything machines" that don't fit traditional sector-specific regulations. They present risks of serious misuse (e.g., bioweapons, cyberattacks) and accidents (e.g., emergent, autonomous behaviors contradicting human intent).
- Solution: A robust, third-party testing and oversight regime is needed to validate safety, build trust, and complement sector-specific rules.
- Goal: Proactively design effective regulation to avoid "extreme, knee-jerk regulatory actions" after a major incident, which could be both stifling and ineffective.
- Scope: The regime should apply only to a narrow set of the most computationally-intensive, large-scale systems. The vast majority of AI systems would not be within its scope.
- Key Ingredients:
- Effective, broadly-trusted tests for measuring behavior and potential misuses.
- Trusted, legitimate third parties to administer tests and audit company procedures.
Why Self-Governance is Insufficient
- Anthropic and others have implemented self-governance (e.g., Anthropic's Responsible Scaling Policy (RSP)), but this relies on decisions by single, private actors.
- "Ultimately, testing will need to be done in a way which is broadly trusted, and it will need to be applied to everyone developing frontier systems." This is standard in sectors like food, medicine, and aerospace.
Blueprint for a Robust Testing Regime
An effective regime requires:
- A shared understanding across industry, government, and academia.
- An initial period of practice runs with third-party oversight.
- A two-stage testing regime:
- A fast, automated first stage biased toward avoiding false negatives.
- A more thorough secondary stage using expert human-led elicitation if problems are spotted.
- Increased government resources for oversight and validation.
- A carefully scoped, small set of legally mandated tests to avoid excessive regulatory burden and capture.
- A balance between safety assurance and ease of administration.
Potential Testers & Ecosystem
- Private Companies: Subcontracting or auditing (similar to accounting firms).
- Universities: Administering testing initiatives, potentially supervised by government.
- Government Agencies: Directly carrying out a small number of legally mandated tests, especially for national security risks (e.g., capabilities to speed up bioweapon creation or complex cyberattacks).
Anthropic's Commitments to Support the Regime
Anthropic will:
- Prototype a testing regime via its RSP and share learnings.
- Test third-party assessment via contractors and government partners.
- Deepen frontier red teaming work.
- Advocate for government funding for agencies like NIST, the US AI Safety Institute, and the National AI Research Resource.
- Encourage governments to build "National Research Clouds" to develop independent testing capacity.
Broader Policy Priorities Linked to Testing
- Greater Funding for Government AI Testing: Specifically for institutions like NIST.
- Public Sector AI Research Infrastructure: To expand the number of people testing and evaluating AI systems (e.g., support for the CREATE AI Act).
- Tests for National Security Capabilities: Supporting government efforts to develop classified tests for sensitive capabilities.
- Scenario Planning for Advanced Systems: Frontloading work to develop tests for hypothetical future capabilities.
Key Policy Discussions
On Openly-Disseminated / Open-Source Models
- Acknowledges the critical role of openness in AI's progress and security.
- Key Insight: "In the future it may be hard to reconcile a culture of full open dissemination of frontier AI systems with a culture of societal safety."
- If models enable significant misuse or accidents, norms may need to adjust. Developers might need to implement safeguards (e.g., classifiers, "know your customer" rules for fine-tuning).
- Resolution requires legitimate third parties to define unacceptable misuses and test both controlled (API) and openly disseminated models to understand the safety landscape.
On Regulatory Capture
- Any policy can be captured by well-resourced actors.
- Third-party testing is advocated as a solution because it builds independent capacity, identifies concrete harms, and creates a level playing field.
- Conversely, industry-led consortia might favor high-compliance approaches that advantage larger companies.
Conclusion & Call to Action
- Anthropic advocates for a "minimal viable policy" that is practical and focused.
来源
暂无来源