原始文档 ›文章 ›Summary: The Case for Targeted Regulation - Anthropic
Core Argument
Anthropic argues for urgently enacted, narrowly-targeted AI regulation within the next 18 months to mitigate catastrophic risks (e.g., cyber, CBRN misuse) while preserving innovation. Delay risks poorly-designed, reactive regulation.
Key Evidence of Urgency & Risk
- Rapid Capability Growth: AI systems show dramatic improvement in math, reasoning, and coding.
- Cyber Offense Potential: Performance on the SWE-bench coding task improved from 1.96% (Oct 2023) to 49% (Oct 2024). Anthropic's Frontier Red Team finds current models can assist with cyber offense tasks.
- CBRN Knowledge Potential: The UK AI Safety Institute found models provide "expert-level knowledge about biology and chemistry... on par with those given by PhD-level experts."
- Benchmark Progress: Scores on the hardest section of the GPQA benchmark grew from 38.8% (Nov 2023) to 77.3% (Sep 2024), nearing the human expert score of 81.2%.
- Timeline: Anthropic previously warned of real cyber/CBRN risks within 2-3 years; they now believe we are "substantially closer."
Anthropic's Model: The Responsible Scaling Policy (RSP)
An adaptive, internal framework for managing catastrophic risk.
- Core Principles:
- Proportionate: Safety/security measures scale with defined model capability thresholds.
- Iterative: Regularly re-evaluated based on model progress.
- Key Benefits:
- Drives investment in security and safety evaluations ahead of time.
- Forces concrete, specific threat modeling.
- Encourages transparency and helps meet voluntary commitments (e.g., White House, Bletchley Park).
- Conclusion: RSPs are a "workable policy" for companies to remain competitive while managing risk, but are not a substitute for regulation.
Principles for Effective AI Regulation
Based on their RSP experience, Anthropic identifies three key elements:
- Transparency: Require companies to publish RSP-like policies and risk evaluations for new models, with a verification mechanism.
- Incentivizing Better Practices: Regulation should encourage robust RSPs. Mechanisms could include specifying threat models, setting standards, or fostering a "race to the top." Flexibility is critical due to rapid technological change.
- Simplicity and Focus: Regulations must be "surgical" and directly tied to preventing catastrophic risks. Unnecessary burdens or complexity are counterproductive.
Call to Action & Implementation
- Timeline: Critical work needed over the next year.
- Jurisdiction: Prefers federal regulation in the US for uniformity and expertise, but supports state-level action as a backstop given federal pace concerns. Principles are applicable internationally.
- Goal: Develop a framework agreeable to a wide range of stakeholders, even if imperfect initially.
Key FAQ Insights
- Regulation by Use Case vs. Model: Use-case regulation is impractical for general-purpose consumer AI (e.g., Claude.ai). Regulating the underlying model's fundamental properties is more effective and trackable.
- Scope of Risks: This post focuses on catastrophic, frontier-model risks (cyber, CBRN). Near-term risks (deepfakes, child safety) are addressed separately.
- Innovation & Competition: Well-designed, proportionate regulation (like the RSP framework) can manage risk with minimal burden. Safety research can have "unexpected spillover benefits" to AI science, and strong security protects IP.
- Open Source: Regulation should focus on empirically measured risks, not the open/closed-weight distinction. It should neither favor nor disfavor open models unless tests show differing risk levels.
来源
暂无来源