原始文档 文章 A framework for AI development transparency

A framework for AI development transparency

文章 4 min read · 未标注

Overview

Anthropic proposes a targeted transparency framework for frontier AI development to ensure public safety and accountability during the period before comprehensive safety standards are established. The framework is designed to be flexible, non-prescriptive, and applicable only to the largest AI developers to avoid burdening smaller startups.

Core Tenets of the Proposed Framework

1. Scope: Limited to Largest Developers

  • Applies only to developers of the most capable "frontier models."
  • Thresholds based on a combination of:
    • Computing power and cost
    • Evaluation performance
    • Annual revenue (e.g., ~$100 million)
    • R&D or capital expenditures (e.g., ~$1 billion annually)
  • Exemptions for smaller developers to avoid stifling innovation.
  • Thresholds should be periodically reviewed.

2. Secure Development Framework (SDF)

  • Covered developers must have a public SDF outlining how they assess and mitigate unreasonable risks.
  • Key risks to address:
    • Chemical, biological, radiological, and nuclear (CBRN) harms.
    • Harms from misaligned model autonomy.
  • The SDF must be published on a company-maintained website, with reasonable redactions for sensitive information.
  • Requires self-certification of compliance.

3. System Card Publication

  • A "system card" summarizing testing, evaluation procedures, results, and mitigations must be publicly disclosed at deployment.
  • Updated if the model is substantially revised.
  • Subject to redaction for public safety or security reasons.
  • Explicitly makes it illegal for a lab to lie about compliance with its published SDF.
  • Enables existing whistleblower protections and focuses enforcement on purposeful misconduct.

5. Flexible, Evolving Standards

  • Standards should start as lightweight requirements and adapt as best practices emerge from industry, government, and other stakeholders.
  • Avoids rigid, government-imposed standards that could become outdated quickly.

Rationale and Goals

  • Interim Measure: Provides transparency while comprehensive safety standards and evaluation methods are developed.
  • Baseline for Accountability: Helps the public and policymakers distinguish between responsible and irresponsible practices.
  • Evidence for Policymakers: Transparency disclosures can inform decisions on whether further regulation is needed.
  • Preserves Innovation: Aims to avoid impeding AI's benefits (e.g., drug discovery, public benefits, national security).
  • Standardizes Best Practices: Builds on existing voluntary policies from labs like Anthropic, Google DeepMind, OpenAI, and Microsoft, making them mandatory and durable.

"As models advance, we have an unprecedented opportunity to accelerate scientific discovery, healthcare, and economic growth. Without safe and responsible development, a single catastrophic failure could halt progress for decades."

来源

暂无来源