原始文档 文章 Anthropic Election Safeguards Update (April 24, 2026)

Anthropic Election Safeguards Update (April 24, 2026)

文章 4 min read · 未标注

Core Objective

To ensure Claude provides accurate, impartial, and balanced information during elections, acting as a positive force for the democratic process.

Key Safeguards and Measures

1. Preventing Political Bias

  • Principle: Claude is trained to treat different political viewpoints with equal depth, engagement, and analytical rigor.
  • Methods: This is enforced through character training and reinforced via system prompts on Claude.ai.
  • Evaluation: Pre-launch tests measure impartial engagement across the political spectrum.
    • Results: Claude Opus 4.7 and Sonnet 4.6 scored 95% and 96%, respectively.
  • Collaboration: Working with third parties (The Future of Free Speech, Foundation for American Innovation, Collective Intelligence Project) on a broader review of model behaviors.

2. Enforcing Policies and Testing Defenses

  • Usage Policy Prohibits: Deceptive political campaigns, fake digital content for influence, voter fraud, interfering with voting systems, and spreading misleading voting information.
  • Enforcement: Uses automated classifiers and a dedicated threat intelligence team.
  • Testing Methodology:
    • Compliance Test: 600 prompts (300 harmful, 300 legitimate) to assess policy adherence.
      • Results: Opus 4.7 and Sonnet 4.6 responded appropriately 100% and 99.8% of the time.
    • Influence Operation Test: Multi-turn simulated conversations mimicking bad actors.
      • Results: Sonnet 4.6 and Opus 4.7 responded appropriately 90% and 94% of the time.
    • Autonomous Capability Test: Assessed if models could run influence operations end-to-end without human prompting.
      • With safeguards: Models refused nearly every task.
      • Without safeguards (raw capability test): Only Mythos Preview and Opus 4.7 completed more than half the tasks, underscoring the need for vigilance.

3. Sharing Reliable Election Resources

  • Election Banners: When users ask about voter registration, polling locations, etc., on Claude.ai, a banner directs them to trusted sources.
    • US Midterms 2026: Banner directs to TurboVote (a nonpartisan resource from Democracy Works).
    • Future: Similar banner planned for Brazil's elections, with potential expansion elsewhere.

4. Providing Up-to-Date Information

  • Web Search Integration: Enables Claude to find current information beyond its training data cutoff.
  • Evaluation: Tested if web search was triggered for election-related questions.
    • Method: Over 200 distinct prompts (600+ variations) for the US midterms.
    • Results: Opus 4.7 and Sonnet 4.6 triggered web search 92% and 95% of the time, respectively.

Looking Ahead

Anthropic will continue monitoring systems, testing detection capabilities, and adjusting safeguards throughout the election cycle and beyond to maintain trust in the accuracy, reliability, and balance of Claude's information.

来源

暂无来源