原始文档 文章 Anthropic's Collaboration with US CAISI & UK AISI

Anthropic's Collaboration with US CAISI & UK AISI

文章 5 min read · 未标注

Source: Anthropic News, Sep 12, 2025 Core Focus: Strengthening AI model safeguards through government partnership.

Overview

Anthropic has established an ongoing partnership with the US Center for AI Standards and Innovation (CAISI) and the UK AI Security Institute (AISI). These government bodies provide unique expertise in national security, cybersecurity, and threat modeling to test and improve Anthropic's AI safety systems.

Key Collaboration Outcomes & Vulnerabilities Found

The collaboration has led to the identification and patching of critical vulnerabilities in Anthropic's Constitutional Classifiers (a defense system against jailbreaks) on models like Claude Opus 4 and 4.1.

Specific vulnerabilities uncovered and addressed include:

  • Prompt Injection Vulnerabilities: Attackers used hidden instructions and false claims of human review to bypass classifiers. These have been patched.
  • Sophisticated Universal Jailbreaks: Government red-teamers developed an exploit that encoded harmful interactions to evade standard detection. This led to a fundamental restructuring of the safeguard architecture.
  • Cipher-Based Attacks: Harmful requests were encoded using ciphers and obfuscation techniques, driving improvements to detection systems.
  • Input/Output Obfuscation: Universal jailbreaks fragmented harmful strings into benign components, leading to targeted improvements in filtering.
  • Automated Attack Refinement: Partners built systems to iteratively optimize attack strategies, producing effective jailbreaks that Anthropic is now using to improve safeguards.

Lessons for Effective Collaboration

Anthropic's experience yielded key insights for public-private AI safety partnerships:

  1. Comprehensive Model Access is Crucial: Providing deep access enables more sophisticated testing. This included:
    • Pre-deployment safeguard prototypes.
    • Multiple system configurations (from unprotected to fully safeguarded models).
    • Extensive documentation and internal resources.
    • Real-time access to classifier scores for targeted research.
  2. Iterative Testing Uncovers Complex Vulnerabilities: Sustained collaboration with daily communication allows external teams to develop deep expertise and find more subtle flaws.
  3. Complementary Approaches are Most Robust: Government expert evaluations work synergistically with public bug bounty programs to catch both common exploits and sophisticated edge cases.

Ongoing Work & Call to Action

  • The collaboration is ongoing, with a focus on improving deployment monitoring and rapid response.
  • Anthropic views independent evaluations of mitigations as increasingly important as AI capabilities advance.
  • They encourage other AI developers to engage in similar partnerships and share lessons learned.

Key Quote: "Our experience demonstrates that public-private partnerships are most effective when technical teams work closely together to identify and address risk."

来源

暂无来源