June 12, 2024
Overview
Red teaming is the adversarial testing of technological systems to identify vulnerabilities. Anthropic emphasizes that while red teaming is critical for AI safety, the industry currently lacks standardized practices, making it difficult to objectively compare the safety of different models.
1. Red Teaming Methodologies
Anthropic utilizes several distinct approaches depending on the threat model and context:
A. Domain-Specific Expert Red Teaming
- Policy Vulnerability Testing (PVT): In-depth qualitative testing on topics covered by Usage Policies (e.g., child safety, election integrity, radicalization).
- Partners: Thorn, Institute for Strategic Dialogue, Global Project Against Hate and Extremism.
- Frontier Threats (National Security): Focuses on high-consequence risks including CBRN (Chemical, Biological, Radiological, and Nuclear), cybersecurity, and autonomous AI risks.
- Multilingual/Multicultural: Testing beyond English-centric perspectives to identify regional risks.
- Example: Partnered with Singapore's IMDA to test in Tamil, Mandarin, and Malay.
B. Model-Assisted Red Teaming
- Automated Red Teaming: Using a "red team / blue team" dynamic.
- Red Team Model: Generates attacks to elicit target behaviors.
- Blue Team Model: Fine-tuned on those outputs to become more robust.
- Goal: Create an iterative loop to cover more surface area than manual testing.
C. New Modalities
- Multimodal Red Teaming: Specifically for models like Claude 3 that process visual information (photos, sketches, charts).
- Risks: Fraudulent activity, visual threats to child safety, and violent extremism.
D. Open-Ended & General Red Teaming
- Crowdsourced: Using crowdworkers in controlled environments to apply their own judgment on general harms.
- Community-Based: Large-scale events like DEF CON's AI Village and the Generative Red Teaming (GRT) Challenge, which involve thousands of non-technical participants.
2. Transitioning from Qualitative to Quantitative
Anthropic outlines a "meta-challenge": converting ad hoc human testing into compounding organizational value.
The Iterative Loop:
- Ad Hoc Probing: Experts identify a threat model and manually elicit harmful behavior.
- Standardization: Red teamers refine inputs to consistently trigger the vulnerability.
- Automation: Language models generate thousands of variations of those successful inputs.
- Evaluation: Qualitative human insights become thorough, quantitative, and automated benchmarks.
3. Policy Recommendations
To foster a robust AI testing ecosystem, Anthropic proposes five actionable steps for policymakers:
- Standardization Funding: Fund NIST to develop technical standards for safe and effective red teaming.
- Independent Bodies: Support government and non-profit organizations to act as independent red team partners.
- Professional Market: Encourage a market for professional AI red teaming services with a formal certification process.
- Third-Party Access: Facilitate access for vetted/certified outside groups to conduct third-party red teaming.
- Link to Scaling: Encourage companies to tie red teaming results to Responsible Scaling Policies (RSP), determining if a model is safe to release or continue scaling.
Key Insight
"The lack of standardized practices for AI red teaming further complicates the situation... This inconsistency makes it challenging to objectively compare the relative safety of different AI systems."
Red teaming is not a one-time event but a critical component of an iterative safety lifecycle that must evolve alongside model capabilities.
来源
暂无来源