原始文档 ›文章 ›Frontier Threats Red Teaming for AI Safety - Summary
Overview
Anthropic details its approach to "frontier threats red teaming," a specialized form of adversarial testing focused on national security risks (e.g., biosecurity, cybersecurity). This work is part of commitments made at the White House on July 21, 2023. The goal is to evaluate baseline risks and create repeatable, scalable evaluation methods.
Key Findings from Biological Risk Red Teaming
- Methodology: Over 150 hours were spent with top biosecurity experts over six months. Experts used a secure, unmonitored interface to probe model capabilities.
- Core Concern: Current frontier models can sometimes produce sophisticated, accurate, and expert-level knowledge. While infrequent in most areas, this capability is present in some domains and grows with model scale.
- Risk Assessment: Unmitigated LLMs could accelerate a bad actor's misuse of biology relative to having only internet access. These risks are considered near-term (2-3 years), not distant.
- Mitigations Discovered:
- Training Process Changes: Techniques like Constitutional AI help models better distinguish harmful from harmless biological uses.
- Classifier-Based Filters: Deployed in public models to block the chaining of multiple expert-level pieces of information needed for harm.
Methodology for Frontier Threats Red Teaming
- Requires Significant Investment: Involves 100+ hours of collaboration between domain experts and LLM experts.
- Process: Starts by defining threat models (what information is dangerous, how it combines to create harm). Experts then learn to interact with and "jailbreak" models to uncover true capabilities.
- Key Objective: Build new, automated evaluations based on expert knowledge to make testing repeatable and scalable.
- Challenge: Findings are sensitive, requiring trusted third-party partnerships and strong information security.
Future Research & Scaling
- Priority Experiments: Measure the speedup LLMs provide for producing harm compared to a search engine, including with future models (next-gen, tool-using, multimodal).
- Call to Action: Frontier model developers must urgently conduct more analysis and develop stronger mitigations, sharing findings with responsible industry developers and select government agencies.
- Open Model Risk: Bad actors could extract harmful capabilities from smaller, fine-tuned models adapted from openly available, sufficiently capable base models.
- Anthropic's Commitment: They are scaling up their frontier threats red teaming team and establishing a responsible disclosure process between labs and stakeholders.
- Long-Term Vision: This research agenda is applicable to other long-term risks like deception. It will help identify future capabilities models should not have and build corresponding mitigations.
Collaboration & Next Steps
- Anthropic is briefing government and labs on detailed findings.
- They are piloting a responsible disclosure process for the community.
- They advocate for new, impartial third-party organizations to conduct national security evaluations between stakeholders.
- Anthropic is hiring mission-driven technical researchers to join this team and is open to supporting other labs or evaluation organizations.
Key Quote: "Current models are only showing the first very early signs of risks of this kind, which makes this our window to evaluate nascent risks and mitigate them before they become acute."
来源
暂无来源