原始文档 ›文章 ›Building Safeguards for Claude
Claude is designed to amplify human potential while ensuring capabilities are channeled toward beneficial outcomes. The Safeguards team identifies misuse, responds to threats, and builds defenses to keep Claude helpful and safe. The team combines experts in policy, enforcement, product, data science, threat intelligence, and engineering.
Multi-Layered Approach
Safeguards operate across the entire model lifecycle:
- Developing policies
- Influencing model training
- Testing for harmful outputs
- Real-time policy enforcement
- Identifying novel misuses and attacks
Key Components
1. Policy Development
- Usage Policy: Framework defining permitted/prohibited uses, covering areas like child safety, election integrity, cybersecurity, healthcare, and finance.
- Unified Harm Framework: Structured lens assessing potential harms across five dimensions: physical, psychological, economic, societal, and individual autonomy. Considers likelihood and scale of misuse.
- Policy Vulnerability Testing: Partners with external domain experts (terrorism, child safety, mental health) to stress-test policies against challenging prompts.
- Example: Partnered with Institute for Strategic Dialogue during 2024 U.S. election to address outdated information, resulting in a banner directing users to authoritative sources like TurboVote.
2. Claude's Training
- Collaborative process with fine-tuning teams to prevent harmful behavior.
- Evaluation and detection identify harmful outputs, leading to solutions like updating reward models or adjusting system prompts.
- Partners with domain specialists (e.g., ThroughLine for crisis support) to refine responses to sensitive topics like self-harm, ensuring nuanced engagement rather than refusal.
- Claude learns to:
- Decline assistance with harmful illegal activities
- Recognize attempts to generate malicious code, fraudulent content, or plan harmful activities
- Discuss sensitive topics with care while distinguishing from actual harm attempts
3. Testing and Evaluation
Pre-deployment evaluations include:
- Safety evaluations: Assess adherence to Usage Policy across clear violations, ambiguous contexts, and multi-turn conversations. Uses model grading with human review.
- Risk assessments: For high-risk domains (cyber harm, CBRNE), conducts AI capability uplift testing with government/private partners. Defines threat models and assesses safeguards.
- Bias evaluations: Checks for reliable, accurate responses across contexts. Tests political bias with opposing viewpoints and identity attribute biases (gender, race, religion).
- Example: Pre-launch evaluations of computer use tool identified spam generation risks, leading to new detection methods and enforcement mechanisms before launch.
- Results reported in system cards released with each model family.
4. Real-Time Detection and Enforcement
- Uses classifiers (prompted or fine-tuned Claude models) to detect policy violations in real-time. Multiple classifiers can run simultaneously.
- Specialized detection for child sexual abuse material (CSAM) via image hash comparison.
- Enforcement actions:
- Response steering: Adjusts Claude's interpretation/response in real-time (e.g., adding system prompt instructions for spam/malware attempts). Can stop responses entirely in narrow cases.
- Account enforcement: Investigates violation patterns; may issue warnings or terminate accounts. Includes defenses against fraudulent account creation.
- Challenge: Classifiers must process trillions of tokens while limiting compute overhead and false positives.
5. Ongoing Monitoring and Investigation
- Claude insights tool: Measures real-world use via privacy-preserving topic clustering. Informs guardrails based on research (e.g., emotional impacts).
- Hierarchical summarization: Condenses interactions into summaries to identify account-level concerns (e.g., automated influence operations).
- Threat intelligence: Studies severe misuses, identifies adversarial patterns, compares abuse indicators against typical usage, cross-references external threat data (open source, industry reports), and monitors bad actor channels (social media, hacker forums). Findings shared in public threat intelligence reports.
Collaboration and Forward Look
- Actively seeks feedback from users, researchers, policymakers, and civil society.
- Maintains an ongoing bug bounty program for testing defenses.
- Actively seeking people to join the Safeguards team.
来源
暂无来源