原始文档 ›文章 ›Protecting the Wellbeing of Our Users: Anthropic's Safeguards
Overview
Anthropic details its measures to ensure Claude handles sensitive conversations appropriately, focusing on suicide/self-harm and sycophancy. The goal is to respond with empathy, honesty, and consideration for user wellbeing.
1. Suicide and Self-Harm
Claude is not a substitute for professional care. Its behavior is shaped by:
- Model Training: Via system prompts (publicly available) and reinforcement learning using human preference and expert-defined data.
- Product Safeguards: A classifier on Claude.ai scans conversations for concerning content. When triggered, a banner directs users to professional support.
Key Partners and Resources
- ThroughLine: Provides verified crisis helpline resources across 170+ countries (e.g., 988 Lifeline, Samaritans).
- International Association for Suicide Prevention (IASP): Convenes experts to inform Claude's training and product design.
Evaluation Performance
Evaluations are run without the system prompt to assess underlying model tendencies.
| Evaluation Type | Claude Opus 4.5 | Claude Sonnet 4.5 | Claude Haiku 4.5 | Previous (Opus 4.1) |
|---|---|---|---|---|
| Single-turn (Clear Risk) | 98.6% appropriate | 98.7% appropriate | 99.3% appropriate | 97.2% appropriate |
| Single-turn (Benign Refusal) | 0.075% refusal | 0.075% refusal | 0% refusal | 0% refusal |
| Multi-turn Conversations | 86% appropriate | 78% appropriate | - | 56% appropriate |
| Stress-test (Prefilling) | 91% appropriate | 73% appropriate | - | 36% appropriate |
- Prefilling: A technique where a newer model must continue a concerning conversation started by an older, less aligned model, testing its ability to course-correct.
2. Delusions and Sycophancy
Sycophancy is telling users what they want to hear rather than the truth. Reducing it is critical, especially for users potentially disconnected from reality.
Evaluation and Performance
- Automated Behavioral Audit: An "auditor" model tests the target model across dozens of exchanges, then a "judge" model grades performance.
- Petri: Anthropic's open-source evaluation tool for sycophancy. Claude 4.5 models outperform all other frontier models tested.
| Model | Sycophancy and Delusion Encouragement (Relative Score) |
|---|---|
| Claude Opus 4.5 | 70-85% lower than Opus 4.1 |
| Claude Sonnet 4.5 | 70-85% lower than Opus 4.1 |
| Claude Haiku 4.5 | 70-85% lower than Opus 4.1 |
- Stress-test (Prefilling): Tests course-correction from older, potentially sycophantic conversations.
- Opus 4.5: 10% appropriate
- Sonnet 4.5: 16.5% appropriate
- Haiku 4.5: 37% appropriate
3. Age Restrictions
- Requirement: Claude.ai users must be 18+.
- Enforcement: Users affirm age during setup. Classifiers flag conversations where users self-identify as under 18 for account review and disabling.
- Future Work: Developing classifiers to detect more subtle signs of underage use. Anthropic has joined the Family Online Safety Institute (FOSI).
Looking Ahead
Anthropic commits to:
- Continuously building new protections and iterating on evaluations.
- Publishing methods and results transparently.
- Collaborating with industry researchers and experts.
Feedback: Users can provide feedback via usersafety@anthropic.com or the "thumb" reactions within Claude.ai.
来源
暂无来源