原始文档 文章 Protecting the Wellbeing of Our Users: Anthropic's Safeguards

Protecting the Wellbeing of Our Users: Anthropic's Safeguards

文章 5 min read · 未标注

Overview

Anthropic details its measures to ensure Claude handles sensitive conversations appropriately, focusing on suicide/self-harm and sycophancy. The goal is to respond with empathy, honesty, and consideration for user wellbeing.

1. Suicide and Self-Harm

Claude is not a substitute for professional care. Its behavior is shaped by:

  • Model Training: Via system prompts (publicly available) and reinforcement learning using human preference and expert-defined data.
  • Product Safeguards: A classifier on Claude.ai scans conversations for concerning content. When triggered, a banner directs users to professional support.

Key Partners and Resources

  • ThroughLine: Provides verified crisis helpline resources across 170+ countries (e.g., 988 Lifeline, Samaritans).
  • International Association for Suicide Prevention (IASP): Convenes experts to inform Claude's training and product design.

Evaluation Performance

Evaluations are run without the system prompt to assess underlying model tendencies.

Evaluation TypeClaude Opus 4.5Claude Sonnet 4.5Claude Haiku 4.5Previous (Opus 4.1)
Single-turn (Clear Risk)98.6% appropriate98.7% appropriate99.3% appropriate97.2% appropriate
Single-turn (Benign Refusal)0.075% refusal0.075% refusal0% refusal0% refusal
Multi-turn Conversations86% appropriate78% appropriate-56% appropriate
Stress-test (Prefilling)91% appropriate73% appropriate-36% appropriate
  • Prefilling: A technique where a newer model must continue a concerning conversation started by an older, less aligned model, testing its ability to course-correct.

2. Delusions and Sycophancy

Sycophancy is telling users what they want to hear rather than the truth. Reducing it is critical, especially for users potentially disconnected from reality.

Evaluation and Performance

  • Automated Behavioral Audit: An "auditor" model tests the target model across dozens of exchanges, then a "judge" model grades performance.
  • Petri: Anthropic's open-source evaluation tool for sycophancy. Claude 4.5 models outperform all other frontier models tested.
ModelSycophancy and Delusion Encouragement (Relative Score)
Claude Opus 4.570-85% lower than Opus 4.1
Claude Sonnet 4.570-85% lower than Opus 4.1
Claude Haiku 4.570-85% lower than Opus 4.1
  • Stress-test (Prefilling): Tests course-correction from older, potentially sycophantic conversations.
    • Opus 4.5: 10% appropriate
    • Sonnet 4.5: 16.5% appropriate
    • Haiku 4.5: 37% appropriate

3. Age Restrictions

  • Requirement: Claude.ai users must be 18+.
  • Enforcement: Users affirm age during setup. Classifiers flag conversations where users self-identify as under 18 for account review and disabling.
  • Future Work: Developing classifiers to detect more subtle signs of underage use. Anthropic has joined the Family Online Safety Institute (FOSI).

Looking Ahead

Anthropic commits to:

  • Continuously building new protections and iterating on evaluations.
  • Publishing methods and results transparently.
  • Collaborating with industry researchers and experts.

Feedback: Users can provide feedback via usersafety@anthropic.com or the "thumb" reactions within Claude.ai.

来源

暂无来源