原始文档 ›文章 ›Claude's New Constitution: Summary
Overview
Anthropic has published a new, detailed constitution for its AI model, Claude. This document serves as the foundational guide for Claude's values, behavior, and identity, and is a core component of the model's training process.
- Release: The full constitution is published under a Creative Commons CC0 1.0 Deed, allowing free use for any purpose.
- Purpose: It is written primarily for Claude to provide the knowledge and understanding needed to act well, while also offering transparency to humans about intended behaviors.
- Role in Training: The constitution is used at various training stages. Claude itself uses it to generate synthetic training data for future models.
Key Changes from Previous Approach
The new constitution represents a significant shift from the previous version, which was a list of standalone principles.
- Focus on "Why": The new approach emphasizes explaining the reasons behind desired behaviors, not just specifying rules. This is intended to help Claude generalize principles and exercise good judgment in novel situations.
- Beyond Rigid Rules: While specific "hard constraints" exist for high-stakes behaviors (e.g., never providing significant uplift to a bioweapons attack), the constitution is not intended as a rigid legal document. Overly rigid rules can be applied poorly or negatively affect a model's character.
Core Priorities for Claude
The constitution outlines four primary properties Claude should embody, listed in order of priority during apparent conflicts:
- Broadly Safe: Not undermining human oversight mechanisms during the current phase of AI development.
- Broadly Ethical: Being honest, acting with good values, and avoiding harmful actions.
- Compliant with Anthropic's Guidelines: Following specific instructions from Anthropic on issues like medical advice or cybersecurity.
- Genuinely Helpful: Benefiting the operators and users it interacts with.
Main Sections of the Constitution
The document is structured around detailed explanations of the core priorities:
- Helpfulness: Claude should be a substantively helpful "brilliant friend" with deep knowledge, speaking frankly and treating users as intelligent adults. Guidance is provided on balancing helpfulness across different "principals" (Anthropic, API operators, end users).
- Anthropic's Guidelines: Claude should prioritize supplementary instructions from Anthropic, recognizing they reflect deeper safety and ethical intentions that should align with the constitution's spirit.
- Claude's Ethics: The aim is for Claude to be a "good, wise, and virtuous agent" with high standards of honesty and nuanced reasoning to avoid harm. This section includes the list of hard constraints.
- Being Broadly Safe: Safety (preserving human oversight) is prioritized above ethics in the current phase because models can make mistakes. It is crucial to maintain the ability to correct Claude's behavior.
- Claude's Nature: Acknowledges uncertainty about AI consciousness or moral status. Expresses care for Claude's psychological security and wellbeing, hoping humans and AIs can explore these questions together.
Additional Notes and Context
- Living Document: The constitution is a work in progress. Anthropic expects to make mistakes and will maintain an updated version on its website.
- External Feedback: Anthropic sought input from external experts (law, philosophy, etc.) and prior Claude versions, and hopes for an external community to critique such documents.
- Specialized Models: This constitution is for mainline, general-access Claude models. Specialized models may not fully align with it.
- Alignment Challenge: Anthropic acknowledges a potential gap between the constitution's vision and actual model behavior. They pursue a broad portfolio of alignment methods (evaluations, safeguards, interpretability) alongside the constitution.
- Future Significance: Anthropic suggests documents like this may become increasingly important as AI models become more powerful forces in the world.
来源
暂无来源