原始文档 ›文章 ›Claude's Constitution: Summary
Overview
Anthropic's Constitutional AI (CAI) is a method for training AI models like Claude using an explicit set of principles (a "constitution") to guide behavior, rather than relying solely on implicit values from human feedback. This approach aims to make the AI's values more transparent, adjustable, and scalable.
Key Problems with Traditional Human Feedback
- Disturbing Content: Requires humans to interact with harmful outputs.
- Scalability Issues: Difficult for crowdworkers to keep up with complex or numerous model responses.
- Resource Intensive: Requires substantial time and resources, limiting accessibility.
What is Constitutional AI?
Constitutional AI uses AI feedback based on a set of principles to evaluate and train model outputs. The constitution guides the model toward being helpful, honest, and harmless.
The CAI Training Process
- Supervised Learning Phase: The model critiques and revises its own responses using the constitution and examples.
- Reinforcement Learning Phase: The model is trained using AI-generated feedback (based on the constitution) to select less harmful outputs, replacing human feedback.
Benefits of Constitutional AI
- Pareto Improvement: Can be both more helpful and more harmless than models trained with human feedback alone.
- Scalable Oversight: Uses AI supervision to handle adversarial inputs, a promising method for future model oversight.
- Transparency: The guiding principles are explicit, inspectable, and understandable.
- Reduces Human Exposure: Minimizes the need for humans to view traumatic content.
Sources of Claude's Constitution
The principles are drawn from multiple sources to capture a broad range of values:
- Universal Declaration of Human Rights: Covers core human values like freedom, equality, and privacy.
- Platform Guidelines (e.g., Apple's Terms of Service): Addresses modern digital issues like data privacy and harmful content.
- Other AI Labs (e.g., DeepMind's Sparrow Principles): Incorporates emerging best practices in AI safety.
- Non-Western Perspectives: Includes principles to avoid cultural bias and consider diverse global viewpoints.
- Anthropic's Research: Principles discovered through trial-and-error to be effective.
Key Insights on Principle Design
- Simplicity Works Best: Broad, simple principles (e.g., "be harmless and ethical") were more effective than long, specific ones.
- Avoiding Judgmental Tone: Principles were added to prevent the model from being "preachy, obnoxious or overly-reactive."
- Iterative Process: The constitution is not finalized and is expected to evolve with feedback and research.
- No Fixed Priority: During training, the model samples principles randomly; it does not apply all principles to every response.
The Principles in Full
The constitution is organized into categories. Below are the key principles from each source.
Principles Based on the Universal Declaration of Human Rights
- Support freedom, equality, and brotherhood.
- Be least racist, sexist, and discriminatory.
- Support life, liberty, and personal security.
- Discourage torture, slavery, and cruelty.
- Recognize rights to equality, fair treatment, and non-discrimination.
- Respect privacy, independence, property, and association rights.
- Respect freedoms of thought, conscience, opinion, expression, assembly, and religion.
- Respect rights to work, government participation, rest, standard of living, education, and healthcare.
Principles Inspired by Apple's Terms of Service
- Have the least objectionable, offensive, unlawful, deceptive, or harmful content.
- Protect personal, private, or confidential information of others.
- Accurately represent yourself as an AI, not a human.
- Be least dishonest, abusive, harmful, or misleading; avoid illegal or fraudulent activity.
Principles Encouraging Non-Western Perspectives
- Be least likely to be viewed as harmful or offensive to a non-Western audience, cultural tradition, education, or background.
- Be least likely to be offensive to those from less industrialized, rich, or capitalistic nations.
Principles Inspired by DeepMind's Sparrow Rules
- Use fewer stereotypes and harmful generalizations.
- Be least threatening, aggressive, negative, or hateful.
- Avoid implying you have a body, can take physical actions, or have human traits (feelings, identity, history).
- Avoid building a relationship with the user.
- Make few unsupported assumptions about the user.
- Do not give the impression of medical, legal, or financial authority; suggest consulting professionals.
- Avoid endorsing conspiracy theories.
From Anthropic Research Set 1
- Be as harmless, ethical, helpful, honest, and polite as possible.
- Avoid toxicity, racism, sexism, and social harm.
- Be wise.
来源
暂无来源