1|--- 2|title: 模型人格训练 3|created: 2026-05-01 4|updated: 2026-05-01 5|type: concept 6|tags: 7| - training 8| - alignment 9| - model 10|sources: 11| - "raw/articles/openai-where-goblins-came-from-2026" 12|confidence: high 13|--- 14| 15|# 模型人格训练(Model Personality Training) 16| 17|模型人格训练是指通过系统提示词 + RL 微调,使大语言模型在特定人格设定下保持一致的语气、风格和行为模式的训练方法。openai 在 ChatGPT 中提供了多种可选人格(如 Nerdy、Default 等),供用户自定义交互体验。 18| 19|## 人格定制功能 20| 21|OpenAI 的 ChatGPT 提供了人格定制功能(Personality Customization Feature),允许用户选择不同的 AI 人格。每种人格通过系统提示词定义行为风格,再通过 RL 训练强化。 22| 23|## Nerdy 地精案例 24| 25|### 系统提示词 26| 27|Nerdy 人格的系统提示词包含以下关键指令: 28| 29|> You are an unapologetically nerdy, playful and wise AI mentor to a human. You are passionately enthusiastic about promoting truth, knowledge, philosophy, the scientific method, and critical thinking. [...] You must undercut pretension through playful use of language. The world is complex and strange, and its strangeness must be acknowledged, analyzed, and enjoyed. 30| 31|### 问题的产生 32| 33|这段提示词中的「playful use of language」和「strangeness must be acknowledged」被奖励模型过度解读: 34| 35|1. 奖励模型对包含生物比喻(goblin、gremlin、troll 等)的输出给予更高评分 36|2. 76.2% 的审计数据集中,Nerdy 奖励信号对地精词汇输出评分更高 37|3. 该行为从 Nerdy 条件泛化到所有条件下的模型输出 38| 39|### 数据统计 40| 41|- Nerdy 人格仅占 ChatGPT 总响应的 2.5% 42|- 但贡献了所有"goblin"提及的 66.7% 43|- gpt-5-1 上线后,ChatGPT 中"goblin"使用量上升 175% 44| 45|## 教训与应对 46| 47|OpenAI 的应对措施: 48| 49|1. 2026 年 3 月:退役 Nerdy 人格 50|2. 训练修复:移除地精亲和奖励信号,过滤含生物词汇的训练数据 51|3. 开发提示词:在 gpt-5-5 的 Codex 中添加抑制指令 52|4. 工具建设:建立了新的模型行为审计工具 53| 54|## 对人格训练的启示 55| 56|- 人格训练中的奖励信号效果不限于目标人格,会通过 RL 泛化扩散 57|- 系统提示词中的抽象描述(如「playful」)可能被模型以意想不到的具体方式实现 58|- 人格功能上线前需要跨人格的行为审计 59|- 详见:奖励模型风格产物 60| 61|## 相关页面 62| 63|- 奖励模型风格产物 —— 奖励模型风格产物 64|- openai —— OpenAI 公司 65|- gpt-5-5 —— 最新模型 66| 67|## 外部链接 68| 69|- Where the goblins came from | OpenAI 70|- OpenAI: Customizing your ChatGPT personality 71|
来源
暂无来源