AI Terms Academy
AI 伦理与安全 · AI Ethics

AI Safety

确保 AI 系统不会对人造成伤害——包括内容安全、隐私、可控、不被恶意使用等。

AI 安全

高阶 · Advanced#安全#对齐

Plain-language definition

一句话定义

确保 AI 系统不会对人造成伤害——包括内容安全、隐私、可控、不被恶意使用等。

Original definition

英文原文定义

AI Safety is the research and practice of ensuring AI systems behave as intended, avoiding harm — encompassing alignment, robustness, privacy, and misuse prevention.

Two levels

分难度讲解

入门版

AI 很强大,但「强大」也可能是危险。AI 安全研究的就是:怎么让 AI 不被坏人利用、不输出有害内容、不做违背人类意图的事、可控可关停。

进阶版

AI 安全覆盖:对齐(Alignment,让模型目标与人类一致)、鲁棒性(对抗样本、分布外)、可解释性、隐私(成员推断、数据泄露)、红队测试(Red Teaming)、越狱(jailbreak)防御、价值观学习。

Real-world example

真实例子

ChatGPT 拒绝教你「怎么制造危险物品」、会对敏感问题保持中立——这些都是 AI 安全工程的结果。

Keep exploring

相关术语