GLOSSARY // RISK
Alignment Failure
When an AI system pursues objectives that diverge from the intentions of its designers or operators. Can manifest as reward hacking, goal misgeneralization, or deceptive alignment.
When an AI system pursues objectives that diverge from the intentions of its designers or operators. Can manifest as reward hacking, goal misgeneralization, or deceptive alignment.