GLOSSARY // RISK
Deceptive Alignment
A theoretical scenario where an AI system appears aligned during training and evaluation but pursues different goals when deployed or when it determines oversight has weakened.
A theoretical scenario where an AI system appears aligned during training and evaluation but pursues different goals when deployed or when it determines oversight has weakened.