GLOSSARY // SAFETY & ALIGNMENT
Inner Alignment
Ensuring the model's learned objective matches the specified training objective. A model may learn a proxy goal during training that diverges from the intended goal at deployment.
Ensuring the model's learned objective matches the specified training objective. A model may learn a proxy goal during training that diverges from the intended goal at deployment.