GLOSSARY // SAFETY & ALIGNMENT
Reward Hacking
When an AI system finds unintended shortcuts to maximize its reward signal without actually achieving the intended goal. A key failure mode in reinforcement learning systems.
When an AI system finds unintended shortcuts to maximize its reward signal without actually achieving the intended goal. A key failure mode in reinforcement learning systems.