GLOSSARY // TECHNICAL
RLHF
Reinforcement Learning from Human Feedback — a training technique where human preferences are used to fine-tune AI behavior. The primary method used to make LLMs "helpful and harmless." Has known limitations including sycophancy and reward hacking.