GLOSSARY // SAFETY & ALIGNMENT
Interpretability
The ability to understand how an AI system arrives at its outputs. Mechanistic interpretability aims to reverse-engineer the internal computations of neural networks.
The ability to understand how an AI system arrives at its outputs. Mechanistic interpretability aims to reverse-engineer the internal computations of neural networks.