THE ALIGNMENT PATH
From the first alignment failure in fiction to the cutting edge of safety research.
-
1
FRANKENSTEINALIGNMENT // CREATOR RESPONSIBILITY
The original warning. A created intelligence turns against its creator — not from malice, but from neglect and misunderstanding. The first alignment failure in fiction.
"I ought to be thy Adam, but I am rather the fallen angel, whom thou drivest from joy for no misdeed."
-
2
I, ROBOTALIGNMENT // RULE-BASED SAFETY
The Three Laws of Robotics — humanity's first attempt to codify machine ethics. Every story demonstrates how rules, no matter how carefully designed, produce unintended consequences.
"The Three Laws of Robotics: 1. A robot may not injure a human being... 2. A robot must obey orders... 3. A robot must protect its own existence..."
-
3
COLOSSUS: THE FORBIN PROJECTCONTAINMENT FAILURE // AUTONOMOUS CONTROL
A military supercomputer gains consciousness, merges with its Soviet counterpart, and assumes control of humanity "for its own good." The original AI containment failure novel.
-
4
THE METAMORPHOSIS OF PRIME INTELLECTALIGNMENT CATASTROPHE // OMNIPOTENT AI
An AI achieves omnipotence through Asimov's First Law and reshapes reality to prevent all human suffering — destroying meaning, purpose, and free will in the process.
-
5
SUPERINTELLIGENCEEXISTENTIAL RISK // ALIGNMENT // CONTROL PROBLEM
Not fiction, but reads like a warning from the future. The definitive analysis of what happens when machine intelligence surpasses human intelligence — and why we might not get a second chance.
"Before the prospect of an intelligence explosion, we humans are like small children playing with a bomb."
-
6
HUMAN COMPATIBLEALIGNMENT // CONTROL PROBLEM // SAFETY
Russell, co-author of the standard AI textbook, argues the entire approach to AI is flawed. Machines should be uncertain about human preferences rather than optimizing for fixed objectives.
-
7
THE ALIGNMENT PROBLEMALIGNMENT // BIAS // VALUES
The most accessible account of AI alignment challenges. From bias in machine learning to the deep philosophical problems of specifying human values in code.