Human Compatible: Artificial Intelligence and the Problem of Control
I keep this at #43 because one of the clearest accounts of the AI alignment problem from a leading researcher in the field. I have not finished learning what it asks of me.
The long version
Stuart Russell argues that the standard goal of building machines that optimize fixed objectives becomes dangerous as systems grow more capable, because human preferences are complex, uncertain, and easily misspecified. He proposes assistance-oriented AI that remains uncertain about what humans want and learns through interaction rather than pursuing a rigid target. I return to it when I need a book that makes the way intelligence is becoming a tool, a product, and a decision I still have to own personal again.
Why it is here
It presents the AI control problem from one of the field’s most influential researchers in accessible form. The book helps connect technical objective design to political, economic, and civilizational stakes without assuming that intelligence automatically produces aligned judgment. The test is whether it changes what I notice after I close the book.
How to read it
Create separate notes for current harms, foreseeable governance problems, and speculative loss-of-control scenarios. For each, list assumptions, evidence, affected actors, and interventions. Pair the policy chapters with a more skeptical or sociotechnical perspective. I will keep a note of the sentences that make me want to look away.
What it taught me
- 01
A highly capable optimizer can pursue a badly specified objective with extreme effectiveness.
- 02
Uncertainty about human preferences can make systems more deferential and corrigible.
- 03
AI risk includes labor, surveillance, weapons, concentration of power, and long-term control.
Before and after
What makes it easier
Where it leads
The Alignment Problem
Brian Christian
The Precipice
Toby Ord
More about the author
Stuart Russell is a computer scientist at the University of California, Berkeley, known for foundational work in artificial intelligence and for coauthoring a standard AI textbook. He has become a leading advocate for research on provably beneficial AI.
More by the same hand