Cover of Human Compatible: Artificial Intelligence and the Problem of Control

Human Compatible: Artificial Intelligence and the Problem of Control

Stuart Russell

I keep this at #43 because one of the clearest accounts of the AI alignment problem from a leading researcher in the field. I have not finished learning what it asks of me.

The long version

Stuart Russell argues that the standard goal of building machines that optimize fixed objectives becomes dangerous as systems grow more capable, because human preferences are complex, uncertain, and easily misspecified. He proposes assistance-oriented AI that remains uncertain about what humans want and learns through interaction rather than pursuing a rigid target. I return to it when I need a book that makes the way intelligence is becoming a tool, a product, and a decision I still have to own personal again.

Why it is here

It presents the AI control problem from one of the field’s most influential researchers in accessible form. The book helps connect technical objective design to political, economic, and civilizational stakes without assuming that intelligence automatically produces aligned judgment. The test is whether it changes what I notice after I close the book.

How to read it

Create separate notes for current harms, foreseeable governance problems, and speculative loss-of-control scenarios. For each, list assumptions, evidence, affected actors, and interventions. Pair the policy chapters with a more skeptical or sociotechnical perspective. I will keep a note of the sentences that make me want to look away.

What it taught me

  1. 01

    A highly capable optimizer can pursue a badly specified objective with extreme effectiveness.

  2. 02

    Uncertainty about human preferences can make systems more deferential and corrigible.

  3. 03

    AI risk includes labor, surveillance, weapons, concentration of power, and long-term control.

Before and after

More about the author

Stuart Russell is a computer scientist at the University of California, Berkeley, known for foundational work in artificial intelligence and for coauthoring a standard AI textbook. He has become a leading advocate for research on provably beneficial AI.

More by the same hand