Handoff is when a group of humans grant AIs a position of trust, decision-making, or both.
A paradigm example would be a frontier AI company handing off to their own AIs. These AIs might be in charge of: prioritising between research agendas; deciding how to develop and deploy successor AIs; managing relations with governments (foreign and domestic); managing relations with clients, suppliers, competitors, and the public; managing philanthropic ventures; etc.
Handoff varies along several dimensions:
- Trust. If the AIs wanted to "screw us over", how badly could they do so?
- Decision-making. How much scrutiny are humans still applying to the AIs' decisions?
- Scale. Did a single individual hand off? A team within a company? An entire company? A government? A coalition of governments? Humanity as a whole?
- Scope. Which decisions are AIs making? Just R&D, or all decisions within the company?
- Capability. How capable are the AIs at handoff? Greenblatt argues for handing off near the minimum viable capability level (“Min-H”), since less capable AIs are less likely to be scheming and easier to oversee.
- Reversibility. If humans decide to reverse handoff, how easily could they do so?
- Transparency. Who knows that about handoff and its various dimensions?
The central questions include:
- How desirable is handoff in different scenarios, given the AI's capabilities and propensities, and the wider strategic landscape? In these scenarios, what happens after handoff?
- How can we make good decisions about whether, when, and how to hand off? What evidence would justify a handoff?
- How can we increase the chance of good handoff, and decrease the chance of bad handoff?
- What are viable alternatives to handoff?
See also:
- Types of Handoff to AIs (Kokotajlo, 2026)
- How do we (more) safely defer to AIs? (Greenblatt & Stastny, 2026)
- AI 2040: Plan A, Alignment Roadmap supplement, Phase 4.