Part III — Scaling Across Accelerators
What changes when computation and state cross device and machine boundaries: partitioning a model across ranks, routing tokens to experts, moving KV state between pools, keeping reusable prefixes alive beyond one accelerator, and the control plane that places work whose best location keeps changing.
Chapters 13–17