Large-scale cluster management at Google with Borg
Borg asks how one fleet can run latency-sensitive services and opportunistic batch work without wasting the space between their peaks. Read it to see how cells, jobs, tasks, priorities, limits, reservations, placement scoring, and a continuously repaired control plane turn heterogeneous machines into shared infrastructure.
Reading focus: How Borg separates a replicated control plane from per-machine Borglets and continuously reconciles desired jobs with observed task state. Why scheduling first filters feasible machines and then scores them, while priority, limits, and reservations make mixed workloads practical. What Google's measurements do and do not show about packing, reclaimed capacity, availability, isolation, and the cost of segregating workloads.
EuroSys 2015. Verma et al.. 40 min read, easy difficulty.