Notes on ML infrastructure, checkpointing, and how transformers actually spend their compute.
A journey through the history of optimization, from Gauss's lost planet to the future of AI.
A click-through annotated walkthrough of transformer FLOPs and memory: where C ≈ 6ND comes from, why training costs 3× a forward pass, and when attention's quadratic term actually starts to matter.