The activation checkpointing Pareto frontier

DRAFT (unlisted one-off study). Prose below is AI-drafted and awaits a human edit pass.1

Post 02 models activation checkpointing torch_remat-style: every op in the DeepSeek-V3 MoE block is either saved (its output is stashed for backward, if backward reads it) or recomputed (replayed once in backward). With 19 markable ops that is 219 policies; the chart keeps the two all-to-all outputs saved and plots the other 217. Each one lands at a point: forward FLOPs replayed, against bytes stashed. The chart plots all of them (grey), the ones no other policy beats on both axes (the Pareto frontier), and the frontier's lower convex hull. The hull points are the policies that are optimal for some exchange rate between FLOPs and memory.

The hull turns out to be nested: each step recomputes exactly what the previous one did, plus one more op (or a pair of ops that only pays off together). The chart numbers the steps, and the strip under it draws each one as a miniature of the layer diagram: the step's new recompute solid, the earlier ones pale. Hover any point for its policy; click it (or a frame) to load it into 02's diagram below, where recomputed ops are tinted. Toggling ops in the diagram moves the amber ring.

1 AI disclosure: the prose and the widget on this page were drafted by Claude and have not been edited by the author. Modeling notes: the costs are 02's. Recompute is counted in FLOPs, so vector ops (norms, RoPE, SwiGLU, adds) come out nearly free. In reality they are bandwidth-bound, so the vertical drop at the left edge is optimistic. The y axis uses the diagram's scaling: one MoE layer's bf16 stash × 8 layers × 8.5 microbatches in flight × 4,096 tokens. Replaying the a2a dispatch buys a second frontier parallel to this one, 112 KiB/token·layer lower, at the price of one more all-to-all per layer (a cost the x axis doesn't count). Replaying the a2a combine is never Pareto-optimal (both checked over all 219 markings). So the chart keeps both a2a outputs saved and leaves out the presets that replay them (Megatron's moe, and full). When several markings tie exactly, the one with the fewest recompute marks stands in for them. ↩