LAB-M2 — Build and Audit a Generation Loop¶
Goal: Build a small decoder generation loop with a correct KV cache, then explain its parameter, memory, phase, position, and routing behavior from evidence.
Time: 180 min Hardware: CPU for the toy path; 1 CUDA GPU for checkpoint verification
Success criteria¶
- [ ] Full causal attention and cached token-by-token attention agree within
1e-5. - [ ] Measured KV bytes exactly match
2 × layers × batch × tokens × kv_heads × head_dim × element_bytesfor the toy model. - [ ] The learner predicts parameter ownership within 1% for all tensors intentionally included.
- [ ] Prefill and decode are timed separately and reported with their matrix shapes; CPU results remain labeled diagnostic.
- [ ] RoPE positions advance correctly through cache growth and pass a shift-invariance test.
- [ ] If MoE is enabled, total and active parameters plus maximum-to-mean expert load are reported separately.
- [ ] On the checkpoint path, greedy output matches the reference generation token-for-token for at least 32 new tokens.
Setup¶
Synchronize the skill environment from 00-foundations/ with uv sync. Run all six M2 artifact runners once. Select a small decoder checkpoint only for the optional integration path; keep the mechanism path library-free apart from PyTorch.
Steps¶
-
Draw the decoder block and annotate every tensor shape. Use the M2T1 counter to predict parameters before constructing modules.
-
Implement one causal attention layer without a cache. Confirm the causal mask by changing a future token and proving earlier outputs remain unchanged.
-
Add KV caching. Take the tempting wrong path first: append K and V but reset RoPE position to zero at every decode step. Observe that shapes and cache bytes look correct while outputs diverge. Fix positions using absolute cache length.
-
Compare full-sequence and cached outputs after every token. Stop at the first mismatch rather than comparing only the final vector.
-
Sweep MHA, valid GQA groupings, and MQA in the cache calculator. Do not change an already initialized checkpoint’s head count; these are architecture variants for the toy model.
-
Instrument prefill separately from 32 decode steps. Record rows, FLOPs, minimum bytes, elapsed time, and whether the result came from CPU or CUDA.
-
Run the RoPE shift test: move query and key positions by the same offset while preserving their relative distance. Their rotated dot product should remain close within floating-point tolerance.
-
Optional MoE extension: replace the dense MLP with top-2 experts. Record assignment counts and show that per-token active parameters do not equal stored parameters.
-
Checkpoint path: map the same concepts onto a small supported model and compare greedy generation against the reference implementation. Match tokenizer, prompt, cache positions, logits, and stopping conditions.
Stretch¶
- Add a paged cache allocator with fixed-size blocks and measure internal waste for variable sequence lengths.
- Quantize cached K/V values to INT8 with explicit scales and measure output error as context grows.
- Batch sequences with different completion lengths and implement slot replacement when one finishes.
Debrief¶
The wrong-position path is important because it creates a realistic failure: allocation and tensor shapes pass while model behavior fails. Systems correctness is not only “the kernel ran.” Position, cache ordering, masks, and head grouping are part of the serving contract.
Your final report should connect five views of the same loop: tensor shapes, parameter ownership, cache bytes, phase timing, and output equivalence. If any number cannot be derived before measurement, return to the corresponding M2 topic before proceeding.