Skill 02 — Model Lifecycle (the from-scratch sprint)¶
Purpose: destroy the black box. Not to produce a useful model. Timeboxed at 3 weeks. Do not overrun. The exit criterion is understanding, not results.
| Duration | 3 weeks (~40 hrs), hard cap |
| Impact | ★★★☆☆ direct market value, ★★★★★ as the foundation for credibility |
| Prerequisites | Skill 00, Skill 01 M1–M2 |
| Hardware | 4–6 GPUs for the main run (~8–12 hrs, run overnight); 1 GPU for all iteration work |
| Status | planned |
Outcome¶
- You have trained a GPT-2-capability model from raw text, and talked to it
- You can draw the full lifecycle — tokenizer → pretrain → midtrain → SFT → RL → eval → serve — and explain every box from having built it
- You can read a production training codebase (torchtitan) and recognize every technique
- You know, from experience, which stage causes which kind of model behaviour
The framing that matters¶
"Building our own custom LLM" in industry means continued pretraining + post-training on open weights, almost never pretraining from zero. This skill exists so that when you do post-training (Skill 05), you understand what you're modifying rather than turning knobs on a black box.
Modules¶
| M | Module | Hrs | Theme |
|---|---|---|---|
| M1 | Tokenization | 8 | The layer everyone skips and then gets bitten by |
| M2 | Pretraining | 16 | Where capability comes from |
| M3 | The Post-Training Stages | 10 | Where behaviour comes from |
| M4 | Production-Scale Training (read, don't master) | 6 | What the real thing looks like |
M1 — Tokenization¶
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M1T1 | BPE from first principles | Merges, vocabulary construction, byte-level fallback, why not words and not characters | BPE trainer in pure Python, ~150 lines | Gen: merge process as growing units (Pattern C) | Your own tokenizer trained on a corpus |
| M1T2 | Tokenizer quality | Compression ratio (bytes/token), fertility, language and code bias, digit and whitespace handling | Compression-rate evaluator across text/code/multilingual | Chart: bytes-per-token by domain, several tokenizers | A comparison table for 4 real tokenizers |
| M1T3 | Vocab size as a design decision | Embedding params vs sequence length vs softmax cost; the tradeoff curve | Vocab-size sweep on a toy model: params, tokens, loss | Chart: vocab size vs total cost | The compute-optimal vocab argument, reproduced |
| M1T4 | Chat templates & special tokens | Chat template rendering, BOS/EOS/pad, tool-call tokens, the silent template-mismatch bug | Template renderer + a deliberate mismatch demo | Mermaid: message list → token stream | A reproduced template-mismatch failure and its fix |
Why M1T4 matters: template mismatch between training and serving is the single most common cause of "my fine-tune got worse". Teaching it here, with a reproduction, pays off in Skill 05.
M2 — Pretraining¶
Spine: run karpathy/nanochat end to end, then modify it.
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M2T1 | The pretraining data pipeline | Corpus selection, shard formats, streaming dataloaders, packing, deterministic resumption | Dataloader that shards, packs, and resumes correctly | Mermaid: data flow | Throughput in tokens/s, and a verified mid-epoch resume |
| M2T2 | The training loop | Forward/backward, gradient accumulation, mixed precision (bf16 on our hardware), grad clipping, checkpointing | Annotated minimal training loop | Mermaid: one step, annotated with memory events | A toy model training with a clean loss curve |
| M2T3 | Optimizers & schedules | AdamW, Muon, optimizer state memory cost, warmup-stable-decay, weight decay, why LR is the master dial | AdamW vs Muon on identical toy runs | Chart: loss vs step and vs wall-clock, both optimizers | A measured optimizer comparison |
| M2T4 | Scaling laws & compute-optimal sizing | Chinchilla, tokens-per-param, why nanochat derives everything from --depth, loss prediction |
Mini scaling-law sweep across depths | Chart: loss vs compute, fitted power law | Your own scaling curve from 4+ runs |
| M2T5 | Measuring a pretraining run | val_bpb (bits per byte, vocab-invariant), CORE score, MFU, tokens/s, VRAM — and what each tells you |
Metrics logger + W&B/Tensorboard dashboard | Chart: the 4-panel run dashboard | A run you can diagnose from its curves alone |
| M2T6 | The full run | Execute the complete speedrun on 4–6 GPUs; ~8–12 hrs, resumable | runs/ config adapted to our hardware |
Gen: hero image for the milestone | A model you trained, that you can chat with |
| M2T7 | Intervene and observe | Change one thing (attention variant, data mix, optimizer, depth) and measure the effect | Experiment harness with config diffing | Chart: baseline vs intervention | A written experiment report: hypothesis → diff → result |
Hardware note for authoring: the reference run is sized for 8×H100. Adapted configs and a short (<2 hr) variant must both be provided so the first pass isn't blocked overnight.
M3 — The Post-Training Stages¶
Here you meet each stage in its simplest possible form. Skill 05 does it properly.
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M3T1 | Midtraining / continued pretraining | Domain adaptation, capability injection, catastrophic forgetting, replay mixtures | Continued-pretrain the base model on a new domain | Gen: capability shift illustration | Domain gain vs general-capability loss, measured |
| M3T2 | SFT, minimally | Chat formatting, assistant-only loss masking, why a tiny high-quality set beats a large noisy one | chat_sft run + a loss-masking ablation |
Mermaid: loss mask over a token stream | Before/after conversational behaviour, side by side |
| M3T3 | RL, minimally | Rollouts, rewards, advantage, policy update; GRPO in its simplest legible form | chat_rl on a verifiable task (GSM8K-style) |
Gen: rollout-and-reward loop (Pattern C) | Reward curve + task accuracy improvement |
| M3T4 | Evaluating your own model | CORE, MMLU/ARC/GSM8K harnesses, base-vs-chat evaluation differences, contamination | Eval runner across all your checkpoints | Chart: capability by stage (base → mid → SFT → RL) | The stage-by-stage capability progression |
| M3T5 | Inference on your own model | KV cache implementation, sampling, tool execution — then read vLLM's version of the same thing | engine.py walkthrough vs vLLM's paged implementation |
Mermaid: naive vs paged KV, side by side | A written comparison of the two implementations |
M3T5 is the payoff topic of the whole skill. Having written a naive KV cache, vLLM's design stops being magic.
M4 — Production-Scale Training (read, don't master)¶
Read and run tiny configs. Do not chase MFU numbers here.
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M4T1 | torchtitan tour | FSDP2 per-parameter sharding, meta-device init, selective activation checkpointing, distributed checkpointing | Run a small config on 2–4 GPUs | Mermaid: package/extension-point map | A run completed and its memory profile read |
| M4T2 | The parallelism dimensions | FSDP/HSDP, TP + async TP, PP with zero-bubble schedules, CP for long context, EP for MoE — and how they compose | Memory-estimation script across parallelism configs | Gen: multi-dimensional sharding (Pattern A) | A composition table: which dimension for which bottleneck |
| M4T3 | Low-precision training | bf16 today; float8/MXFP8 on newer silicon — conceptual, plus why it's not available to us | Read components/quantization/float8.py; run the bf16 path |
Mermaid: where quantization sits in the step | A written "what we'd gain on Hopper" analysis |
| M4T4 | Reliability at scale | Distributed + async checkpointing, TorchFT, flight recorder, NCCL hangs, straggler detection | Kill a rank mid-run; recover from checkpoint | Mermaid: failure and recovery flow | A successful recovery from an induced failure |
Labs¶
| Lab | Goal | Success criterion |
|---|---|---|
| LAB-M1 | Build a tokenizer | Round-trips correctly; compression within 10% of a reference tokenizer |
| LAB-M2 | The full run | A chattable model you trained; run dashboard interpreted in writing |
| LAB-M3 | One intervention | A completed hypothesis→result experiment report, whether or not it worked |
| LAB-M4 | Break and recover | Induced rank failure recovered from checkpoint with no loss discontinuity |
Capstone¶
Deliverable: lifecycle-teardown — a single document + repo that:
- Contains your trained model checkpoint and its full metric history
- Traces one input string through every stage: raw bytes → tokens → embeddings → transformer → logits → sampled token → detokenized output, with real tensor shapes at each step
- Documents your one intervention experiment with a real conclusion
- Compares your naive inference engine against vLLM's design, feature by feature
- States, in one page, exactly what a lab does that you didn't, and why it matters
Rubric: a reader with no prior LLM knowledge should be able to follow the trace in (2) and understand what a transformer does. If they can't, it isn't done.
Assessment¶
| Tier | Count | Example |
|---|---|---|
| Recall | 6 | What is bits-per-byte and why is it preferred to raw loss for comparing models? |
| Apply | 10 | This loss curve plateaus then spikes at step 4k. Give three plausible causes and how to distinguish them |
| Design | 4 | You have a 200M-token domain corpus and an 8B base model. Design the training plan and justify each stage |
Practical challenge: given a broken training config (deliberately introduced fault — bad LR, wrong template, broken loss mask, silent data corruption), diagnose and fix it from the curves. 2 hours.
Asset inventory¶
| Type | Count |
|---|---|
| Theory files | 20 |
| Code artifacts | 18 |
| Mermaid diagrams | ~22 |
| Generated images | ~10 |
| Charts | ~18 |
| Labs | 4 |
| Capstone | 1 |
Primary sources¶
karpathy/nanochat— README,dev/LEADERBOARD.md, and the Discussions (esp. Beating GPT-2 for <$100 and the miniseries write-ups)karpathy/nanoGPT+ Let's build GPT video;KellerJordan/modded-nanogptfor the optimizer/speedrun ideaspytorch/torchtitan+ the TorchTitan ICLR 2025 paper; the PyTorch distributed blog series (FSDP2, async TP, zero-bubble PP, context parallel)- Chinchilla (Hoffmann et al.); DCLM and FineWeb data papers
- Muon optimizer write-ups
- Tulu 3 (Allen AI) — the most transparent published post-training recipe