Skip to content

Skill 02 — Model Lifecycle (the from-scratch sprint)

Purpose: destroy the black box. Not to produce a useful model. Timeboxed at 3 weeks. Do not overrun. The exit criterion is understanding, not results.

Duration 3 weeks (~40 hrs), hard cap
Impact ★★★☆☆ direct market value, ★★★★★ as the foundation for credibility
Prerequisites Skill 00, Skill 01 M1–M2
Hardware 4–6 GPUs for the main run (~8–12 hrs, run overnight); 1 GPU for all iteration work
Status planned

Outcome

  • You have trained a GPT-2-capability model from raw text, and talked to it
  • You can draw the full lifecycle — tokenizer → pretrain → midtrain → SFT → RL → eval → serve — and explain every box from having built it
  • You can read a production training codebase (torchtitan) and recognize every technique
  • You know, from experience, which stage causes which kind of model behaviour

The framing that matters

"Building our own custom LLM" in industry means continued pretraining + post-training on open weights, almost never pretraining from zero. This skill exists so that when you do post-training (Skill 05), you understand what you're modifying rather than turning knobs on a black box.


Modules

M Module Hrs Theme
M1 Tokenization 8 The layer everyone skips and then gets bitten by
M2 Pretraining 16 Where capability comes from
M3 The Post-Training Stages 10 Where behaviour comes from
M4 Production-Scale Training (read, don't master) 6 What the real thing looks like

M1 — Tokenization

ID Topic Key concepts Code artifact Visual Evidence
M1T1 BPE from first principles Merges, vocabulary construction, byte-level fallback, why not words and not characters BPE trainer in pure Python, ~150 lines Gen: merge process as growing units (Pattern C) Your own tokenizer trained on a corpus
M1T2 Tokenizer quality Compression ratio (bytes/token), fertility, language and code bias, digit and whitespace handling Compression-rate evaluator across text/code/multilingual Chart: bytes-per-token by domain, several tokenizers A comparison table for 4 real tokenizers
M1T3 Vocab size as a design decision Embedding params vs sequence length vs softmax cost; the tradeoff curve Vocab-size sweep on a toy model: params, tokens, loss Chart: vocab size vs total cost The compute-optimal vocab argument, reproduced
M1T4 Chat templates & special tokens Chat template rendering, BOS/EOS/pad, tool-call tokens, the silent template-mismatch bug Template renderer + a deliberate mismatch demo Mermaid: message list → token stream A reproduced template-mismatch failure and its fix

Why M1T4 matters: template mismatch between training and serving is the single most common cause of "my fine-tune got worse". Teaching it here, with a reproduction, pays off in Skill 05.


M2 — Pretraining

Spine: run karpathy/nanochat end to end, then modify it.

ID Topic Key concepts Code artifact Visual Evidence
M2T1 The pretraining data pipeline Corpus selection, shard formats, streaming dataloaders, packing, deterministic resumption Dataloader that shards, packs, and resumes correctly Mermaid: data flow Throughput in tokens/s, and a verified mid-epoch resume
M2T2 The training loop Forward/backward, gradient accumulation, mixed precision (bf16 on our hardware), grad clipping, checkpointing Annotated minimal training loop Mermaid: one step, annotated with memory events A toy model training with a clean loss curve
M2T3 Optimizers & schedules AdamW, Muon, optimizer state memory cost, warmup-stable-decay, weight decay, why LR is the master dial AdamW vs Muon on identical toy runs Chart: loss vs step and vs wall-clock, both optimizers A measured optimizer comparison
M2T4 Scaling laws & compute-optimal sizing Chinchilla, tokens-per-param, why nanochat derives everything from --depth, loss prediction Mini scaling-law sweep across depths Chart: loss vs compute, fitted power law Your own scaling curve from 4+ runs
M2T5 Measuring a pretraining run val_bpb (bits per byte, vocab-invariant), CORE score, MFU, tokens/s, VRAM — and what each tells you Metrics logger + W&B/Tensorboard dashboard Chart: the 4-panel run dashboard A run you can diagnose from its curves alone
M2T6 The full run Execute the complete speedrun on 4–6 GPUs; ~8–12 hrs, resumable runs/ config adapted to our hardware Gen: hero image for the milestone A model you trained, that you can chat with
M2T7 Intervene and observe Change one thing (attention variant, data mix, optimizer, depth) and measure the effect Experiment harness with config diffing Chart: baseline vs intervention A written experiment report: hypothesis → diff → result

Hardware note for authoring: the reference run is sized for 8×H100. Adapted configs and a short (<2 hr) variant must both be provided so the first pass isn't blocked overnight.


M3 — The Post-Training Stages

Here you meet each stage in its simplest possible form. Skill 05 does it properly.

ID Topic Key concepts Code artifact Visual Evidence
M3T1 Midtraining / continued pretraining Domain adaptation, capability injection, catastrophic forgetting, replay mixtures Continued-pretrain the base model on a new domain Gen: capability shift illustration Domain gain vs general-capability loss, measured
M3T2 SFT, minimally Chat formatting, assistant-only loss masking, why a tiny high-quality set beats a large noisy one chat_sft run + a loss-masking ablation Mermaid: loss mask over a token stream Before/after conversational behaviour, side by side
M3T3 RL, minimally Rollouts, rewards, advantage, policy update; GRPO in its simplest legible form chat_rl on a verifiable task (GSM8K-style) Gen: rollout-and-reward loop (Pattern C) Reward curve + task accuracy improvement
M3T4 Evaluating your own model CORE, MMLU/ARC/GSM8K harnesses, base-vs-chat evaluation differences, contamination Eval runner across all your checkpoints Chart: capability by stage (base → mid → SFT → RL) The stage-by-stage capability progression
M3T5 Inference on your own model KV cache implementation, sampling, tool execution — then read vLLM's version of the same thing engine.py walkthrough vs vLLM's paged implementation Mermaid: naive vs paged KV, side by side A written comparison of the two implementations

M3T5 is the payoff topic of the whole skill. Having written a naive KV cache, vLLM's design stops being magic.


M4 — Production-Scale Training (read, don't master)

Read and run tiny configs. Do not chase MFU numbers here.

ID Topic Key concepts Code artifact Visual Evidence
M4T1 torchtitan tour FSDP2 per-parameter sharding, meta-device init, selective activation checkpointing, distributed checkpointing Run a small config on 2–4 GPUs Mermaid: package/extension-point map A run completed and its memory profile read
M4T2 The parallelism dimensions FSDP/HSDP, TP + async TP, PP with zero-bubble schedules, CP for long context, EP for MoE — and how they compose Memory-estimation script across parallelism configs Gen: multi-dimensional sharding (Pattern A) A composition table: which dimension for which bottleneck
M4T3 Low-precision training bf16 today; float8/MXFP8 on newer silicon — conceptual, plus why it's not available to us Read components/quantization/float8.py; run the bf16 path Mermaid: where quantization sits in the step A written "what we'd gain on Hopper" analysis
M4T4 Reliability at scale Distributed + async checkpointing, TorchFT, flight recorder, NCCL hangs, straggler detection Kill a rank mid-run; recover from checkpoint Mermaid: failure and recovery flow A successful recovery from an induced failure

Labs

Lab Goal Success criterion
LAB-M1 Build a tokenizer Round-trips correctly; compression within 10% of a reference tokenizer
LAB-M2 The full run A chattable model you trained; run dashboard interpreted in writing
LAB-M3 One intervention A completed hypothesis→result experiment report, whether or not it worked
LAB-M4 Break and recover Induced rank failure recovered from checkpoint with no loss discontinuity

Capstone

Deliverable: lifecycle-teardown — a single document + repo that:

  1. Contains your trained model checkpoint and its full metric history
  2. Traces one input string through every stage: raw bytes → tokens → embeddings → transformer → logits → sampled token → detokenized output, with real tensor shapes at each step
  3. Documents your one intervention experiment with a real conclusion
  4. Compares your naive inference engine against vLLM's design, feature by feature
  5. States, in one page, exactly what a lab does that you didn't, and why it matters

Rubric: a reader with no prior LLM knowledge should be able to follow the trace in (2) and understand what a transformer does. If they can't, it isn't done.


Assessment

Tier Count Example
Recall 6 What is bits-per-byte and why is it preferred to raw loss for comparing models?
Apply 10 This loss curve plateaus then spikes at step 4k. Give three plausible causes and how to distinguish them
Design 4 You have a 200M-token domain corpus and an 8B base model. Design the training plan and justify each stage

Practical challenge: given a broken training config (deliberately introduced fault — bad LR, wrong template, broken loss mask, silent data corruption), diagnose and fix it from the curves. 2 hours.


Asset inventory

Type Count
Theory files 20
Code artifacts 18
Mermaid diagrams ~22
Generated images ~10
Charts ~18
Labs 4
Capstone 1

Primary sources

  • karpathy/nanochat — README, dev/LEADERBOARD.md, and the Discussions (esp. Beating GPT-2 for <$100 and the miniseries write-ups)
  • karpathy/nanoGPT + Let's build GPT video; KellerJordan/modded-nanogpt for the optimizer/speedrun ideas
  • pytorch/torchtitan + the TorchTitan ICLR 2025 paper; the PyTorch distributed blog series (FSDP2, async TP, zero-bubble PP, context parallel)
  • Chinchilla (Hoffmann et al.); DCLM and FineWeb data papers
  • Muon optimizer write-ups
  • Tulu 3 (Allen AI) — the most transparent published post-training recipe