Skip to content

Skill 05 — Post-Training

This is what "building our own custom LLM" actually means in industry. Not pretraining from zero — taking open weights and bending them to your task with SFT, preference optimization, and RL on verifiable rewards.

Duration 5 weeks (~65 hrs)
Impact ★★★★☆ — and the RL half is the fastest-moving frontier in the stack
Prerequisites Skill 02 (you must have built a model once), Skill 04 (you need a real dataset), Skill 01 M6 (multi-LoRA serving)
Hardware LoRA on 1 GPU · full-FT and QLoRA-70B on 4 · GRPO needs 2 GPUs for generation + 2–4 for training simultaneously
Status planned

Outcome

  • You can take a business task and produce a measurably better model than the base, and prove it
  • You can choose between prompting, RAG, SFT, preference tuning, RL, and distillation — and say why
  • You can run GRPO with a verifiable reward and recognize reward hacking when it happens
  • You can serve 20 tenant-specific adapters off one base model and know exactly what it costs

The decision that comes first

Most "we need to fine-tune" requests shouldn't be fine-tunes. Teach this before teaching any technique:

flowchart TD
    A[Model isn't good enough] --> B{What kind of gap?}
    B -->|Missing facts| C[RAG. Not fine-tuning.]
    B -->|Wrong format or style| D[Prompt, then SFT if it must be reliable]
    B -->|Missing domain vocabulary| E[Continued pretraining, then SFT]
    B -->|Right answers exist but model won't prefer them| F[Preference tuning: DPO/KTO]
    B -->|Answer is checkable by a program| G[RL with verifiable rewards: GRPO]
    B -->|Too slow or expensive| H[Distillation + quantization. Skill 03.]

Modules

M Module Hrs Theme
M1 Parameter-Efficient Fine-Tuning 14 LoRA and what it actually does
M2 Supervised Fine-Tuning at Quality 14 Making SFT reliably work
M3 Preference Optimization 12 Teaching taste without a reward model
M4 Reinforcement Learning 17 GRPO, RLVR, and the training/serving merge
M5 Serving What You Trained 8 Multi-LoRA, merging, and the deployment loop

M1 — Parameter-Efficient Fine-Tuning

ID Topic Key concepts Code artifact Visual Evidence
M1T1 LoRA from first principles Low-rank decomposition ΔW = BA, why gradients concentrate in a low-rank subspace, rank vs capacity LoRA implemented from scratch on a small model, ~80 lines Gen sequence: full matrix → decomposition → adapter injection Your LoRA matching PEFT's output on the same seed
M1T2 The knobs that matter Rank, alpha (and the alpha/r scaling), dropout, which modules to target, why attention-only is usually wrong Ablation grid: rank × target-modules Chart: eval score vs rank vs targets The rank at which returns flatten, on your task
M1T3 QLoRA NF4 base quantization, double quantization, paged optimizers, the accuracy/memory tradeoff 70B QLoRA fine-tune on 4 GPUs Gen: frozen quantized base + live adapter (Pattern B) 70B fine-tuned on hardware that can't hold it in bf16
M1T4 The LoRA variants DoRA, rsLoRA, LoRA+, PiSSA — what each fixes, and whether it's worth the complexity Head-to-head on one task Chart: variant vs eval vs training time A recommendation with data, not vibes
M1T5 When full fine-tuning wins Large domain shift, big datasets, capability injection; the memory arithmetic of full-FT Full-FT 8B with FSDP2 on 4 GPUs Chart: LoRA vs full-FT, eval and cost The crossover point on your data

M2 — Supervised Fine-Tuning at Quality

ID Topic Key concepts Code artifact Visual Evidence
M2T1 The frameworks TRL SFTTrainer (canonical), Axolotl (YAML, broadest features), Unsloth (fastest single-GPU), torchtune — and how to pick Same fine-tune in all three Mermaid: framework capability map Identical results across frameworks, or an explained discrepancy
M2T2 Configuring a real run LR and schedule, epochs vs overfitting, batch size and grad accumulation, packing, eval cadence, checkpoint selection Annotated production-grade config Chart: the four curves you must watch A run you can diagnose from its curves
M2T3 Memory and speed Gradient checkpointing, Flash Attention 2, Liger kernels, cut cross-entropy, 8-bit optimizers — what each buys Ablation: memory and step-time per optimization Chart: VRAM and tokens/s per technique Your own optimization stack, measured
M2T4 Distributed SFT DeepSpeed ZeRO 1/2/3 vs FSDP2, sharding strategies, sequence parallelism for long context Same job under ZeRO-3 and FSDP2 Gen: what gets sharded at each ZeRO stage (progressive sequence) Scaling efficiency 1→2→4 GPUs
M2T5 Catastrophic forgetting Measuring general-capability loss, replay mixtures, LR sensitivity, adapter isolation as mitigation Forgetting measurement harness across checkpoints Chart: target gain vs general loss, per config The forgetting curve, and your chosen mitigation
M2T6 Debugging a bad fine-tune Loss won't drop / drops too fast / diverges; template mismatch; broken masks; data leakage; bad LR Fault-injection suite with five deliberately broken runs Mermaid: symptom → cause → fix decision tree All five faults diagnosed from curves alone

M2T6 is the highest-value topic in this module — it's what people actually spend their time on.


M3 — Preference Optimization

ID Topic Key concepts Code artifact Visual Evidence
M3T1 Why SFT isn't enough SFT teaches "a good answer", preference teaches "better than that one"; the ranking signal Failure demo: an SFT model that can't discriminate quality Gen: single-target vs relative-preference (Pattern B) A capability gap SFT provably can't close
M3T2 DPO The implicit reward, the reference model, beta, why no reward model is needed; the derivation, briefly DPO on your Skill 04 preference set Gen sequence: RLHF pipeline → DPO's collapse of it Win-rate improvement vs the SFT checkpoint
M3T3 The DPO family KTO (unpaired, binary), ORPO (no reference model, merged with SFT), CPO, SimPO — data requirements drive the choice Same task, three methods Chart: method vs win rate vs data cost A selection rule based on what data you can get
M3T4 Failure modes Reward over-optimization, verbosity bias, degeneration, the KL leash, the length-exploitation problem Deliberate over-optimization run Chart: win rate vs KL divergence, with the collapse point A model you broke by pushing beta too far
M3T5 Reward modelling Outcome RM vs Process RM (PRM), Bradley-Terry, calibration, when you need an explicit RM at all Train an RM; measure agreement with held-out human labels Mermaid: RM training and use RM accuracy + calibration plot

M4 — Reinforcement Learning

The frontier. Also where training infrastructure and serving infrastructure merge — which is why this skill sits after Skill 01.

ID Topic Key concepts Code artifact Visual Evidence
M4T1 The RL loop for LLMs Policy, rollout, reward, advantage, update; why sequence-level RL is different from classic RL Minimal RL loop, readable, ~150 lines Gen sequence: rollout → score → advantage → update (Pattern C, one image per stage) A working loop on a toy verifiable task
M4T2 PPO and why it was replaced Value network cost, clipping, GAE; the memory bill of four models in memory at once PPO run for comparison Gen: four-model memory footprint Measured cost of PPO vs GRPO on identical hardware
M4T3 GRPO Group-relative advantage with no value network, group size, the DeepSeek-R1 result, why it won GRPO on a math/code task Gen: group of rollouts scored relative to each other Task accuracy lift + the reward curve
M4T4 RLVR — verifiable rewards Programmatic verifiers (unit tests, math checkers, schema validators), why verifiable > learned reward Verifier suite + GRPO against it Mermaid: task → rollout → verify → reward Accuracy improvement on a genuinely held-out set
M4T5 Reward hacking Degenerate strategies, length gaming, verifier exploitation, detection and mitigation Induce a reward hack, then defend against it Gen: intended vs exploited path (Pattern B) A documented hack you caused and then fixed
M4T6 The generation/training co-location problem vLLM inside the training loop, weight synchronization, co-located vs disaggregated RL, GPU allocation GRPO with co-located vLLM on our GPU budget Mermaid: weight sync sequence + Gen: GPU allocation split Throughput comparison: co-located vs sequential
M4T7 Scaling out verl, OpenRLHF, NeMo-RL, TRL at scale — architectures and when you'd need them Read + run one small verl config Mermaid: framework architecture comparison A written framework selection rationale
M4T8 Agentic RL & environments Multi-turn credit assignment, tool-use RL, OpenEnv/Harbor-style environments with environment-owned rewards, sandboxing Build one environment with a reward function Gen: agent-environment loop over multiple turns A trained tool-using policy beating the base model

M4T8 is the forward bet. It's where the field is heading and where the least competition exists.


M5 — Serving What You Trained

ID Topic Key concepts Code artifact Visual Evidence
M5T1 Merge or serve separately Merging LoRA into base weights, quality of merged vs adapter, when merging is wrong Merge, then compare merged vs adapter-served Gen: merge operation (Pattern B) Quality and latency delta, merged vs adapter
M5T2 Multi-LoRA in production One base + N adapters, dynamic loading, per-adapter latency overhead, tenant isolation Serve 5+ adapters on one vLLM instance Gen sequence: N deployments → one base + N adapters Cost per tenant vs separate deployments
M5T3 Model merging techniques TIES, DARE, SLERP, task arithmetic — combining capabilities without retraining Merge two task-specific models; evaluate both tasks Gen: weight-space interpolation Whether merging preserved both capabilities
M5T4 The post-training release loop Versioning, eval gating, canary, rollback, registry, reproducibility manifests Release pipeline script with an eval gate Mermaid: train → eval → gate → canary → promote A model promoted through a real gate

Labs

Lab Goal Success criterion
LAB-M1 LoRA by hand Hand-rolled LoRA matching PEFT numerically; rank ablation with a defended choice
LAB-M2 Five broken runs Diagnose all five injected faults from curves and configs alone
LAB-M3 Push DPO until it breaks Induce reward over-optimization, then find the beta that holds
LAB-M4 GRPO on a verifiable task ≥10 point accuracy gain on held-out, with no detectable reward hacking
LAB-M5 Ten tenants, one GPU 10 adapters served from one base, per-tenant cost measured

Capstone

Deliverable: custom-model-end-to-end — the flagship artifact of the entire curriculum.

Pick one real task you care about. Then:

  1. Justify the approach using the decision tree above — including why RAG/prompting wasn't enough
  2. Dataset from Skill 04, with its datasheet and contamination audit
  3. SFT with a defended configuration and an ablation supporting it
  4. Preference tuning pass with measured win-rate improvement
  5. GRPO with a verifiable reward and a documented reward-hacking check
  6. Compress it (Skill 03) with an accuracy delta report
  7. Serve it as one adapter among several on shared infrastructure
  8. Evaluate against base, against a frontier API model, and against a strong prompted baseline
  9. Cost report: total training cost + $/1M tokens served
  10. Honest limitations section — where it's worse than the base model

Rubric: could someone reproduce every number from your repo? Does the eval prove the model is better on the task, not just on the training distribution? Is the limitations section brave?


Assessment

Tier Count Example
Recall 7 What does beta control in DPO, and what happens at both extremes?
Apply 13 GRPO reward is climbing but held-out accuracy is flat. Give three causes and how to distinguish them
Design 5 200 labelled examples, a code-generation task, 4×A100, two weeks. Design the full post-training plan

Practical challenge: given a task, a base model, and 6 hours of GPU time, produce a measurably improved model with the evidence to prove it.


Asset inventory

Type Count (minimum)
Theory articles 28
Code artifacts 26
Mermaid diagrams ~25
Generated images ~110
Charts ~35
Labs 5
Capstone 1

Primary sources

  • HF TRL docs — the trainer taxonomy (SFT, DPO, KTO, GRPO, RLOO, PRM, GKD) and the "reducing memory usage" / "speeding up training" how-tos
  • Axolotl docs — config reference, axolotl agent-docs grpo|sft|preference_tuning, multi-GPU and multi-node guides
  • Unsloth docs; pytorch/torchtune
  • Papers: LoRA, QLoRA, DoRA, DPO, KTO, ORPO, SimPO, PPO, DeepSeek-R1 (GRPO), RLOO, Tulu 3
  • Open-R1 — the open reproduction of R1, with full recipes
  • volcengine/verl docs; OpenRLHF
  • TRL blog posts: NO GPU left behind: co-located vLLM in TRL, Liger GRPO meets TRL, OpenEnv
  • Model merging: TIES, DARE, mergekit