Skill 05 — Post-Training
This is what "building our own custom LLM" actually means in industry.
Not pretraining from zero — taking open weights and bending them to your task with SFT, preference
optimization, and RL on verifiable rewards.
|
|
| Duration |
5 weeks (~65 hrs) |
| Impact |
★★★★☆ — and the RL half is the fastest-moving frontier in the stack |
| Prerequisites |
Skill 02 (you must have built a model once), Skill 04 (you need a real dataset), Skill 01 M6 (multi-LoRA serving) |
| Hardware |
LoRA on 1 GPU · full-FT and QLoRA-70B on 4 · GRPO needs 2 GPUs for generation + 2–4 for training simultaneously |
| Status |
planned |
Outcome
- You can take a business task and produce a measurably better model than the base, and prove it
- You can choose between prompting, RAG, SFT, preference tuning, RL, and distillation — and say why
- You can run GRPO with a verifiable reward and recognize reward hacking when it happens
- You can serve 20 tenant-specific adapters off one base model and know exactly what it costs
The decision that comes first
Most "we need to fine-tune" requests shouldn't be fine-tunes. Teach this before teaching any technique:
flowchart TD
A[Model isn't good enough] --> B{What kind of gap?}
B -->|Missing facts| C[RAG. Not fine-tuning.]
B -->|Wrong format or style| D[Prompt, then SFT if it must be reliable]
B -->|Missing domain vocabulary| E[Continued pretraining, then SFT]
B -->|Right answers exist but model won't prefer them| F[Preference tuning: DPO/KTO]
B -->|Answer is checkable by a program| G[RL with verifiable rewards: GRPO]
B -->|Too slow or expensive| H[Distillation + quantization. Skill 03.]
Modules
| M |
Module |
Hrs |
Theme |
| M1 |
Parameter-Efficient Fine-Tuning |
14 |
LoRA and what it actually does |
| M2 |
Supervised Fine-Tuning at Quality |
14 |
Making SFT reliably work |
| M3 |
Preference Optimization |
12 |
Teaching taste without a reward model |
| M4 |
Reinforcement Learning |
17 |
GRPO, RLVR, and the training/serving merge |
| M5 |
Serving What You Trained |
8 |
Multi-LoRA, merging, and the deployment loop |
M1 — Parameter-Efficient Fine-Tuning
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M1T1 |
LoRA from first principles |
Low-rank decomposition ΔW = BA, why gradients concentrate in a low-rank subspace, rank vs capacity |
LoRA implemented from scratch on a small model, ~80 lines |
Gen sequence: full matrix → decomposition → adapter injection |
Your LoRA matching PEFT's output on the same seed |
| M1T2 |
The knobs that matter |
Rank, alpha (and the alpha/r scaling), dropout, which modules to target, why attention-only is usually wrong |
Ablation grid: rank × target-modules |
Chart: eval score vs rank vs targets |
The rank at which returns flatten, on your task |
| M1T3 |
QLoRA |
NF4 base quantization, double quantization, paged optimizers, the accuracy/memory tradeoff |
70B QLoRA fine-tune on 4 GPUs |
Gen: frozen quantized base + live adapter (Pattern B) |
70B fine-tuned on hardware that can't hold it in bf16 |
| M1T4 |
The LoRA variants |
DoRA, rsLoRA, LoRA+, PiSSA — what each fixes, and whether it's worth the complexity |
Head-to-head on one task |
Chart: variant vs eval vs training time |
A recommendation with data, not vibes |
| M1T5 |
When full fine-tuning wins |
Large domain shift, big datasets, capability injection; the memory arithmetic of full-FT |
Full-FT 8B with FSDP2 on 4 GPUs |
Chart: LoRA vs full-FT, eval and cost |
The crossover point on your data |
M2 — Supervised Fine-Tuning at Quality
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M2T1 |
The frameworks |
TRL SFTTrainer (canonical), Axolotl (YAML, broadest features), Unsloth (fastest single-GPU), torchtune — and how to pick |
Same fine-tune in all three |
Mermaid: framework capability map |
Identical results across frameworks, or an explained discrepancy |
| M2T2 |
Configuring a real run |
LR and schedule, epochs vs overfitting, batch size and grad accumulation, packing, eval cadence, checkpoint selection |
Annotated production-grade config |
Chart: the four curves you must watch |
A run you can diagnose from its curves |
| M2T3 |
Memory and speed |
Gradient checkpointing, Flash Attention 2, Liger kernels, cut cross-entropy, 8-bit optimizers — what each buys |
Ablation: memory and step-time per optimization |
Chart: VRAM and tokens/s per technique |
Your own optimization stack, measured |
| M2T4 |
Distributed SFT |
DeepSpeed ZeRO 1/2/3 vs FSDP2, sharding strategies, sequence parallelism for long context |
Same job under ZeRO-3 and FSDP2 |
Gen: what gets sharded at each ZeRO stage (progressive sequence) |
Scaling efficiency 1→2→4 GPUs |
| M2T5 |
Catastrophic forgetting |
Measuring general-capability loss, replay mixtures, LR sensitivity, adapter isolation as mitigation |
Forgetting measurement harness across checkpoints |
Chart: target gain vs general loss, per config |
The forgetting curve, and your chosen mitigation |
| M2T6 |
Debugging a bad fine-tune |
Loss won't drop / drops too fast / diverges; template mismatch; broken masks; data leakage; bad LR |
Fault-injection suite with five deliberately broken runs |
Mermaid: symptom → cause → fix decision tree |
All five faults diagnosed from curves alone |
M2T6 is the highest-value topic in this module — it's what people actually spend their time on.
M3 — Preference Optimization
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M3T1 |
Why SFT isn't enough |
SFT teaches "a good answer", preference teaches "better than that one"; the ranking signal |
Failure demo: an SFT model that can't discriminate quality |
Gen: single-target vs relative-preference (Pattern B) |
A capability gap SFT provably can't close |
| M3T2 |
DPO |
The implicit reward, the reference model, beta, why no reward model is needed; the derivation, briefly |
DPO on your Skill 04 preference set |
Gen sequence: RLHF pipeline → DPO's collapse of it |
Win-rate improvement vs the SFT checkpoint |
| M3T3 |
The DPO family |
KTO (unpaired, binary), ORPO (no reference model, merged with SFT), CPO, SimPO — data requirements drive the choice |
Same task, three methods |
Chart: method vs win rate vs data cost |
A selection rule based on what data you can get |
| M3T4 |
Failure modes |
Reward over-optimization, verbosity bias, degeneration, the KL leash, the length-exploitation problem |
Deliberate over-optimization run |
Chart: win rate vs KL divergence, with the collapse point |
A model you broke by pushing beta too far |
| M3T5 |
Reward modelling |
Outcome RM vs Process RM (PRM), Bradley-Terry, calibration, when you need an explicit RM at all |
Train an RM; measure agreement with held-out human labels |
Mermaid: RM training and use |
RM accuracy + calibration plot |
M4 — Reinforcement Learning
The frontier. Also where training infrastructure and serving infrastructure merge — which is why this
skill sits after Skill 01.
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M4T1 |
The RL loop for LLMs |
Policy, rollout, reward, advantage, update; why sequence-level RL is different from classic RL |
Minimal RL loop, readable, ~150 lines |
Gen sequence: rollout → score → advantage → update (Pattern C, one image per stage) |
A working loop on a toy verifiable task |
| M4T2 |
PPO and why it was replaced |
Value network cost, clipping, GAE; the memory bill of four models in memory at once |
PPO run for comparison |
Gen: four-model memory footprint |
Measured cost of PPO vs GRPO on identical hardware |
| M4T3 |
GRPO |
Group-relative advantage with no value network, group size, the DeepSeek-R1 result, why it won |
GRPO on a math/code task |
Gen: group of rollouts scored relative to each other |
Task accuracy lift + the reward curve |
| M4T4 |
RLVR — verifiable rewards |
Programmatic verifiers (unit tests, math checkers, schema validators), why verifiable > learned reward |
Verifier suite + GRPO against it |
Mermaid: task → rollout → verify → reward |
Accuracy improvement on a genuinely held-out set |
| M4T5 |
Reward hacking |
Degenerate strategies, length gaming, verifier exploitation, detection and mitigation |
Induce a reward hack, then defend against it |
Gen: intended vs exploited path (Pattern B) |
A documented hack you caused and then fixed |
| M4T6 |
The generation/training co-location problem |
vLLM inside the training loop, weight synchronization, co-located vs disaggregated RL, GPU allocation |
GRPO with co-located vLLM on our GPU budget |
Mermaid: weight sync sequence + Gen: GPU allocation split |
Throughput comparison: co-located vs sequential |
| M4T7 |
Scaling out |
verl, OpenRLHF, NeMo-RL, TRL at scale — architectures and when you'd need them |
Read + run one small verl config |
Mermaid: framework architecture comparison |
A written framework selection rationale |
| M4T8 |
Agentic RL & environments |
Multi-turn credit assignment, tool-use RL, OpenEnv/Harbor-style environments with environment-owned rewards, sandboxing |
Build one environment with a reward function |
Gen: agent-environment loop over multiple turns |
A trained tool-using policy beating the base model |
M4T8 is the forward bet. It's where the field is heading and where the least competition exists.
M5 — Serving What You Trained
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M5T1 |
Merge or serve separately |
Merging LoRA into base weights, quality of merged vs adapter, when merging is wrong |
Merge, then compare merged vs adapter-served |
Gen: merge operation (Pattern B) |
Quality and latency delta, merged vs adapter |
| M5T2 |
Multi-LoRA in production |
One base + N adapters, dynamic loading, per-adapter latency overhead, tenant isolation |
Serve 5+ adapters on one vLLM instance |
Gen sequence: N deployments → one base + N adapters |
Cost per tenant vs separate deployments |
| M5T3 |
Model merging techniques |
TIES, DARE, SLERP, task arithmetic — combining capabilities without retraining |
Merge two task-specific models; evaluate both tasks |
Gen: weight-space interpolation |
Whether merging preserved both capabilities |
| M5T4 |
The post-training release loop |
Versioning, eval gating, canary, rollback, registry, reproducibility manifests |
Release pipeline script with an eval gate |
Mermaid: train → eval → gate → canary → promote |
A model promoted through a real gate |
Labs
| Lab |
Goal |
Success criterion |
| LAB-M1 |
LoRA by hand |
Hand-rolled LoRA matching PEFT numerically; rank ablation with a defended choice |
| LAB-M2 |
Five broken runs |
Diagnose all five injected faults from curves and configs alone |
| LAB-M3 |
Push DPO until it breaks |
Induce reward over-optimization, then find the beta that holds |
| LAB-M4 |
GRPO on a verifiable task |
≥10 point accuracy gain on held-out, with no detectable reward hacking |
| LAB-M5 |
Ten tenants, one GPU |
10 adapters served from one base, per-tenant cost measured |
Capstone
Deliverable: custom-model-end-to-end — the flagship artifact of the entire curriculum.
Pick one real task you care about. Then:
- Justify the approach using the decision tree above — including why RAG/prompting wasn't enough
- Dataset from Skill 04, with its datasheet and contamination audit
- SFT with a defended configuration and an ablation supporting it
- Preference tuning pass with measured win-rate improvement
- GRPO with a verifiable reward and a documented reward-hacking check
- Compress it (Skill 03) with an accuracy delta report
- Serve it as one adapter among several on shared infrastructure
- Evaluate against base, against a frontier API model, and against a strong prompted baseline
- Cost report: total training cost + $/1M tokens served
- Honest limitations section — where it's worse than the base model
Rubric: could someone reproduce every number from your repo? Does the eval prove the model is
better on the task, not just on the training distribution? Is the limitations section brave?
Assessment
| Tier |
Count |
Example |
| Recall |
7 |
What does beta control in DPO, and what happens at both extremes? |
| Apply |
13 |
GRPO reward is climbing but held-out accuracy is flat. Give three causes and how to distinguish them |
| Design |
5 |
200 labelled examples, a code-generation task, 4×A100, two weeks. Design the full post-training plan |
Practical challenge: given a task, a base model, and 6 hours of GPU time, produce a measurably
improved model with the evidence to prove it.
Asset inventory
| Type |
Count (minimum) |
| Theory articles |
28 |
| Code artifacts |
26 |
| Mermaid diagrams |
~25 |
| Generated images |
~110 |
| Charts |
~35 |
| Labs |
5 |
| Capstone |
1 |
Primary sources
- HF TRL docs — the trainer taxonomy (SFT, DPO, KTO, GRPO, RLOO, PRM, GKD) and the "reducing memory usage" / "speeding up training" how-tos
- Axolotl docs — config reference,
axolotl agent-docs grpo|sft|preference_tuning, multi-GPU and multi-node guides
- Unsloth docs;
pytorch/torchtune
- Papers: LoRA, QLoRA, DoRA, DPO, KTO, ORPO, SimPO, PPO, DeepSeek-R1 (GRPO), RLOO, Tulu 3
- Open-R1 — the open reproduction of R1, with full recipes
volcengine/verl docs; OpenRLHF
- TRL blog posts: NO GPU left behind: co-located vLLM in TRL, Liger GRPO meets TRL, OpenEnv
- Model merging: TIES, DARE,
mergekit