Skill 03 — Compression & Optimization¶
The highest ROI-per-hour skill after serving. This is the one that converts directly into "we now need half the GPUs" — the single most requested outcome from anyone paying for inference.
| Duration | 3 weeks (~40 hrs) |
| Impact | ★★★★☆ |
| Prerequisites | Skill 00, Skill 01 (esp. M3 benchmarking) |
| Hardware | 1–2 GPUs for quantization; 2–4 for the 70B stress lab |
| Status | planned |
Outcome¶
- You can quantize any model and prove the accuracy cost, not guess at it
- You can pick a scheme from hardware + workload + accuracy budget, and defend it
- You understand which techniques are latency wins, which are throughput wins, and which are neither
- You never ship a compressed model without a delta report — and you can explain why anyone who does is negligent
The methodology this skill really teaches¶
Quantization algorithms churn. The method doesn't:
flowchart LR
A[Define accuracy budget<br/>on task-relevant evals] --> B[Baseline: measure<br/>accuracy + latency + cost]
B --> C[Apply scheme]
C --> D[Measure the same three]
D --> E{Within budget?}
E -->|no| F[Adjust: calibration set,<br/>group size, excluded layers]
F --> C
E -->|yes| G[Report the delta.<br/>Ship with the report attached.]
Learn this loop and every future format slots straight into it.
Modules¶
| M | Module | Hrs | Theme |
|---|---|---|---|
| M1 | Quantization Theory | 10 | Why it works and where it breaks |
| M2 | Post-Training Quantization in Practice | 14 | The llm-compressor workflow, end to end |
| M3 | Beyond Weights | 8 | KV cache, activations, and the other levers |
| M4 | Decoding-Time Optimization | 8 | Speculative decoding and friends |
M1 — Quantization Theory¶
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M1T1 | What quantization actually is | Scale/zero-point, symmetric vs asymmetric, per-tensor vs per-channel vs per-group, the group-size tradeoff | Quantize/dequantize a real weight tensor at every granularity; plot error | Gen: continuous values snapping to a grid | Error-vs-granularity table for real weights |
| M1T2 | The scheme notation | W4A16, W8A8-INT8, W8A8-FP8, W4AFP8, weight-only vs weight+activation — and what each requires from hardware | Scheme→hardware compatibility checker | Mermaid: decision tree, hardware-gated | A correct compatibility matrix for your own GPUs |
| M1T3 | Why activations are harder than weights | Outlier channels, dynamic range, why naive INT8 activations collapse, SmoothQuant's migration trick | Activation-distribution profiler over calibration data; find the outlier channels | Chart: per-channel activation ranges, outliers marked | The outlier channels identified in a real model |
| M1T4 | The algorithm families | RTN, GPTQ (Hessian-weighted, sequential), AWQ (activation-aware, salient channels), SmoothQuant, SpinQuant/QuIP (rotations), AutoRound | Side-by-side on one small model | Chart: accuracy vs algorithm at matched bit-width | An algorithm ranking on your model and task |
| M1T5 | Weight-only ≠ faster at scale | Dequant overhead, why W4A16 is a latency/memory win but not a throughput win at high batch; Marlin-style kernels | Batch-size sweep: W4A16 vs bf16 throughput | Chart: throughput vs batch, crossover marked | The batch size where the win disappears |
M1T5 is the topic that separates people who understand quantization from people who read a blog post about it.
M2 — Post-Training Quantization in Practice¶
Spine: llm-compressor → vLLM deployment.
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M2T1 | The llm-compressor model |
Recipes, modifiers, observers, oneshot entrypoint, compressed-tensors checkpoint format |
First quantization: 8B → W4A16, then serve it | Mermaid: recipe → modifier → checkpoint pipeline | A quantized checkpoint served by vLLM |
| M2T2 | Calibration set design | Size, domain match, sequence length, the silent accuracy killer; why 512 samples of the wrong data ruins a model | Ablate calibration sets: size × domain × length | Chart: accuracy vs calibration set choice | Proof that calibration choice moves accuracy by >X points |
| M2T3 | INT8 W8A8 with SmoothQuant | The primary hands-on activation-quantization track on Ampere; alpha tuning | 8B → W8A8-INT8, alpha sweep | Gen: outlier migration weights↔activations (Pattern B) | Accuracy + throughput vs alpha |
| M2T4 | GPTQ and AWQ at W4A16 | Sequential Hessian updates vs salient-channel scaling; group size 128 vs 64 vs 32; act-order | Both algorithms, multiple group sizes, one matrix of results | Chart: accuracy vs VRAM vs throughput, all variants | A complete decision matrix |
| M2T5 | Big models & distributed quantization | Sequential/layer-wise pipelines, offloading, model-free PTQ, memory requirements | Quantize a 70B on limited VRAM | Mermaid: sequential pipeline memory profile | 70B quantized and served on 2 GPUs |
| M2T6 | MoE-specific quantization | Expert quantization, why MoE is mostly-memory and therefore the best quantization target | Quantize a small MoE; measure VRAM collapse | Chart: VRAM and throughput, dense vs MoE gains | MoE vs dense quantization payoff, compared |
| M2T7 | Newer-silicon formats | FP8 (Ada/Hopper), NVFP4/MXFP8 (Blackwell) — the schemes, the hardware gate, what we'd choose and why | Config walkthrough only; no execution | Gen: format bit-layouts (Pattern B) | A written "next hardware purchase" recommendation |
M2T7 is deliberately non-executable on our hardware. It is taught as a decision framework, and the theory file must say so plainly rather than pretending.
M3 — Beyond Weights¶
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M3T1 | KV cache quantization | INT8 KV cache; why this is often the biggest long-context win; quality impact on long generations | Enable INT8 KV; sweep context length | Chart: max concurrency vs context, quantized vs not | Concurrency gain + long-context quality delta |
| M3T2 | Distillation | Logit vs sequence-level KD, GKD, MiniLLM, on-policy distillation; why a distilled small model often beats a quantized large one | Distill a 7B teacher into a ~1B student on a narrow task | Gen: teacher→student transfer (Pattern C) | Student vs quantized-teacher: accuracy, latency, cost |
| M3T3 | Pruning & sparsity — and why it faded | Magnitude, Wanda, SparseGPT, 2:4 structured sparsity; the hardware-support problem that killed it in practice | Prune a model, measure the accuracy/speed reality | Chart: sparsity vs accuracy vs actual speedup | Evidence for why this is deprioritized today |
| M3T4 | QAT when PTQ isn't enough | Fake quant in the forward pass, STE, when the accuracy budget forces it, cost vs PTQ | QAT fine-tune recovering PTQ loss on one task | Chart: PTQ vs QAT accuracy recovery | The recovered accuracy, and the training cost it took |
| M3T5 | Architectural compression | Layer pruning/depth reduction, width reduction, MHA→GQA conversion, healing with continued training | Drop layers + heal; measure | Chart: layers removed vs accuracy, before/after healing | How many layers you can remove for <2 points |
M4 — Decoding-Time Optimization¶
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M4T1 | Speculative decoding mechanics | Draft/verify, rejection sampling, acceptance rate, expected tokens per step — the math | Toy implementation from scratch, ~120 lines | Gen: draft-and-verify (Pattern C) | Measured acceptance rate matching your predicted value |
| M4T2 | The method zoo | n-gram/prompt lookup, draft models, EAGLE, Medusa, MTP — cost, quality, and where each fits | Compare 3 methods on the same workload | Chart: speedup vs method vs workload type | A method-selection table |
| M4T3 | When it hurts | The compute-for-latency trade; why speculation dies at high batch; GPU saturation | Batch-size sweep to find your crossover | Chart: speedup vs batch, crossing below 1.0 | Your crossover point, and the rule you'd give an ops team |
| M4T4 | Prompt-side optimization | Prompt compression, context pruning, cache-aware prompt design, system-prompt stability | Restructure prompts for prefix reuse; measure hit rate | Gen: shared vs shattered prefix (Pattern B) | TTFT win from prompt restructuring alone |
Labs¶
| Lab | Goal | Success criterion |
|---|---|---|
| LAB-M1 | Quantize by hand | Hand-rolled GPTQ on one layer matching the library's output within tolerance |
| LAB-M2 | The calibration trap | Deliberately ruin a model with a bad calibration set, then diagnose and fix it |
| LAB-M3 | 70B on two GPUs | Serve a 70B on 2×A100 with <2 points accuracy loss on a chosen benchmark |
| LAB-M4 | The speculation crossover | Locate the batch size where speculative decoding stops paying, and explain it |
Capstone¶
Deliverable: compression-tradeoff-report — the portfolio-grade artifact of this skill.
One model family, every viable variant on your hardware:
| Variant | Accuracy Δ | VRAM | TTFT p95 | Throughput | Max concurrency | $/1M tok |
|---|---|---|---|---|---|---|
| bf16 baseline | — | |||||
| W8A8-INT8 (SmoothQuant) | ||||||
| W4A16 (GPTQ, g128) | ||||||
| W4A16 (AWQ, g128) | ||||||
| W4A16 + INT8 KV | ||||||
| Distilled small model |
Plus: methodology, calibration design, eval suite justification, a recommendation with a stated accuracy budget, and a one-page "on Hopper/Blackwell this changes as follows".
Rubric: would a skeptical engineer trust this enough to change production on it? That's the bar.
Assessment¶
| Tier | Count | Example |
|---|---|---|
| Recall | 6 | What does group size control in W4A16, and what's the tradeoff? |
| Apply | 12 | This model lost 8 points on GSM8K after W4A16 but only 1 on MMLU. Explain, and propose two fixes |
| Design | 5 | Latency-sensitive, low-concurrency chat on Ampere, ≤1 point accuracy budget. Choose a scheme, justify, state the risk |
Practical challenge: given a model, a target GPU budget, and an accuracy budget, deliver a serving-ready compressed checkpoint plus a delta report. 4 hours.
Asset inventory¶
| Type | Count |
|---|---|
| Theory files | 21 |
| Code artifacts | 20 |
| Mermaid diagrams | ~18 |
| Generated images | ~12 |
| Charts | ~28 |
| Labs | 4 |
| Capstone | 1 |
Primary sources¶
llm-compressordocs — compression schemes, observers, memory requirements, big-models guide, deploying with vLLM- Papers: GPTQ, AWQ, SmoothQuant, SpinQuant, QuIP#, AutoRound, LLM.int8()
- Speculative decoding: original spec-decode paper, Medusa, EAGLE-1/2/3, DeepSeek MTP
- Distillation: MiniLLM, GKD, and the DeepSeek-R1 distillation appendix
- Pruning: Wanda, SparseGPT — plus the vLLM discussion on dropping sparse support (a good lesson in hardware reality)
- vLLM quantization docs + the Marlin/Machete kernel write-ups
compressed-tensorsformat spec