Skip to content

Skill 03 — Compression & Optimization

The highest ROI-per-hour skill after serving. This is the one that converts directly into "we now need half the GPUs" — the single most requested outcome from anyone paying for inference.

Duration 3 weeks (~40 hrs)
Impact ★★★★☆
Prerequisites Skill 00, Skill 01 (esp. M3 benchmarking)
Hardware 1–2 GPUs for quantization; 2–4 for the 70B stress lab
Status planned

Outcome

  • You can quantize any model and prove the accuracy cost, not guess at it
  • You can pick a scheme from hardware + workload + accuracy budget, and defend it
  • You understand which techniques are latency wins, which are throughput wins, and which are neither
  • You never ship a compressed model without a delta report — and you can explain why anyone who does is negligent

The methodology this skill really teaches

Quantization algorithms churn. The method doesn't:

flowchart LR
    A[Define accuracy budget<br/>on task-relevant evals] --> B[Baseline: measure<br/>accuracy + latency + cost]
    B --> C[Apply scheme]
    C --> D[Measure the same three]
    D --> E{Within budget?}
    E -->|no| F[Adjust: calibration set,<br/>group size, excluded layers]
    F --> C
    E -->|yes| G[Report the delta.<br/>Ship with the report attached.]

Learn this loop and every future format slots straight into it.


Modules

M Module Hrs Theme
M1 Quantization Theory 10 Why it works and where it breaks
M2 Post-Training Quantization in Practice 14 The llm-compressor workflow, end to end
M3 Beyond Weights 8 KV cache, activations, and the other levers
M4 Decoding-Time Optimization 8 Speculative decoding and friends

M1 — Quantization Theory

ID Topic Key concepts Code artifact Visual Evidence
M1T1 What quantization actually is Scale/zero-point, symmetric vs asymmetric, per-tensor vs per-channel vs per-group, the group-size tradeoff Quantize/dequantize a real weight tensor at every granularity; plot error Gen: continuous values snapping to a grid Error-vs-granularity table for real weights
M1T2 The scheme notation W4A16, W8A8-INT8, W8A8-FP8, W4AFP8, weight-only vs weight+activation — and what each requires from hardware Scheme→hardware compatibility checker Mermaid: decision tree, hardware-gated A correct compatibility matrix for your own GPUs
M1T3 Why activations are harder than weights Outlier channels, dynamic range, why naive INT8 activations collapse, SmoothQuant's migration trick Activation-distribution profiler over calibration data; find the outlier channels Chart: per-channel activation ranges, outliers marked The outlier channels identified in a real model
M1T4 The algorithm families RTN, GPTQ (Hessian-weighted, sequential), AWQ (activation-aware, salient channels), SmoothQuant, SpinQuant/QuIP (rotations), AutoRound Side-by-side on one small model Chart: accuracy vs algorithm at matched bit-width An algorithm ranking on your model and task
M1T5 Weight-only ≠ faster at scale Dequant overhead, why W4A16 is a latency/memory win but not a throughput win at high batch; Marlin-style kernels Batch-size sweep: W4A16 vs bf16 throughput Chart: throughput vs batch, crossover marked The batch size where the win disappears

M1T5 is the topic that separates people who understand quantization from people who read a blog post about it.


M2 — Post-Training Quantization in Practice

Spine: llm-compressor → vLLM deployment.

ID Topic Key concepts Code artifact Visual Evidence
M2T1 The llm-compressor model Recipes, modifiers, observers, oneshot entrypoint, compressed-tensors checkpoint format First quantization: 8B → W4A16, then serve it Mermaid: recipe → modifier → checkpoint pipeline A quantized checkpoint served by vLLM
M2T2 Calibration set design Size, domain match, sequence length, the silent accuracy killer; why 512 samples of the wrong data ruins a model Ablate calibration sets: size × domain × length Chart: accuracy vs calibration set choice Proof that calibration choice moves accuracy by >X points
M2T3 INT8 W8A8 with SmoothQuant The primary hands-on activation-quantization track on Ampere; alpha tuning 8B → W8A8-INT8, alpha sweep Gen: outlier migration weights↔activations (Pattern B) Accuracy + throughput vs alpha
M2T4 GPTQ and AWQ at W4A16 Sequential Hessian updates vs salient-channel scaling; group size 128 vs 64 vs 32; act-order Both algorithms, multiple group sizes, one matrix of results Chart: accuracy vs VRAM vs throughput, all variants A complete decision matrix
M2T5 Big models & distributed quantization Sequential/layer-wise pipelines, offloading, model-free PTQ, memory requirements Quantize a 70B on limited VRAM Mermaid: sequential pipeline memory profile 70B quantized and served on 2 GPUs
M2T6 MoE-specific quantization Expert quantization, why MoE is mostly-memory and therefore the best quantization target Quantize a small MoE; measure VRAM collapse Chart: VRAM and throughput, dense vs MoE gains MoE vs dense quantization payoff, compared
M2T7 Newer-silicon formats FP8 (Ada/Hopper), NVFP4/MXFP8 (Blackwell) — the schemes, the hardware gate, what we'd choose and why Config walkthrough only; no execution Gen: format bit-layouts (Pattern B) A written "next hardware purchase" recommendation

M2T7 is deliberately non-executable on our hardware. It is taught as a decision framework, and the theory file must say so plainly rather than pretending.


M3 — Beyond Weights

ID Topic Key concepts Code artifact Visual Evidence
M3T1 KV cache quantization INT8 KV cache; why this is often the biggest long-context win; quality impact on long generations Enable INT8 KV; sweep context length Chart: max concurrency vs context, quantized vs not Concurrency gain + long-context quality delta
M3T2 Distillation Logit vs sequence-level KD, GKD, MiniLLM, on-policy distillation; why a distilled small model often beats a quantized large one Distill a 7B teacher into a ~1B student on a narrow task Gen: teacher→student transfer (Pattern C) Student vs quantized-teacher: accuracy, latency, cost
M3T3 Pruning & sparsity — and why it faded Magnitude, Wanda, SparseGPT, 2:4 structured sparsity; the hardware-support problem that killed it in practice Prune a model, measure the accuracy/speed reality Chart: sparsity vs accuracy vs actual speedup Evidence for why this is deprioritized today
M3T4 QAT when PTQ isn't enough Fake quant in the forward pass, STE, when the accuracy budget forces it, cost vs PTQ QAT fine-tune recovering PTQ loss on one task Chart: PTQ vs QAT accuracy recovery The recovered accuracy, and the training cost it took
M3T5 Architectural compression Layer pruning/depth reduction, width reduction, MHA→GQA conversion, healing with continued training Drop layers + heal; measure Chart: layers removed vs accuracy, before/after healing How many layers you can remove for <2 points

M4 — Decoding-Time Optimization

ID Topic Key concepts Code artifact Visual Evidence
M4T1 Speculative decoding mechanics Draft/verify, rejection sampling, acceptance rate, expected tokens per step — the math Toy implementation from scratch, ~120 lines Gen: draft-and-verify (Pattern C) Measured acceptance rate matching your predicted value
M4T2 The method zoo n-gram/prompt lookup, draft models, EAGLE, Medusa, MTP — cost, quality, and where each fits Compare 3 methods on the same workload Chart: speedup vs method vs workload type A method-selection table
M4T3 When it hurts The compute-for-latency trade; why speculation dies at high batch; GPU saturation Batch-size sweep to find your crossover Chart: speedup vs batch, crossing below 1.0 Your crossover point, and the rule you'd give an ops team
M4T4 Prompt-side optimization Prompt compression, context pruning, cache-aware prompt design, system-prompt stability Restructure prompts for prefix reuse; measure hit rate Gen: shared vs shattered prefix (Pattern B) TTFT win from prompt restructuring alone

Labs

Lab Goal Success criterion
LAB-M1 Quantize by hand Hand-rolled GPTQ on one layer matching the library's output within tolerance
LAB-M2 The calibration trap Deliberately ruin a model with a bad calibration set, then diagnose and fix it
LAB-M3 70B on two GPUs Serve a 70B on 2×A100 with <2 points accuracy loss on a chosen benchmark
LAB-M4 The speculation crossover Locate the batch size where speculative decoding stops paying, and explain it

Capstone

Deliverable: compression-tradeoff-report — the portfolio-grade artifact of this skill.

One model family, every viable variant on your hardware:

Variant Accuracy Δ VRAM TTFT p95 Throughput Max concurrency $/1M tok
bf16 baseline —
W8A8-INT8 (SmoothQuant)
W4A16 (GPTQ, g128)
W4A16 (AWQ, g128)
W4A16 + INT8 KV
Distilled small model

Plus: methodology, calibration design, eval suite justification, a recommendation with a stated accuracy budget, and a one-page "on Hopper/Blackwell this changes as follows".

Rubric: would a skeptical engineer trust this enough to change production on it? That's the bar.


Assessment

Tier Count Example
Recall 6 What does group size control in W4A16, and what's the tradeoff?
Apply 12 This model lost 8 points on GSM8K after W4A16 but only 1 on MMLU. Explain, and propose two fixes
Design 5 Latency-sensitive, low-concurrency chat on Ampere, ≤1 point accuracy budget. Choose a scheme, justify, state the risk

Practical challenge: given a model, a target GPU budget, and an accuracy budget, deliver a serving-ready compressed checkpoint plus a delta report. 4 hours.


Asset inventory

Type Count
Theory files 21
Code artifacts 20
Mermaid diagrams ~18
Generated images ~12
Charts ~28
Labs 4
Capstone 1

Primary sources

  • llm-compressor docs — compression schemes, observers, memory requirements, big-models guide, deploying with vLLM
  • Papers: GPTQ, AWQ, SmoothQuant, SpinQuant, QuIP#, AutoRound, LLM.int8()
  • Speculative decoding: original spec-decode paper, Medusa, EAGLE-1/2/3, DeepSeek MTP
  • Distillation: MiniLLM, GKD, and the DeepSeek-R1 distillation appendix
  • Pruning: Wanda, SparseGPT — plus the vLLM discussion on dropping sparse support (a good lesson in hardware reality)
  • vLLM quantization docs + the Marlin/Machete kernel write-ups
  • compressed-tensors format spec