Skill 01 — Inference & Serving
The core investment. This is where the money, the jobs, and the compounding expertise are.
If you only ever master one skill in this repo, make it this one.
|
|
| Duration |
6 weeks (~80 hrs) |
| Impact |
★★★★★ — highest demand × highest scarcity × most durable |
| Prerequisites |
Skill 00 |
| Hardware |
1–4 GPUs. Workhorse 7–8B default; 30B and 70B-quantized for parallelism modules |
| Status |
planned |
Outcome
You can take any open-weight model and a workload description and:
- Choose the engine and defend it with your own measurements
- Tune it to a stated TTFT/throughput/cost target and prove you hit it
- Explain every knob you turned in terms of memory bandwidth, cache behaviour, or scheduling
- Diagnose a degraded production server from its metrics alone
- Design the parallelism layout and know when to disaggregate prefill from decode
Modules
| M |
Module |
Hrs |
Theme |
| M1 |
The Engine — vLLM in depth |
16 |
What actually happens to a request |
| M2 |
Scheduling, Batching, KV Management |
14 |
The mechanisms that create throughput |
| M3 |
Benchmarking Like an Engineer |
12 |
Producing numbers you can defend |
| M4 |
Parallelism & Scaling |
14 |
Making big models fit and go fast |
| M5 |
The Engine Landscape |
10 |
Choosing correctly, with evidence |
| M6 |
Advanced Serving Architectures |
14 |
Where the field is going |
M1 — The Engine: vLLM in depth
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M1T1 |
Request lifecycle end-to-end |
Entrypoint → tokenizer → scheduler → model runner → sampler → detokenizer → stream; the V1 engine core loop |
Traced request with timing at each stage |
Mermaid: sequence diagram, full path |
Stage-by-stage timing breakdown of one request |
| M1T2 |
Offline vs online serving |
LLM class vs OpenAI-compatible server; when batch offline inference is the right answer |
Same workload run both ways, throughput compared |
Mermaid: two deployment shapes |
Throughput delta offline vs online, explained |
| M1T3 |
The memory allocator |
--gpu-memory-utilization, KV cache block allocation, profiling run at startup, preemption and recompute |
Sweep gpu-mem-util, plot resulting KV blocks and max concurrency |
Gen: budget partitioning (Pattern A) |
Chart: KV blocks vs gpu-memory-utilization |
| M1T4 |
Sampling and structured output |
Temperature/top-p/top-k/min-p, logits processors, guided decoding via xgrammar, the latency cost of constraints |
Constrained JSON generation with and without grammar; measure overhead |
Mermaid: logits pipeline |
Latency overhead of grammar-constrained decoding |
| M1T5 |
Tool calling & reasoning parsers |
Parser plugins, streaming partial tool calls, reasoning-token separation |
Serve a tool-calling model, parse a streamed call correctly |
Mermaid: parser state machine |
Working streamed tool-call handler |
| M1T6 |
Reading the source |
Where to look: v1/core/, v1/worker/, model_executor/layers/, attention backends |
Guided source tour with annotated excerpts |
Mermaid: package map of the codebase |
A written trace of one request through actual source files |
M2 — Scheduling, Batching, KV Management
The intellectual heart of the skill.
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M2T1 |
Static vs continuous batching |
Head-of-line blocking, iteration-level scheduling, why this was a 20×+ unlock |
Simulator: static vs continuous batching on a length distribution |
Gen: before/after batching timeline (Pattern B) |
Simulated + measured throughput ratio |
| M2T2 |
PagedAttention |
Fragmentation in naive allocation, block tables, copy-on-write sharing, near-zero waste |
Block-table visualizer over a live server's cache state |
Gen: fragmented vs paged memory (Pattern B) |
Measured memory waste, naive vs paged |
| M2T3 |
Chunked prefill |
Long prefills starving decode; interleaving; the max-num-batched-tokens dial and its ITL/TTFT tradeoff |
Sweep chunked-prefill settings under mixed load |
Chart: TTFT vs ITL Pareto across chunk sizes |
The tradeoff curve, with a chosen operating point |
| M2T4 |
Prefix caching |
Automatic prefix caching, hash-based block reuse, RadixAttention; why RAG and agents benefit enormously |
Shared-system-prompt workload with cache on/off; measure hit rate |
Gen: shared prefix tree (Pattern A) |
TTFT reduction vs cache hit rate curve |
| M2T5 |
Preemption, priority, and fairness |
Swap vs recompute, priority scheduling, starvation, what happens when you oversubscribe |
Force preemption storms; observe and interpret the metrics |
Mermaid: scheduler state machine |
Preemption-rate vs load curve; the cliff located |
| M2T6 |
KV offloading & tiering |
GPU → CPU → SSD tiers, LMCache-style connectors, when offload helps and when it's a trap |
Enable CPU offload, measure the bandwidth penalty honestly |
Gen: tiered storage hierarchy (Pattern D) |
Hit-rate vs latency-penalty table per tier |
M3 — Benchmarking Like an Engineer
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M3T1 |
The metric vocabulary |
TTFT, TPOT/ITL, E2E latency, throughput, goodput, queue time, cache hit rate, percentiles vs means |
Metric extractor from server logs + Prometheus endpoint |
Mermaid: where each metric is measured on the timeline |
A glossary you wrote, with the measurement point of each |
| M3T2 |
Designing a workload |
Input/output length distributions, arrival process (Poisson vs closed-loop), why synthetic uniform load lies |
Workload generator with realistic length distributions |
Chart: distribution of a real vs synthetic workload |
Two benchmark runs showing how workload shape changes conclusions |
| M3T3 |
Running a defensible benchmark |
Warmup, steady state, run length, seed control, isolating variables, reporting confidence |
Reusable harness wrapping vllm bench serve / GuideLLM |
— |
A rerunnable harness with config-hashed results |
| M3T4 |
Reading the Pareto frontier |
Latency vs throughput as a curve, not a point; choosing an operating point from an SLO |
Pareto plotter over a config sweep |
Chart: the frontier, annotated with configs |
Your model's frontier with a justified operating point |
| M3T5 |
From benchmark to $/1M tokens |
Utilization, SLO attainment, replica count, the cost of headroom |
Cost model consuming benchmark output |
Chart: $/1M tokens vs SLO tightness |
A cost table for 3 SLO tiers |
This module produces the artifact you reuse for the rest of the curriculum. Every later skill benchmarks against it.
M4 — Parallelism & Scaling
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M4T1 |
Tensor parallelism |
Column/row splits, all-reduce per layer, head divisibility, NVLink dependence, diminishing returns |
TP=1/2/4 sweep on the same model |
Gen: weight matrix sharding (Pattern A) + Mermaid comms |
TP scaling efficiency curve; the point it stops paying |
| M4T2 |
Pipeline parallelism |
Stage splits, the bubble, micro-batching, why PP is for cross-node not intra-node |
PP=2 run; measure and visualize the bubble |
Gen: pipeline bubble timeline (Pattern C) |
Measured bubble fraction |
| M4T3 |
Data parallelism & replicas |
Replicas vs bigger parallel groups, when DP beats TP, router in front |
4 replicas vs TP=4, same total GPUs, same workload |
Mermaid: two topologies compared |
Head-to-head: which wins, and at what concurrency |
| M4T4 |
Expert parallelism for MoE |
Expert sharding, all-to-all, expert-load imbalance, EPLB |
EP=2/4 on a small MoE; plot expert load |
Gen: expert routing across devices |
Expert-load imbalance and its throughput cost |
| M4T5 |
Context parallelism |
Long-context sharding, ring attention family, when it's the only option |
One long-context CP config, demonstrated |
Mermaid: sequence sharding |
Max context achieved with vs without CP |
| M4T6 |
Choosing a layout |
Decision framework: model size × GPU count × interconnect × workload shape |
Layout recommender consuming the Skill 00 sizing tool |
Mermaid: decision tree |
Recommendations validated against M4T1–T5 measurements |
M5 — The Engine Landscape
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M5T1 |
SGLang |
RadixAttention, structured output strength, agentic/multi-turn reuse, scheduler differences |
Same workload as M3 harness, on SGLang |
— |
Head-to-head vs vLLM on identical workload |
| M5T2 |
TensorRT-LLM |
Ahead-of-time engine build, kernel specialization, the ergonomics tax, when the perf is worth it |
Build and serve one engine; time the build |
Mermaid: build vs runtime pipeline compared |
Perf delta vs build-time and flexibility cost |
| M5T3 |
llama.cpp / GGUF |
CPU/edge/consumer serving, GGUF quant tiers, when this is genuinely the right answer |
Serve a GGUF model on CPU; measure |
— |
CPU-serving cost/latency vs GPU, honestly compared |
| M5T4 |
The rest of the field |
TGI, LMDeploy, MLC, Ollama, TorchServe — what each is actually for |
— |
Mermaid: positioning map |
A one-page landscape you can defend |
| M5T5 |
The selection framework |
Criteria: perf, model coverage, feature velocity, ops burden, community, licence |
Scoring rubric applied to your workload |
Mermaid: decision tree |
A written engine ADR (architecture decision record) |
M6 — Advanced Serving Architectures
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M6T1 |
Speculative decoding |
Draft-then-verify, acceptance rate, n-gram / EAGLE / MTP, why it hurts at high batch |
Enable spec decode; sweep batch size and measure acceptance |
Gen: draft-and-verify (Pattern C) |
Speedup vs batch-size curve, including the crossover to harm |
| M6T2 |
Disaggregated prefill/decode |
Separate worker pools, KV transfer (NIXL/Mooncake/LMCache), why opposite hardware profiles demand separation |
2 prefill + 2 decode + router, on our own GPUs |
Mermaid: KV handoff sequence + Gen: split pipeline |
TTFT/ITL vs colocated baseline |
| M6T3 |
Multi-LoRA serving |
Adapter-per-tenant, dynamic loading, punica kernels, the multi-tenant cost unlock |
Serve one base + 5 adapters; measure per-adapter overhead |
Gen: shared base with swappable adapters |
Cost per tenant: multi-LoRA vs separate deployments |
| M6T4 |
Multimodal & embedding serving |
Vision encoder cost, encoder caching, pooling models, why embeddings serve differently |
Serve a VLM and an embedding model |
Mermaid: multimodal request path |
Prefill cost breakdown: image tokens vs text tokens |
| M6T5 |
Long-context serving economics |
Quadratic prefill, KV growth, chunking strategies, when RAG beats long context on pure cost |
Cost-per-request vs context length sweep |
Chart: cost vs context length, with the RAG crossover |
The context length where RAG becomes cheaper |
| M6T6 |
Production hardening |
Health checks, graceful drain, request timeouts, back-pressure, OOM recovery, version pinning |
Chaos script: kill/starve/overload a live server |
Mermaid: failure/recovery states |
A runbook written from observed failures |
Labs
| Lab |
Goal |
Success criterion |
| LAB-M1 |
First serve |
Model serving with streaming + tool calls; explain every startup log line |
| LAB-M2 |
Cache detective |
Achieve >70% prefix-cache hit rate on a RAG-shaped workload and quantify the TTFT win |
| LAB-M3 |
The benchmark harness |
Reproduce your own results twice within 5% variance |
| LAB-M4 |
Parallelism bake-off |
TP=4 vs 4×DP on identical hardware/workload; declare a winner with data |
| LAB-M5 |
Engine bake-off |
vLLM vs SGLang on your workload; a written ADR |
| LAB-M6 |
Disaggregate it |
Prefill/decode split beating the colocated baseline on TTFT p95 under mixed load |
Each lab includes a wrong path — a plausible approach that fails, which the learner must diagnose.
Capstone
Deliverable: serving-decision-report — a genuine consulting artifact for one realistic workload (pick: RAG assistant, coding agent, or batch document extraction).
Required contents:
1. Workload characterization with measured length distributions
2. Sizing prediction from the Skill 00 tool, then measured reality, then the explained gap
3. Engine selection ADR with head-to-head data
4. Parallelism layout with scaling curves justifying it
5. Tuned configuration with a Pareto frontier and a chosen operating point against a stated SLO
6. $/1M input and output tokens, and replica count for a target QPS
7. Failure-mode runbook
8. "What I'd do differently on Hopper/Blackwell" — one page
Rubric (0–4 each): measurement rigour · reproducibility · depth of causal explanation · quality of tradeoff reasoning · communicability to a non-specialist.
This document is the single strongest thing you'll have to show for the whole curriculum. Treat it accordingly.
Assessment
| Tier |
Count |
Example |
| Recall |
8 |
What does --max-num-batched-tokens control, and what breaks if it's too low? |
| Apply |
12 |
Given this metrics dashboard, diagnose why p99 TTFT tripled while throughput stayed flat |
| Design |
5 |
4×A100 80GB, 8k-in/200-out RAG, 50 concurrent users, p95 TTFT ≤800ms. Full architecture, with justification |
Practical challenge: an unfamiliar model + an SLO. Reach it, or prove it's impossible on the given hardware. 4 hours.
Asset inventory
| Type |
Count |
| Theory files |
33 |
| Code artifacts |
30 |
| Mermaid diagrams |
~40 |
| Generated images |
~16 |
| Charts |
~30 |
| Labs |
6 |
| Capstone |
1 |
Primary sources
- vLLM docs → Design Documents (not just the user guide): Architecture Overview, Paged Attention, Automatic Prefix Caching, Hybrid KV Cache Manager, Metrics, Optimization Levels, torch.compile integration
- vLLM paper, SOSP 2023 (arXiv 2309.06180)
- Anyscale — How continuous batching enables 23x throughput in LLM inference
- SGLang docs + the RadixAttention paper
- Sarathi-Serve paper (chunked prefill), DistServe & Splitwise papers (disaggregation), Mooncake paper (KV-centric architecture)
- EAGLE / Medusa / MTP papers (speculative decoding)
- NVIDIA TensorRT-LLM docs; NVIDIA Dynamo design docs
- vLLM source:
vllm/v1/core/sched/, vllm/v1/core/kv_cache_manager.py, vllm/v1/worker/