Skip to content

Skill 01 — Inference & Serving

The core investment. This is where the money, the jobs, and the compounding expertise are. If you only ever master one skill in this repo, make it this one.

Duration 6 weeks (~80 hrs)
Impact ★★★★★ — highest demand × highest scarcity × most durable
Prerequisites Skill 00
Hardware 1–4 GPUs. Workhorse 7–8B default; 30B and 70B-quantized for parallelism modules
Status planned

Outcome

You can take any open-weight model and a workload description and:

  • Choose the engine and defend it with your own measurements
  • Tune it to a stated TTFT/throughput/cost target and prove you hit it
  • Explain every knob you turned in terms of memory bandwidth, cache behaviour, or scheduling
  • Diagnose a degraded production server from its metrics alone
  • Design the parallelism layout and know when to disaggregate prefill from decode

Modules

M Module Hrs Theme
M1 The Engine — vLLM in depth 16 What actually happens to a request
M2 Scheduling, Batching, KV Management 14 The mechanisms that create throughput
M3 Benchmarking Like an Engineer 12 Producing numbers you can defend
M4 Parallelism & Scaling 14 Making big models fit and go fast
M5 The Engine Landscape 10 Choosing correctly, with evidence
M6 Advanced Serving Architectures 14 Where the field is going

M1 — The Engine: vLLM in depth

ID Topic Key concepts Code artifact Visual Evidence
M1T1 Request lifecycle end-to-end Entrypoint → tokenizer → scheduler → model runner → sampler → detokenizer → stream; the V1 engine core loop Traced request with timing at each stage Mermaid: sequence diagram, full path Stage-by-stage timing breakdown of one request
M1T2 Offline vs online serving LLM class vs OpenAI-compatible server; when batch offline inference is the right answer Same workload run both ways, throughput compared Mermaid: two deployment shapes Throughput delta offline vs online, explained
M1T3 The memory allocator --gpu-memory-utilization, KV cache block allocation, profiling run at startup, preemption and recompute Sweep gpu-mem-util, plot resulting KV blocks and max concurrency Gen: budget partitioning (Pattern A) Chart: KV blocks vs gpu-memory-utilization
M1T4 Sampling and structured output Temperature/top-p/top-k/min-p, logits processors, guided decoding via xgrammar, the latency cost of constraints Constrained JSON generation with and without grammar; measure overhead Mermaid: logits pipeline Latency overhead of grammar-constrained decoding
M1T5 Tool calling & reasoning parsers Parser plugins, streaming partial tool calls, reasoning-token separation Serve a tool-calling model, parse a streamed call correctly Mermaid: parser state machine Working streamed tool-call handler
M1T6 Reading the source Where to look: v1/core/, v1/worker/, model_executor/layers/, attention backends Guided source tour with annotated excerpts Mermaid: package map of the codebase A written trace of one request through actual source files

M2 — Scheduling, Batching, KV Management

The intellectual heart of the skill.

ID Topic Key concepts Code artifact Visual Evidence
M2T1 Static vs continuous batching Head-of-line blocking, iteration-level scheduling, why this was a 20×+ unlock Simulator: static vs continuous batching on a length distribution Gen: before/after batching timeline (Pattern B) Simulated + measured throughput ratio
M2T2 PagedAttention Fragmentation in naive allocation, block tables, copy-on-write sharing, near-zero waste Block-table visualizer over a live server's cache state Gen: fragmented vs paged memory (Pattern B) Measured memory waste, naive vs paged
M2T3 Chunked prefill Long prefills starving decode; interleaving; the max-num-batched-tokens dial and its ITL/TTFT tradeoff Sweep chunked-prefill settings under mixed load Chart: TTFT vs ITL Pareto across chunk sizes The tradeoff curve, with a chosen operating point
M2T4 Prefix caching Automatic prefix caching, hash-based block reuse, RadixAttention; why RAG and agents benefit enormously Shared-system-prompt workload with cache on/off; measure hit rate Gen: shared prefix tree (Pattern A) TTFT reduction vs cache hit rate curve
M2T5 Preemption, priority, and fairness Swap vs recompute, priority scheduling, starvation, what happens when you oversubscribe Force preemption storms; observe and interpret the metrics Mermaid: scheduler state machine Preemption-rate vs load curve; the cliff located
M2T6 KV offloading & tiering GPU → CPU → SSD tiers, LMCache-style connectors, when offload helps and when it's a trap Enable CPU offload, measure the bandwidth penalty honestly Gen: tiered storage hierarchy (Pattern D) Hit-rate vs latency-penalty table per tier

M3 — Benchmarking Like an Engineer

ID Topic Key concepts Code artifact Visual Evidence
M3T1 The metric vocabulary TTFT, TPOT/ITL, E2E latency, throughput, goodput, queue time, cache hit rate, percentiles vs means Metric extractor from server logs + Prometheus endpoint Mermaid: where each metric is measured on the timeline A glossary you wrote, with the measurement point of each
M3T2 Designing a workload Input/output length distributions, arrival process (Poisson vs closed-loop), why synthetic uniform load lies Workload generator with realistic length distributions Chart: distribution of a real vs synthetic workload Two benchmark runs showing how workload shape changes conclusions
M3T3 Running a defensible benchmark Warmup, steady state, run length, seed control, isolating variables, reporting confidence Reusable harness wrapping vllm bench serve / GuideLLM — A rerunnable harness with config-hashed results
M3T4 Reading the Pareto frontier Latency vs throughput as a curve, not a point; choosing an operating point from an SLO Pareto plotter over a config sweep Chart: the frontier, annotated with configs Your model's frontier with a justified operating point
M3T5 From benchmark to $/1M tokens Utilization, SLO attainment, replica count, the cost of headroom Cost model consuming benchmark output Chart: $/1M tokens vs SLO tightness A cost table for 3 SLO tiers

This module produces the artifact you reuse for the rest of the curriculum. Every later skill benchmarks against it.


M4 — Parallelism & Scaling

ID Topic Key concepts Code artifact Visual Evidence
M4T1 Tensor parallelism Column/row splits, all-reduce per layer, head divisibility, NVLink dependence, diminishing returns TP=1/2/4 sweep on the same model Gen: weight matrix sharding (Pattern A) + Mermaid comms TP scaling efficiency curve; the point it stops paying
M4T2 Pipeline parallelism Stage splits, the bubble, micro-batching, why PP is for cross-node not intra-node PP=2 run; measure and visualize the bubble Gen: pipeline bubble timeline (Pattern C) Measured bubble fraction
M4T3 Data parallelism & replicas Replicas vs bigger parallel groups, when DP beats TP, router in front 4 replicas vs TP=4, same total GPUs, same workload Mermaid: two topologies compared Head-to-head: which wins, and at what concurrency
M4T4 Expert parallelism for MoE Expert sharding, all-to-all, expert-load imbalance, EPLB EP=2/4 on a small MoE; plot expert load Gen: expert routing across devices Expert-load imbalance and its throughput cost
M4T5 Context parallelism Long-context sharding, ring attention family, when it's the only option One long-context CP config, demonstrated Mermaid: sequence sharding Max context achieved with vs without CP
M4T6 Choosing a layout Decision framework: model size × GPU count × interconnect × workload shape Layout recommender consuming the Skill 00 sizing tool Mermaid: decision tree Recommendations validated against M4T1–T5 measurements

M5 — The Engine Landscape

ID Topic Key concepts Code artifact Visual Evidence
M5T1 SGLang RadixAttention, structured output strength, agentic/multi-turn reuse, scheduler differences Same workload as M3 harness, on SGLang — Head-to-head vs vLLM on identical workload
M5T2 TensorRT-LLM Ahead-of-time engine build, kernel specialization, the ergonomics tax, when the perf is worth it Build and serve one engine; time the build Mermaid: build vs runtime pipeline compared Perf delta vs build-time and flexibility cost
M5T3 llama.cpp / GGUF CPU/edge/consumer serving, GGUF quant tiers, when this is genuinely the right answer Serve a GGUF model on CPU; measure — CPU-serving cost/latency vs GPU, honestly compared
M5T4 The rest of the field TGI, LMDeploy, MLC, Ollama, TorchServe — what each is actually for — Mermaid: positioning map A one-page landscape you can defend
M5T5 The selection framework Criteria: perf, model coverage, feature velocity, ops burden, community, licence Scoring rubric applied to your workload Mermaid: decision tree A written engine ADR (architecture decision record)

M6 — Advanced Serving Architectures

ID Topic Key concepts Code artifact Visual Evidence
M6T1 Speculative decoding Draft-then-verify, acceptance rate, n-gram / EAGLE / MTP, why it hurts at high batch Enable spec decode; sweep batch size and measure acceptance Gen: draft-and-verify (Pattern C) Speedup vs batch-size curve, including the crossover to harm
M6T2 Disaggregated prefill/decode Separate worker pools, KV transfer (NIXL/Mooncake/LMCache), why opposite hardware profiles demand separation 2 prefill + 2 decode + router, on our own GPUs Mermaid: KV handoff sequence + Gen: split pipeline TTFT/ITL vs colocated baseline
M6T3 Multi-LoRA serving Adapter-per-tenant, dynamic loading, punica kernels, the multi-tenant cost unlock Serve one base + 5 adapters; measure per-adapter overhead Gen: shared base with swappable adapters Cost per tenant: multi-LoRA vs separate deployments
M6T4 Multimodal & embedding serving Vision encoder cost, encoder caching, pooling models, why embeddings serve differently Serve a VLM and an embedding model Mermaid: multimodal request path Prefill cost breakdown: image tokens vs text tokens
M6T5 Long-context serving economics Quadratic prefill, KV growth, chunking strategies, when RAG beats long context on pure cost Cost-per-request vs context length sweep Chart: cost vs context length, with the RAG crossover The context length where RAG becomes cheaper
M6T6 Production hardening Health checks, graceful drain, request timeouts, back-pressure, OOM recovery, version pinning Chaos script: kill/starve/overload a live server Mermaid: failure/recovery states A runbook written from observed failures

Labs

Lab Goal Success criterion
LAB-M1 First serve Model serving with streaming + tool calls; explain every startup log line
LAB-M2 Cache detective Achieve >70% prefix-cache hit rate on a RAG-shaped workload and quantify the TTFT win
LAB-M3 The benchmark harness Reproduce your own results twice within 5% variance
LAB-M4 Parallelism bake-off TP=4 vs 4×DP on identical hardware/workload; declare a winner with data
LAB-M5 Engine bake-off vLLM vs SGLang on your workload; a written ADR
LAB-M6 Disaggregate it Prefill/decode split beating the colocated baseline on TTFT p95 under mixed load

Each lab includes a wrong path — a plausible approach that fails, which the learner must diagnose.


Capstone

Deliverable: serving-decision-report — a genuine consulting artifact for one realistic workload (pick: RAG assistant, coding agent, or batch document extraction).

Required contents: 1. Workload characterization with measured length distributions 2. Sizing prediction from the Skill 00 tool, then measured reality, then the explained gap 3. Engine selection ADR with head-to-head data 4. Parallelism layout with scaling curves justifying it 5. Tuned configuration with a Pareto frontier and a chosen operating point against a stated SLO 6. $/1M input and output tokens, and replica count for a target QPS 7. Failure-mode runbook 8. "What I'd do differently on Hopper/Blackwell" — one page

Rubric (0–4 each): measurement rigour · reproducibility · depth of causal explanation · quality of tradeoff reasoning · communicability to a non-specialist.

This document is the single strongest thing you'll have to show for the whole curriculum. Treat it accordingly.


Assessment

Tier Count Example
Recall 8 What does --max-num-batched-tokens control, and what breaks if it's too low?
Apply 12 Given this metrics dashboard, diagnose why p99 TTFT tripled while throughput stayed flat
Design 5 4×A100 80GB, 8k-in/200-out RAG, 50 concurrent users, p95 TTFT ≤800ms. Full architecture, with justification

Practical challenge: an unfamiliar model + an SLO. Reach it, or prove it's impossible on the given hardware. 4 hours.


Asset inventory

Type Count
Theory files 33
Code artifacts 30
Mermaid diagrams ~40
Generated images ~16
Charts ~30
Labs 6
Capstone 1

Primary sources

  • vLLM docs → Design Documents (not just the user guide): Architecture Overview, Paged Attention, Automatic Prefix Caching, Hybrid KV Cache Manager, Metrics, Optimization Levels, torch.compile integration
  • vLLM paper, SOSP 2023 (arXiv 2309.06180)
  • Anyscale — How continuous batching enables 23x throughput in LLM inference
  • SGLang docs + the RadixAttention paper
  • Sarathi-Serve paper (chunked prefill), DistServe & Splitwise papers (disaggregation), Mooncake paper (KV-centric architecture)
  • EAGLE / Medusa / MTP papers (speculative decoding)
  • NVIDIA TensorRT-LLM docs; NVIDIA Dynamo design docs
  • vLLM source: vllm/v1/core/sched/, vllm/v1/core/kv_cache_manager.py, vllm/v1/worker/