The scarcest skill in the stack right now. Plenty of people can run vLLM. Very few can run an
inference platform — multi-model, multi-tenant, cache-aware, autoscaled, and safely upgradable.
|
|
| Duration |
4 weeks (~55 hrs) |
| Impact |
★★★★★ — the 2025→2027 standardization wave, and salaries reflect the scarcity |
| Prerequisites |
Skill 01 (all), Skill 03, basic Kubernetes (Deployments, Services, ConfigMaps) |
| Hardware |
Single-node k3s/kind on the GPU host. MIG carves 5–6 A100s into ~35 schedulable instances — that's the unlock that makes fleet-scale labs real |
| Status |
planned |
Outcome
- You can stand up a production inference platform: gateway, model pools, routing, autoscaling, rollouts
- You know why LLM serving breaks every assumption of normal Kubernetes autoscaling — and what to do instead
- You can implement KV-cache-aware routing and prove it beats round-robin
- You can run disaggregated prefill/decode on real hardware and explain the KV transfer path
Why LLM serving breaks standard Kubernetes
This is the framing for the whole skill. Teach it in the first hour.
| Standard web service |
LLM inference service |
| Requests are ~uniform cost |
Costs vary 1000× (100-token vs 100k-token prompt) |
| Stateless — any replica will do |
Stateful — the replica holding your prefix in KV cache is 10× cheaper |
| Scale on CPU% |
CPU% is meaningless; scale on queue depth and KV utilization |
| Startup in seconds |
Startup in minutes (weight loading) — HPA reacts far too late |
| Round-robin LB is fine |
Round-robin destroys cache locality and tail latency |
| Rollout = swap containers |
Rollout = re-download 140GB of weights |
Every component in this skill exists because of a row in that table.
Modules
| M |
Module |
Hrs |
Theme |
| M1 |
GPUs on Kubernetes |
12 |
Making accelerators schedulable and shareable |
| M2 |
The Inference Gateway |
14 |
Routing that understands LLMs |
| M3 |
Platforms & Operators |
12 |
llm-d, KServe, Dynamo, Ray Serve |
| M4 |
Autoscaling, Rollouts, Reliability |
17 |
Keeping it up and affordable |
M1 — GPUs on Kubernetes
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M1T1 |
The GPU stack in a cluster |
NVIDIA GPU Operator, device plugin, node feature discovery, driver containers, DCGM exporter |
Cluster bring-up: k3s + GPU Operator, verified end to end |
Mermaid: operator component map |
A GPU-schedulable cluster, nvidia.com/gpu allocatable |
| M1T2 |
GPU sharing: MIG, time-slicing, MPS |
Hard vs soft partitioning, memory isolation, why time-slicing is dangerous for latency SLOs |
Configure MIG profiles; expose them as distinct resources |
Gen sequence: whole GPU → time-sliced → MIG-partitioned |
~35 schedulable instances from 5 GPUs, isolation verified |
| M1T3 |
Scheduling for inference |
Resource requests, topology-aware placement, NUMA, node pools, taints/tolerations, extended resources |
Scheduling policies + a placement test suite |
Mermaid: scheduling decision flow |
Pods landing on NVLink-connected GPUs, verified |
| M1T4 |
The weight-loading problem |
Multi-GB images vs volume-mounted weights, model caching, streaming loaders, warm nodes, init containers |
Three loading strategies benchmarked |
Chart: pod-ready time by strategy |
Cold-start time cut by ≥3× |
| M1T5 |
Gang scheduling for multi-GPU jobs |
All-or-nothing placement, Kueue / Volcano, LeaderWorkerSet for TP-sharded serving, deadlock avoidance |
LeaderWorkerSet deploying a TP=4 model |
Mermaid: leader/worker topology |
A multi-GPU model served as one logical unit |
M2 — The Inference Gateway
The heart of the skill. Built on the official Kubernetes Gateway API Inference Extension.
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M2T1 |
The resource model |
InferencePool as an "LLM-optimized Service", InferenceObjective, variants via pod labels, roles and personas |
Deploy Gateway API + Inference Extension; expose one pool |
Mermaid: resource model + Gen: pool-of-servers concept |
An InferencePool serving traffic |
| M2T2 |
The Endpoint Picker |
Envoy ext-proc, scoring endpoints on live metrics, the model-server protocol, why the gateway must read engine internals |
Trace a request through the EPP with logging |
Mermaid: full request flow, gateway → EPP → pod |
A request traced through every hop |
| M2T3 |
Routing policies |
Round-robin vs least-queue vs prefix-cache-aware; scoring and filtering; the model server metrics that drive it |
Three policies, identical workload |
Chart: TTFT p95 and cache hit rate per policy |
Cache-aware routing beating round-robin, quantified |
| M2T4 |
KV-cache indexing |
Event-driven tracking of which pod holds which prefix, precise vs heuristic indexing, index staleness |
Cache index observer + hit-rate dashboard |
Gen sequence: request → prefix hash → index lookup → pod choice |
Index accuracy vs routing quality |
| M2T5 |
Model-aware routing & rollouts |
Routing by model name, LoRA-adapter-aware routing, traffic splitting for model versions, serving priority |
Canary a new model version by traffic split |
Mermaid: traffic-split topology |
A 5%→100% canary with metrics at each step |
| M2T6 |
AI gateways above the inference gateway |
LiteLLM / Envoy AI Gateway: multi-provider routing, virtual keys, budgets, rate limits, fallback to hosted APIs |
Gateway with quota enforcement and API fallback |
Mermaid: two-layer gateway architecture |
Quota enforcement + graceful failover demonstrated |
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M3T1 |
llm-d architecture |
Router (proxy + EPP), InferencePool variants, well-lit paths, how the pieces compose |
Deploy an llm-d well-lit path |
Mermaid: llm-d component architecture |
A working llm-d deployment |
| M3T2 |
Disaggregated serving on Kubernetes |
Prefill and decode as separate variants, router coordinating both, KV transfer via NIXL/connectors |
2 prefill + 2 decode pods + router |
Gen sequence: unified pod → split roles → KV handoff; Mermaid: transfer sequence |
TTFT/ITL vs colocated, on the same hardware |
| M3T3 |
KV offloading at platform scale |
Tiered CPU/SSD KV storage, shared cache across pods, LMCache-style connectors, when the network beats recompute |
Enable a KV offload tier; measure honestly |
Gen: tiered hierarchy (Pattern D) |
Hit rate and latency per tier — including the disappointing parts |
| M3T4 |
KServe |
InferenceService CRD, ServingRuntime, scale-to-zero, Knative vs raw deployment, transformers/predictors |
Deploy the same model via KServe |
Mermaid: KServe reconciliation flow |
Scale-to-zero with measured cold-start cost |
| M3T5 |
NVIDIA Dynamo |
KV-aware router, disaggregation, the Planner, NIXL — and how it compares to llm-d |
Architecture study + a minimal deployment |
Mermaid: Dynamo vs llm-d, side by side |
A written comparison with a recommendation |
| M3T6 |
Ray Serve LLM |
Python-native composition, mixed workloads, autoscaling model, when it beats a CRD-based platform |
Multi-stage pipeline: embed → retrieve → generate |
Mermaid: deployment graph |
A composite service with independently scaled stages |
| M3T7 |
Choosing a platform |
Decision framework: team skills, existing stack, scale, model count, ops budget |
Scoring rubric |
Mermaid: decision tree |
A platform ADR for a stated scenario |
M4 — Autoscaling, Rollouts, Reliability
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M4T1 |
Why CPU-based autoscaling fails |
The GPU-utilization lie, saturation vs utilization, lag vs load, the death spiral |
Reproduce the death spiral with HPA-on-GPU% |
Gen: correct vs wrong scaling signal (Pattern B) |
A documented failure you caused deliberately |
| M4T2 |
Scaling on the right signals |
Queue depth, KV cache utilization, TTFT, SLO burn rate; KEDA with Prometheus triggers |
KEDA scaler on vLLM metrics |
Chart: load, replicas, and p95 latency over time |
Stable scaling under a step-load test |
| M4T3 |
Cold start & scale-to-zero |
Weight-load time, warm pools, over-provision factor, the economics of keeping one replica hot |
Cold-start economics model |
Chart: $/month vs p99 latency across warm-pool sizes |
The break-even point for scale-to-zero |
| M4T4 |
SLO-aware and predictive scaling |
Latency prediction for routing and scaling, workload-variant autoscaling, cost-optimal replica placement |
SLO-driven scaling policy |
Mermaid: control loop |
SLO attainment vs replica count |
| M4T5 |
Rollouts & rollbacks |
Canary by model name, shadow traffic, blue/green with expensive weights, instant rollback, eval gates in CD |
Full rollout pipeline with an automated gate |
Mermaid: rollout state machine |
A bad model automatically blocked by the gate |
| M4T6 |
Multi-tenancy & fairness |
Priority classes, quotas, fair-share scheduling, noisy-neighbour isolation, chargeback |
Priority-based admission + per-tenant accounting |
Chart: latency per tenant tier under contention |
High-priority SLO held while low-priority degrades |
| M4T7 |
Capacity planning |
Headroom, peak-to-average ratio, burst-to-cloud with SkyPilot, spot economics, the buy-vs-rent calculation |
Capacity model consuming Skill 01 benchmarks |
Chart: cost vs SLO attainment vs capacity |
A 12-month capacity plan for a stated growth curve |
| M4T8 |
Operating it |
Health/readiness probes that mean something, graceful drain, back-pressure, OOM recovery, upgrade choreography, on-call runbook |
Chaos suite against the live platform |
Mermaid: failure taxonomy → response |
A runbook written from real induced failures |
Labs
| Lab |
Goal |
Success criterion |
| LAB-M1 |
Build the cluster |
GPU-schedulable k3s with MIG; ~35 instances allocatable |
| LAB-M2 |
Cache-aware routing |
≥30% TTFT p95 improvement over round-robin on a shared-prefix workload |
| LAB-M3 |
Disaggregate on K8s |
Prefill/decode split deployed and beating colocated under mixed load |
| LAB-M4 |
Break the autoscaler |
Cause a scaling death spiral, diagnose it, then fix it with correct signals |
| LAB-M5 |
Ship safely |
A deliberately-degraded model blocked automatically by the eval gate |
Capstone
Deliverable: inference-platform — a running platform plus the documentation to hand it to an ops team.
Must include:
1. GPU-enabled cluster with MIG partitioning and topology-aware scheduling
2. Inference gateway with an InferencePool and prefix-cache-aware endpoint picking
3. At least three models served, one of them with multiple LoRA adapters
4. Disaggregated prefill/decode for at least one model
5. KEDA autoscaling on queue depth with a documented tuning rationale
6. Canary rollout pipeline with an automated eval gate
7. Multi-tenant priority tiers with per-tenant cost accounting
8. Full observability: Grafana dashboards for engine, GPU, and gateway metrics
9. Chaos-tested runbook covering ≥6 failure modes you actually induced
10. An architecture document with the decision rationale for every component choice
Rubric: hand items 9 and 10 to someone who wasn't involved. Can they operate it? Can they explain
why it's built this way? That's the bar.
Assessment
| Tier |
Count |
Example |
| Recall |
7 |
What is an InferencePool and how does it differ from a Service? |
| Apply |
13 |
Replicas are scaling up but p95 latency keeps rising. Walk through the diagnosis |
| Design |
5 |
8 models, 40 tenants, 24 GPUs, mixed interactive and batch, 99.5% SLO. Design the platform |
Practical challenge: deploy an unfamiliar model to the platform with cache-aware routing,
autoscaling, and a canary rollout, meeting a stated SLO. 4 hours.
Asset inventory
| Type |
Count (minimum) |
| Theory articles |
26 |
| Code artifacts |
24 |
| Mermaid diagrams |
~40 |
| Generated images |
~100 |
| Charts |
~25 |
| Labs |
5 |
| Capstone |
1 |
Primary sources
- Kubernetes Gateway API Inference Extension docs — API overview, roles & personas, request flow, conformance, InferencePool/InferenceObjective specs
- llm-d docs — Architecture (Router, EPP, InferencePool, Model Servers), KV Cache Management (prefix-aware routing, KV indexer, KV offloader), Disaggregation, Autoscaling, Well-Lit Paths
- NVIDIA Dynamo design docs — overall architecture, architecture flow, feature matrix
- KServe docs; Ray Serve LLM docs; SkyPilot docs
- NVIDIA GPU Operator + MIG user guides; DCGM exporter
- KEDA docs (Prometheus scaler); Kueue and LeaderWorkerSet docs
- Kubernetes WG-Serving working-group materials — where this is all being standardized