Skip to content

Skill 06 — Platform & Kubernetes

The scarcest skill in the stack right now. Plenty of people can run vLLM. Very few can run an inference platform — multi-model, multi-tenant, cache-aware, autoscaled, and safely upgradable.

Duration 4 weeks (~55 hrs)
Impact ★★★★★ — the 2025→2027 standardization wave, and salaries reflect the scarcity
Prerequisites Skill 01 (all), Skill 03, basic Kubernetes (Deployments, Services, ConfigMaps)
Hardware Single-node k3s/kind on the GPU host. MIG carves 5–6 A100s into ~35 schedulable instances — that's the unlock that makes fleet-scale labs real
Status planned

Outcome

  • You can stand up a production inference platform: gateway, model pools, routing, autoscaling, rollouts
  • You know why LLM serving breaks every assumption of normal Kubernetes autoscaling — and what to do instead
  • You can implement KV-cache-aware routing and prove it beats round-robin
  • You can run disaggregated prefill/decode on real hardware and explain the KV transfer path

Why LLM serving breaks standard Kubernetes

This is the framing for the whole skill. Teach it in the first hour.

Standard web service LLM inference service
Requests are ~uniform cost Costs vary 1000× (100-token vs 100k-token prompt)
Stateless — any replica will do Stateful — the replica holding your prefix in KV cache is 10× cheaper
Scale on CPU% CPU% is meaningless; scale on queue depth and KV utilization
Startup in seconds Startup in minutes (weight loading) — HPA reacts far too late
Round-robin LB is fine Round-robin destroys cache locality and tail latency
Rollout = swap containers Rollout = re-download 140GB of weights

Every component in this skill exists because of a row in that table.


Modules

M Module Hrs Theme
M1 GPUs on Kubernetes 12 Making accelerators schedulable and shareable
M2 The Inference Gateway 14 Routing that understands LLMs
M3 Platforms & Operators 12 llm-d, KServe, Dynamo, Ray Serve
M4 Autoscaling, Rollouts, Reliability 17 Keeping it up and affordable

M1 — GPUs on Kubernetes

ID Topic Key concepts Code artifact Visual Evidence
M1T1 The GPU stack in a cluster NVIDIA GPU Operator, device plugin, node feature discovery, driver containers, DCGM exporter Cluster bring-up: k3s + GPU Operator, verified end to end Mermaid: operator component map A GPU-schedulable cluster, nvidia.com/gpu allocatable
M1T2 GPU sharing: MIG, time-slicing, MPS Hard vs soft partitioning, memory isolation, why time-slicing is dangerous for latency SLOs Configure MIG profiles; expose them as distinct resources Gen sequence: whole GPU → time-sliced → MIG-partitioned ~35 schedulable instances from 5 GPUs, isolation verified
M1T3 Scheduling for inference Resource requests, topology-aware placement, NUMA, node pools, taints/tolerations, extended resources Scheduling policies + a placement test suite Mermaid: scheduling decision flow Pods landing on NVLink-connected GPUs, verified
M1T4 The weight-loading problem Multi-GB images vs volume-mounted weights, model caching, streaming loaders, warm nodes, init containers Three loading strategies benchmarked Chart: pod-ready time by strategy Cold-start time cut by ≥3×
M1T5 Gang scheduling for multi-GPU jobs All-or-nothing placement, Kueue / Volcano, LeaderWorkerSet for TP-sharded serving, deadlock avoidance LeaderWorkerSet deploying a TP=4 model Mermaid: leader/worker topology A multi-GPU model served as one logical unit

M2 — The Inference Gateway

The heart of the skill. Built on the official Kubernetes Gateway API Inference Extension.

ID Topic Key concepts Code artifact Visual Evidence
M2T1 The resource model InferencePool as an "LLM-optimized Service", InferenceObjective, variants via pod labels, roles and personas Deploy Gateway API + Inference Extension; expose one pool Mermaid: resource model + Gen: pool-of-servers concept An InferencePool serving traffic
M2T2 The Endpoint Picker Envoy ext-proc, scoring endpoints on live metrics, the model-server protocol, why the gateway must read engine internals Trace a request through the EPP with logging Mermaid: full request flow, gateway → EPP → pod A request traced through every hop
M2T3 Routing policies Round-robin vs least-queue vs prefix-cache-aware; scoring and filtering; the model server metrics that drive it Three policies, identical workload Chart: TTFT p95 and cache hit rate per policy Cache-aware routing beating round-robin, quantified
M2T4 KV-cache indexing Event-driven tracking of which pod holds which prefix, precise vs heuristic indexing, index staleness Cache index observer + hit-rate dashboard Gen sequence: request → prefix hash → index lookup → pod choice Index accuracy vs routing quality
M2T5 Model-aware routing & rollouts Routing by model name, LoRA-adapter-aware routing, traffic splitting for model versions, serving priority Canary a new model version by traffic split Mermaid: traffic-split topology A 5%→100% canary with metrics at each step
M2T6 AI gateways above the inference gateway LiteLLM / Envoy AI Gateway: multi-provider routing, virtual keys, budgets, rate limits, fallback to hosted APIs Gateway with quota enforcement and API fallback Mermaid: two-layer gateway architecture Quota enforcement + graceful failover demonstrated

M3 — Platforms & Operators

ID Topic Key concepts Code artifact Visual Evidence
M3T1 llm-d architecture Router (proxy + EPP), InferencePool variants, well-lit paths, how the pieces compose Deploy an llm-d well-lit path Mermaid: llm-d component architecture A working llm-d deployment
M3T2 Disaggregated serving on Kubernetes Prefill and decode as separate variants, router coordinating both, KV transfer via NIXL/connectors 2 prefill + 2 decode pods + router Gen sequence: unified pod → split roles → KV handoff; Mermaid: transfer sequence TTFT/ITL vs colocated, on the same hardware
M3T3 KV offloading at platform scale Tiered CPU/SSD KV storage, shared cache across pods, LMCache-style connectors, when the network beats recompute Enable a KV offload tier; measure honestly Gen: tiered hierarchy (Pattern D) Hit rate and latency per tier — including the disappointing parts
M3T4 KServe InferenceService CRD, ServingRuntime, scale-to-zero, Knative vs raw deployment, transformers/predictors Deploy the same model via KServe Mermaid: KServe reconciliation flow Scale-to-zero with measured cold-start cost
M3T5 NVIDIA Dynamo KV-aware router, disaggregation, the Planner, NIXL — and how it compares to llm-d Architecture study + a minimal deployment Mermaid: Dynamo vs llm-d, side by side A written comparison with a recommendation
M3T6 Ray Serve LLM Python-native composition, mixed workloads, autoscaling model, when it beats a CRD-based platform Multi-stage pipeline: embed → retrieve → generate Mermaid: deployment graph A composite service with independently scaled stages
M3T7 Choosing a platform Decision framework: team skills, existing stack, scale, model count, ops budget Scoring rubric Mermaid: decision tree A platform ADR for a stated scenario

M4 — Autoscaling, Rollouts, Reliability

ID Topic Key concepts Code artifact Visual Evidence
M4T1 Why CPU-based autoscaling fails The GPU-utilization lie, saturation vs utilization, lag vs load, the death spiral Reproduce the death spiral with HPA-on-GPU% Gen: correct vs wrong scaling signal (Pattern B) A documented failure you caused deliberately
M4T2 Scaling on the right signals Queue depth, KV cache utilization, TTFT, SLO burn rate; KEDA with Prometheus triggers KEDA scaler on vLLM metrics Chart: load, replicas, and p95 latency over time Stable scaling under a step-load test
M4T3 Cold start & scale-to-zero Weight-load time, warm pools, over-provision factor, the economics of keeping one replica hot Cold-start economics model Chart: $/month vs p99 latency across warm-pool sizes The break-even point for scale-to-zero
M4T4 SLO-aware and predictive scaling Latency prediction for routing and scaling, workload-variant autoscaling, cost-optimal replica placement SLO-driven scaling policy Mermaid: control loop SLO attainment vs replica count
M4T5 Rollouts & rollbacks Canary by model name, shadow traffic, blue/green with expensive weights, instant rollback, eval gates in CD Full rollout pipeline with an automated gate Mermaid: rollout state machine A bad model automatically blocked by the gate
M4T6 Multi-tenancy & fairness Priority classes, quotas, fair-share scheduling, noisy-neighbour isolation, chargeback Priority-based admission + per-tenant accounting Chart: latency per tenant tier under contention High-priority SLO held while low-priority degrades
M4T7 Capacity planning Headroom, peak-to-average ratio, burst-to-cloud with SkyPilot, spot economics, the buy-vs-rent calculation Capacity model consuming Skill 01 benchmarks Chart: cost vs SLO attainment vs capacity A 12-month capacity plan for a stated growth curve
M4T8 Operating it Health/readiness probes that mean something, graceful drain, back-pressure, OOM recovery, upgrade choreography, on-call runbook Chaos suite against the live platform Mermaid: failure taxonomy → response A runbook written from real induced failures

Labs

Lab Goal Success criterion
LAB-M1 Build the cluster GPU-schedulable k3s with MIG; ~35 instances allocatable
LAB-M2 Cache-aware routing ≥30% TTFT p95 improvement over round-robin on a shared-prefix workload
LAB-M3 Disaggregate on K8s Prefill/decode split deployed and beating colocated under mixed load
LAB-M4 Break the autoscaler Cause a scaling death spiral, diagnose it, then fix it with correct signals
LAB-M5 Ship safely A deliberately-degraded model blocked automatically by the eval gate

Capstone

Deliverable: inference-platform — a running platform plus the documentation to hand it to an ops team.

Must include: 1. GPU-enabled cluster with MIG partitioning and topology-aware scheduling 2. Inference gateway with an InferencePool and prefix-cache-aware endpoint picking 3. At least three models served, one of them with multiple LoRA adapters 4. Disaggregated prefill/decode for at least one model 5. KEDA autoscaling on queue depth with a documented tuning rationale 6. Canary rollout pipeline with an automated eval gate 7. Multi-tenant priority tiers with per-tenant cost accounting 8. Full observability: Grafana dashboards for engine, GPU, and gateway metrics 9. Chaos-tested runbook covering ≥6 failure modes you actually induced 10. An architecture document with the decision rationale for every component choice

Rubric: hand items 9 and 10 to someone who wasn't involved. Can they operate it? Can they explain why it's built this way? That's the bar.


Assessment

Tier Count Example
Recall 7 What is an InferencePool and how does it differ from a Service?
Apply 13 Replicas are scaling up but p95 latency keeps rising. Walk through the diagnosis
Design 5 8 models, 40 tenants, 24 GPUs, mixed interactive and batch, 99.5% SLO. Design the platform

Practical challenge: deploy an unfamiliar model to the platform with cache-aware routing, autoscaling, and a canary rollout, meeting a stated SLO. 4 hours.


Asset inventory

Type Count (minimum)
Theory articles 26
Code artifacts 24
Mermaid diagrams ~40
Generated images ~100
Charts ~25
Labs 5
Capstone 1

Primary sources

  • Kubernetes Gateway API Inference Extension docs — API overview, roles & personas, request flow, conformance, InferencePool/InferenceObjective specs
  • llm-d docs — Architecture (Router, EPP, InferencePool, Model Servers), KV Cache Management (prefix-aware routing, KV indexer, KV offloader), Disaggregation, Autoscaling, Well-Lit Paths
  • NVIDIA Dynamo design docs — overall architecture, architecture flow, feature matrix
  • KServe docs; Ray Serve LLM docs; SkyPilot docs
  • NVIDIA GPU Operator + MIG user guides; DCGM exporter
  • KEDA docs (Prometheus scaler); Kueue and LeaderWorkerSet docs
  • Kubernetes WG-Serving working-group materials — where this is all being standardized