Open-Source LLM Infrastructure: Expert Learning Path¶
Starting point: competent LLM user (RAG, agents, prompt/tool orchestration). Target: person who can custom-build, host, serve, optimize, monitor and post-train open-weight models on owned GPUs — and defend the architecture decisions. Assumed effort: 12–15 hrs/week. Horizon: 6 months to strong practitioner, 12 months to differentiated expert.
0. The strategic answer first¶
"Should I learn/build from scratch, or jump straight into hosting, inference, parallelism, Kubernetes?"
Do not start from scratch. Start at the serving layer and spiral outward. Reasons:
- Serving is where the money and the jobs are. Pretraining a frontier model is a ~10-lab activity. Serving, optimizing, and post-training open weights is a ~every-enterprise activity. The scarce skill in 2026 is "make Qwen/Llama/GLM/DeepSpeek-class weights run at target latency and cost on our GPUs, safely, and keep them there."
- Serving gives fastest feedback loops. A vLLM config change shows up in a p99 latency chart in minutes. A pretraining hypothesis takes hours-to-days.
- Serving teaches the internals anyway. You cannot tune KV cache, prefix caching, chunked prefill, tensor/expert parallelism, or speculative decoding without actually understanding attention, the transformer memory layout, and the prefill/decode split. You'll learn the architecture because you need it, which sticks far better.
- But you do need one from-scratch build — timeboxed. Without it, you'll be a config-tuner with a ceiling. The fix is a 2–3 week "pretrain a tiny model end-to-end" sprint (Phase 2), not a 6-month one.
The shape of the path:
flowchart LR
P0[P0: Foundations<br/>GPU + transformer math] --> P1[P1: Serving & inference<br/>vLLM/SGLang deep]
P1 --> P2[P2: From scratch sprint<br/>nanochat, timeboxed]
P1 --> P3[P3: Compression<br/>quantization, spec-decode]
P2 --> P4[P4: Post-training<br/>SFT/LoRA -> DPO -> GRPO]
P3 --> P5[P5: Platform<br/>K8s, llm-d, autoscaling]
P4 --> P5
P5 --> P6[P6: Evals + Observability<br/>+ SLO ops]
P6 --> P7[P7: Specialize<br/>kernels / RL envs / multi-tenant]
1. Skill impact ranking (what actually pays, 2026 → 2029)¶
Ranked by (market demand × scarcity × durability). This should drive how you allocate time.
| # | Skill cluster | Impact | Why it lasts | Time to allocate |
|---|---|---|---|---|
| 1 | Inference engine mastery + GPU unit economics — vLLM/SGLang internals, KV cache management, continuous batching, chunked prefill, prefix caching, TTFT/ITL/throughput tradeoffs, cost-per-million-tokens modeling | ★★★★★ | Every self-hosting org needs it; models change, the memory-bandwidth physics doesn't | 25% |
| 2 | Kubernetes-native inference platform — InferencePool / Gateway API Inference Extension, llm-d, KServe, NVIDIA Dynamo, KV-aware routing, disaggregated prefill/decode, GPU operator, MIG, KEDA/HPA on LLM metrics | ★★★★★ | This is the 2025→2027 standardization wave; scarce skill, high salary | 20% |
| 3 | Model compression — FP8/NVFP4/MXFP4, AWQ/GPTQ/SmoothQuant, KV-cache quant, QAT, speculative decoding (EAGLE/MTP), distillation | ★★★★☆ | Directly converts to GPU cost reduction — the #1 exec ask | 15% |
| 4 | Post-training pipeline — SFT/LoRA/QLoRA → preference (DPO/KTO/ORPO) → RL (GRPO/RLOO/RLVR), reward modelling, multi-LoRA serving | ★★★★☆ | Moving from "fine-tune a chatbot" to "RL on verifiable rewards + agentic environments" — the current research→product frontier | 15% |
| 5 | Evals + observability + SLO operations — lm-eval-harness, task-specific eval design, LLM-as-judge calibration, OTel GenAI semconv, Langfuse, Prometheus/Grafana on engine metrics, regression gating | ★★★★☆ | Massively underrated; the thing that makes everything else trustworthy. Nearly everyone is weak here | 10% |
| 6 | Data engineering for post-training — dataset curation, dedup, decontamination, synthetic data generation, filtering, packing | ★★★★☆ | Quality of a fine-tune is 80% data; least glamorous, highest real leverage | 8% |
| 7 | Pretraining / distributed training at scale — FSDP2, TP/PP/CP/EP, torchtitan, DeepSpeed, MFU optimization | ★★★☆☆ | Low job count, but conceptual payoff is huge and it's the credential for "architect" roles | 5% |
| 8 | Kernel-level work — Triton, CUDA, FlashAttention variants, custom fused MoE | ★★★☆☆ (★★★★★ if you go deep) | Small market, extremely high pay, steep curve. A specialization, not a starting point | 2% (Phase 7 branch) |
Emerging bets worth tracking (2026–2029): disaggregated serving as default; KV-cache as a first-class distributed tier (offload to CPU/SSD/network); SLO-aware / latency-predictive routing; agentic RL environments as a product category; on-prem sovereign/regulated deployments; MoE-specific serving (expert parallelism, expert-load balancing); hybrid attention/SSM models changing KV assumptions.
2. Phase-by-phase plan¶
Phase 0 — Foundations you can't skip (Weeks 1–2)¶
Not a course. A checklist. If you can already answer these cold, skip.
GPU & systems literacy
- GPU memory hierarchy (HBM vs SRAM), why LLM decode is memory-bandwidth bound and prefill is compute bound
- Arithmetic intensity, roofline model, MFU (model FLOPs utilization)
- nvidia-smi, DCGM, NVLink vs PCIe, NCCL basics, CUDA/driver/toolkit version hell
- Precision formats: fp32 / tf32 / bf16 / fp16 / fp8(e4m3,e5m2) / int8 / int4 / nvfp4 / mxfp4
Transformer internals (from the inference angle)
- Where the parameters live: params ≈ 12·d²·L for dense; KV cache size = 2 · L · n_kv_heads · d_head · seq_len · batch · bytes_per_elem
- MHA vs GQA vs MQA vs MLA — and why each is a KV-cache decision
- Prefill vs decode phase; why they have opposite hardware profiles
- RoPE, RMSNorm, SwiGLU, MoE routing (top-k experts, shared experts)
Do: 1. Write a spreadsheet: for a model you own, compute weights VRAM, KV cache per 1k tokens, max concurrent sequences at 80GB. Verify against reality later. 2. Implement attention + KV cache in ~100 lines of raw PyTorch, no libraries.
Resources: Karpathy "Let's build GPT" + nanoGPT; the vLLM PagedAttention paper (arXiv 2309.06180); the Anyscale continuous-batching post; GPU MODE lecture series.
Phase 1 — Inference & serving, deep (Weeks 3–8) ← the core investment¶
This is where you spend the most time and where expertise compounds.
1.1 Get fluent with vLLM (weeks 3–4)
- Serve a 7–14B model with the OpenAI-compatible server. Then a 30B+ MoE.
- Systematically learn each knob and measure its effect:
--gpu-memory-utilization, --max-model-len, --max-num-seqs, --max-num-batched-tokens, --enable-prefix-caching, chunked prefill, --tensor-parallel-size, --pipeline-parallel-size, --data-parallel-size, CUDA graph / compile levels, --enable-lora + multi-LoRA
- Structured output (xgrammar/guidance), tool-calling parsers, reasoning parsers
- Read the design docs, not just the user guide: Architecture Overview, Paged Attention, Automatic Prefix Caching, Hybrid KV Cache Manager, Metrics, torch.compile integration.
1.2 Benchmark like an engineer (week 5)
- Use vllm bench serve / GuideLLM to build a repeatable harness.
- Learn the metric vocabulary and never say "it's fast" again: TTFT, TPOT/ITL, end-to-end latency, output tokens/s per user, total throughput, goodput (requests meeting SLO), queue time, KV cache hit rate, GPU utilization, MFU.
- Sweep: input/output length distributions × concurrency × batch settings. Produce latency-vs-throughput Pareto curves.
- Build a cost model: $/GPU-hr → $/1M input tokens and $/1M output tokens at a given SLO. This one artifact makes you credible with leadership.
1.3 Compare engines (week 6) - SGLang — RadixAttention prefix caching, structured output, strong on agentic/multi-turn reuse. - TensorRT-LLM — best raw NVIDIA perf, worst ergonomics; know when it's worth it. - llama.cpp / GGUF — CPU/edge/consumer, quant format ecosystem. - Hugging Face TGI, LMDeploy, MLC — know they exist and their niche. - Deliverable: a written comparison on your hardware with your workload shape. Opinionated, with numbers.
1.4 Advanced serving concepts (weeks 7–8) - Prefix caching / KV reuse — huge for RAG and agents with long shared system prompts. Measure the hit-rate → TTFT relationship. - Speculative decoding — n-gram, EAGLE, MTP/multi-token-prediction. Measure acceptance rate vs speedup; know when it hurts (high batch sizes). - Parallelism strategies — TP (intra-layer, needs NVLink), PP (cross-node, bubble cost), DP (replicas), EP (MoE experts), CP (long context). Know which to reach for and why. - Disaggregated prefill/decode — separate prefill and decode workers, transfer KV between them (NIXL/Mooncake/LMCache connectors). Understand why: prefill and decode have opposite hardware profiles and interfere with each other in a shared batch. This is becoming the default production architecture. - KV offloading — tiering KV cache to CPU RAM / SSD / remote store.
Capstone P1: A benchmarking repo + a written architecture decision record: "For workload X on hardware Y, use engine Z with config C, delivering TTFT p95 = …, throughput = …, at $/1M tokens = …"
Phase 2 — The from-scratch sprint (Weeks 9–11) ← timeboxed, do not overrun¶
Purpose: destroy the black box. Not to produce a useful model.
- Run
karpathy/nanochatend to end on your GPUs: tokenizer training → pretraining → SFT → RL → eval → inference. It's a single-node, minimal, readable codebase covering every stage. A GPT-2-capability model is ~2 hrs on 8×H100 (proportionally longer on smaller boxes; scale down--depth). - Then modify it: change the attention variant, swap the optimizer (Muon vs AdamW), change the data mixture, and observe
val_bpb/ CORE score / MFU move. - Read
nanochat/engine.py(KV-cache inference) after having read vLLM's — the contrast is the lesson. - Skim scaling laws (Chinchilla, DCLM/FineWeb data papers) so you can reason about compute-optimal sizing.
Then look at pytorch/torchtitan to see what production-grade pretraining looks like: FSDP2 per-parameter sharding, TP/PP/CP/EP composition, async TP, activation checkpointing, distributed checkpointing, float8/MXFP8 training, torch.compile. You are reading and running, not mastering.
Exit criteria: you can draw the full lifecycle on a whiteboard and explain every box. Move on.
Phase 3 — Compression & optimization (Weeks 12–14)¶
The highest-ROI-per-hour skill after serving.
llm-compressor(vLLM project) — the standard open path from HF checkpoint → quantized, vLLM-deployable model.- Algorithms: RTN, GPTQ, AWQ, SmoothQuant, SpinQuant, QuIP, AutoRound
- Schemes: W4A16, W8A8-INT8, W8A8-FP8 (Ada+), MXFP8 / NVFP4 (Blackwell), W4AFP8 (Hopper+), FP8 KV cache
- Learn calibration-set design — this is where people silently destroy accuracy.
- KV cache quantization — often the biggest win for long-context serving.
- QAT (int8/int4/FP8/NVFP4) when PTQ loses too much — via Axolotl or torchao.
- Distillation — GKD / MiniLLM (in TRL) to make a small model imitate a large one. Increasingly the right answer over "fine-tune the big one".
- Rule you must internalize: never ship a quantized model without an accuracy delta report on task-relevant evals. Perplexity is not enough.
Capstone P3: Take one model, produce FP8 / W4A16 / NVFP4 variants; publish a table of accuracy delta × latency × throughput × VRAM × $/1M tokens. This artifact alone is portfolio-grade.
Phase 4 — Post-training / custom models (Weeks 15–19)¶
"Building our own custom LLM" in practice = continued pretraining + post-training on open weights, not pretraining from zero. Frame it that way.
4.1 Data first (week 15) — resist skipping this. - Dataset formats (chat/ShareGPT/instruct), templating, loss masking (assistant-only), packing/multipack - Dedup, decontamination against your eval sets, quality filtering - Synthetic data generation with a strong teacher model + rejection sampling - Build a small, excellent dataset rather than a large mediocre one
4.2 Supervised fine-tuning (weeks 16–17)
- LoRA / QLoRA / DoRA — rank, alpha, target modules, merge-back, when full FT wins
- Frameworks: Axolotl (YAML-driven, broadest features), HF TRL (SFTTrainer, most canonical), Unsloth (fastest single-GPU), torchtune
- Efficiency: Flash Attention 2/3, Liger kernels, cut cross-entropy, gradient checkpointing, FSDP2/DeepSpeed ZeRO-3, sequence parallelism
- Multi-LoRA: train several adapters, then serve them all from one vLLM instance with dynamic adapter loading. This is the killer cost pattern for multi-tenant.
4.3 Preference & RL (weeks 18–19) - Offline preference: DPO, KTO, ORPO, CPO, SimPO — know the tradeoffs, DPO/KTO are the workhorses - Reward modelling: outcome RM and process RM (PRM) - Online RL: GRPO (the DeepSeek-R1 method, now the default), RLOO, PPO, Online DPO - RLVR — RL with verifiable rewards (math, code, unit tests, tool-call correctness). This is the highest-signal current technique and where "custom model" work is heading. - Serving-in-the-loop: TRL/Axolotl co-locate vLLM for generation during RL — understand this architecture, it's where training and inference infra merge. - Scale-out: verl, OpenRLHF, NeMo-RL for multi-node RL. - Emerging: agentic RL with environments (OpenEnv/Harbor) — per-example environments with environment-owned rewards. Track this closely.
Capstone P4: Take a real task you care about. Build the dataset → SFT with LoRA → DPO pass → GRPO with a verifiable reward → quantize → serve with multi-LoRA → eval against base. Document every delta. This is a complete, hireable portfolio project.
Phase 5 — Platform: Kubernetes & production hosting (Weeks 20–23)¶
Only now does Kubernetes make sense — you now know what you're orchestrating.
5.1 GPU-on-Kubernetes fundamentals
- NVIDIA GPU Operator, device plugin, DCGM exporter, node feature discovery
- GPU sharing: MIG, time-slicing, MPS — and when each is a bad idea for LLMs
- Topology-aware scheduling, node pools, taints/tolerations, huge model image/weight loading (this is a real bottleneck — look at model caching, run:ai streamer, CSI-backed weight volumes)
- Gang scheduling for multi-node (Kueue, Volcano, LeaderWorkerSet)
5.2 Inference platforms — learn the layers
- Gateway API Inference Extension (official Kubernetes, SIG-Network WG-Serving) — the emerging standard. Key concepts: InferencePool ("LLM-optimized Service"), Endpoint Picker (EPP) via Envoy ext-proc, model-aware routing, serving priority, model rollouts by traffic split, LoRA-adapter-aware routing.
- llm-d (CNCF sandbox, Red Hat/Google/IBM) — the production reference architecture on top of it: Router (Envoy proxy + EPP), InferencePool with prefill/decode variants, prefix-cache-aware routing, KV-cache indexing, KV offloading tiers, disaggregated serving, latency-predictor-based routing, batch gateway, and autoscaling via HPA/KEDA or a Workload Variant Autoscaler.
- NVIDIA Dynamo — the NVIDIA equivalent: KV-aware router, disaggregated prefill/decode, planner-based scaling, NIXL transfer library.
- KServe — CRD-driven InferenceService, canary rollouts, scale-to-zero; broader ML scope.
- Ray Serve LLM — Python-native, good for composite pipelines and mixed workloads.
- AI gateways — LiteLLM / Envoy AI Gateway / kgateway for multi-provider routing, budgets, rate limits, virtual keys.
5.3 Autoscaling & reliability - Scale on the right signals: queue depth, KV cache utilization, TTFT — not CPU or GPU%. This is the single most common production mistake. - Cold start mitigation: preloaded weights, warm pools, scale-to-zero economics - Rollouts: canary by model name, shadow traffic, instant rollback - SkyPilot for burst-to-cloud / spot GPU orchestration when your own GPUs are saturated
Capstone P5: A working cluster serving 2+ models with an inference gateway, prefix-cache-aware routing, KEDA autoscaling on queue depth, canary rollout, and a Grafana dashboard. Bonus: disaggregated prefill/decode.
Phase 6 — Evals, observability, safety, governance (Weeks 24–26)¶
Run this concurrently from Phase 3 onward, then formalize here.
Evaluation
- Harnesses: lm-evaluation-harness, HELM, evalchemy; task-specific suites
- Build your own domain eval set — public benchmarks are contaminated and rarely predict your use case
- LLM-as-judge: rubric design, position/verbosity bias, judge calibration against human labels, agreement metrics
- Regression gating in CI: no model or config ships without passing the eval suite
- Eval for agents/RAG: tool-call accuracy, trajectory-level scoring, retrieval quality separated from generation quality
Observability (two distinct layers — don't conflate them)
1. Infrastructure/serving layer — vLLM Prometheus metrics (TTFT, TPOT, running/waiting requests, KV cache usage, preemptions), DCGM GPU metrics, Grafana dashboards, alerting on SLO burn
2. Application/semantic layer — OpenTelemetry GenAI semantic conventions (now in the dedicated semantic-conventions-genai repo) as the vendor-neutral standard; Langfuse (open-source, self-hostable, OTel-based) for traces, sessions, prompt management/versioning, datasets, experiments, online LLM-as-judge scoring; OpenLLMetry/OpenLIT as alternatives
Safety & governance
- Guardrails: input/output filtering, PII redaction, jailbreak & prompt-injection detection (Llama Guard, NeMo Guardrails, Granite Guardian)
- OWASP Top 10 for LLM Applications — read it properly
- Model provenance: licences (Apache-2.0 vs Llama Community vs research-only), model cards, SBOM, supply-chain (never trust_remote_code blindly; prefer safetensors)
- EU AI Act obligations for deployers/providers if that applies to you
Capstone P6: An end-to-end golden path: request → gateway → engine → response, with correlated OTel traces, engine metrics on Grafana, automated eval regression gate in CI, and a guardrail layer. Deliberately break things (OOM, preemption storm, bad quant, cache thrash) and write the runbook.
Phase 7 — Specialize (Months 7–12)¶
Pick one or two. Generalists plateau here; specialists don't.
| Branch | What you learn | Who needs it |
|---|---|---|
| Kernels & compilers | Triton, CUDA, FlashAttention variants, fused MoE kernels, CUTLASS/CuTeDSL, torch.compile internals, custom vLLM attention backends |
Engine vendors, HFT-like latency shops. Highest pay ceiling. |
| Distributed training at scale | torchtitan/Megatron/DeepSpeed, 5D parallelism, MFU tuning, fault-tolerant training (TorchFT), multi-node debugging | Labs, national/sovereign AI programs |
| Agentic RL & environments | GRPO at scale with verl, environment design, reward hacking mitigation, tool-use RL, long-horizon credit assignment | Frontier product teams. Fastest-growing. |
| Serving platform engineering | Contribute to vLLM / llm-d / Gateway API Inference Ext.; build KV connectors, EPP scorers, custom schedulers | Platform teams; strongest public credential |
| Multi-tenant / cost engineering | Fair-share scheduling, priority tiers, chargeback, capacity planning, MoE expert-load balancing, spot-instance strategies | Enterprise platform / FinOps-for-GPU |
Highest-leverage move in this phase: contribute upstream. A merged vLLM or llm-d PR is worth more than any certificate.
3. Sequencing summary¶
| Weeks | Phase | Primary output |
|---|---|---|
| 1–2 | Foundations | KV-cache/VRAM math sheet; hand-rolled attention |
| 3–8 | Inference & serving | Benchmark harness + engine comparison + cost model |
| 9–11 | From-scratch sprint (timeboxed) | Own GPT-2-class model, full lifecycle |
| 12–14 | Compression | Quantization tradeoff report |
| 15–19 | Post-training | SFT→DPO→GRPO custom model, multi-LoRA served |
| 20–23 | Kubernetes platform | Inference gateway + autoscaling + canary |
| 24–26 | Evals & observability | Golden path + CI eval gate + runbook |
| 27–52 | Specialization | Upstream contributions / deep branch |
4. Rules of engagement¶
- Measure everything. A claim without a p95 number is not an engineering claim. Every phase produces a chart.
- One model family, deep. Pick one (e.g. Qwen3.x or GLM or Llama) and know it intimately rather than sampling twenty.
- Keep a lab notebook in git. Every experiment: hypothesis, config diff, result, conclusion. This becomes your portfolio and your promotion case.
- Timebox the shiny things. The field ships weekly. Read release notes, not every paper. Fundamentals (memory bandwidth, batching, KV cache) don't churn — chase those.
- Reproduce before you innovate. Reproduce a published benchmark number on your hardware first. The gap you find is the learning.
- Build the boring layer. Config management, weight caching, image builds, eval CI. This is what separates a hobbyist from a platform engineer.
- Teach it. Given the workspace this lives in — turn each capstone into a training module. Teaching forces the gaps into the open.
5. Curated resource list¶
Inference / serving - vLLM docs — especially Design Documents: Architecture Overview, Paged Attention, Automatic Prefix Caching, Hybrid KV Cache Manager, Metrics, Optimization Levels - vLLM paper (SOSP 2023, arXiv 2309.06180); vLLM blog; vLLM Office Hours (recorded) - SGLang docs + RadixAttention paper - Anyscale: "How continuous batching enables 23x throughput…" - NVIDIA Dynamo design docs; TensorRT-LLM docs
Kubernetes / platform
- Kubernetes Gateway API Inference Extension docs (gateway-api-inference-extension.sigs.k8s.io) — API overview, roles & personas, request flow
- llm-d docs — Architecture (Router/EPP, InferencePool, Model Server) and Well-Lit Paths
- KServe docs; Ray Serve LLM docs; SkyPilot docs
- Kubernetes WG-Serving working group notes
Compression
- llm-compressor docs (compression schemes, observers, memory requirements, deploying with vLLM)
- Papers: GPTQ, AWQ, SmoothQuant, SpinQuant, EAGLE, Medusa/MTP
Training / post-training
- karpathy/nanochat + its Discussions (the "Beating GPT-2 for <$100" write-up is excellent)
- pytorch/torchtitan + the TorchTitan ICLR 2025 paper
- HF TRL docs (trainer taxonomy: SFT/DPO/KTO/GRPO/RLOO/PRM/GKD) + HF smol-course
- Axolotl docs (config reference, multipack, multi-GPU/multi-node, axolotl agent-docs)
- Unsloth docs; verl docs; Open-R1
- Papers: LoRA, QLoRA, DPO, DeepSeek-R1 (GRPO), Tulu 3 (post-training recipe transparency)
Evals & observability
- OpenTelemetry GenAI semantic conventions (open-telemetry/semantic-conventions-genai)
- Langfuse docs (observability, prompt management, evaluation, datasets/experiments)
- EleutherAI lm-evaluation-harness; HELM
- OWASP Top 10 for LLM Applications
Communities worth being in - vLLM Slack/forum, llm-d Slack, GPU MODE Discord, EleutherAI Discord, Hugging Face forums, r/LocalLLaMA (noisy but fast signal)
6. What "expert" looks like at the end¶
You can walk into a room and: - Size hardware for a workload from first principles, and justify it with a $/1M-token model - Choose between vLLM / SGLang / TRT-LLM and defend it with your own measurements - Decide dense vs MoE, quant scheme, TP/PP/EP layout, and disaggregate-or-not — with reasons - Take a business task → dataset → SFT → preference → RL → quantized → multi-LoRA served → evaluated → monitored, unaided - Design the K8s serving topology with cache-aware routing, correct autoscaling signals, and safe rollouts - Explain every layer down to the KV cache block, because you built a model from scratch once