Skip to content

Open-Source LLM Infrastructure: Expert Learning Path

Starting point: competent LLM user (RAG, agents, prompt/tool orchestration). Target: person who can custom-build, host, serve, optimize, monitor and post-train open-weight models on owned GPUs — and defend the architecture decisions. Assumed effort: 12–15 hrs/week. Horizon: 6 months to strong practitioner, 12 months to differentiated expert.


0. The strategic answer first

"Should I learn/build from scratch, or jump straight into hosting, inference, parallelism, Kubernetes?"

Do not start from scratch. Start at the serving layer and spiral outward. Reasons:

  1. Serving is where the money and the jobs are. Pretraining a frontier model is a ~10-lab activity. Serving, optimizing, and post-training open weights is a ~every-enterprise activity. The scarce skill in 2026 is "make Qwen/Llama/GLM/DeepSpeek-class weights run at target latency and cost on our GPUs, safely, and keep them there."
  2. Serving gives fastest feedback loops. A vLLM config change shows up in a p99 latency chart in minutes. A pretraining hypothesis takes hours-to-days.
  3. Serving teaches the internals anyway. You cannot tune KV cache, prefix caching, chunked prefill, tensor/expert parallelism, or speculative decoding without actually understanding attention, the transformer memory layout, and the prefill/decode split. You'll learn the architecture because you need it, which sticks far better.
  4. But you do need one from-scratch build — timeboxed. Without it, you'll be a config-tuner with a ceiling. The fix is a 2–3 week "pretrain a tiny model end-to-end" sprint (Phase 2), not a 6-month one.

The shape of the path:

flowchart LR
    P0[P0: Foundations<br/>GPU + transformer math] --> P1[P1: Serving & inference<br/>vLLM/SGLang deep]
    P1 --> P2[P2: From scratch sprint<br/>nanochat, timeboxed]
    P1 --> P3[P3: Compression<br/>quantization, spec-decode]
    P2 --> P4[P4: Post-training<br/>SFT/LoRA -> DPO -> GRPO]
    P3 --> P5[P5: Platform<br/>K8s, llm-d, autoscaling]
    P4 --> P5
    P5 --> P6[P6: Evals + Observability<br/>+ SLO ops]
    P6 --> P7[P7: Specialize<br/>kernels / RL envs / multi-tenant]

1. Skill impact ranking (what actually pays, 2026 → 2029)

Ranked by (market demand × scarcity × durability). This should drive how you allocate time.

# Skill cluster Impact Why it lasts Time to allocate
1 Inference engine mastery + GPU unit economics — vLLM/SGLang internals, KV cache management, continuous batching, chunked prefill, prefix caching, TTFT/ITL/throughput tradeoffs, cost-per-million-tokens modeling ★★★★★ Every self-hosting org needs it; models change, the memory-bandwidth physics doesn't 25%
2 Kubernetes-native inference platform — InferencePool / Gateway API Inference Extension, llm-d, KServe, NVIDIA Dynamo, KV-aware routing, disaggregated prefill/decode, GPU operator, MIG, KEDA/HPA on LLM metrics ★★★★★ This is the 2025→2027 standardization wave; scarce skill, high salary 20%
3 Model compression — FP8/NVFP4/MXFP4, AWQ/GPTQ/SmoothQuant, KV-cache quant, QAT, speculative decoding (EAGLE/MTP), distillation ★★★★☆ Directly converts to GPU cost reduction — the #1 exec ask 15%
4 Post-training pipeline — SFT/LoRA/QLoRA → preference (DPO/KTO/ORPO) → RL (GRPO/RLOO/RLVR), reward modelling, multi-LoRA serving ★★★★☆ Moving from "fine-tune a chatbot" to "RL on verifiable rewards + agentic environments" — the current research→product frontier 15%
5 Evals + observability + SLO operations — lm-eval-harness, task-specific eval design, LLM-as-judge calibration, OTel GenAI semconv, Langfuse, Prometheus/Grafana on engine metrics, regression gating ★★★★☆ Massively underrated; the thing that makes everything else trustworthy. Nearly everyone is weak here 10%
6 Data engineering for post-training — dataset curation, dedup, decontamination, synthetic data generation, filtering, packing ★★★★☆ Quality of a fine-tune is 80% data; least glamorous, highest real leverage 8%
7 Pretraining / distributed training at scale — FSDP2, TP/PP/CP/EP, torchtitan, DeepSpeed, MFU optimization ★★★☆☆ Low job count, but conceptual payoff is huge and it's the credential for "architect" roles 5%
8 Kernel-level work — Triton, CUDA, FlashAttention variants, custom fused MoE ★★★☆☆ (★★★★★ if you go deep) Small market, extremely high pay, steep curve. A specialization, not a starting point 2% (Phase 7 branch)

Emerging bets worth tracking (2026–2029): disaggregated serving as default; KV-cache as a first-class distributed tier (offload to CPU/SSD/network); SLO-aware / latency-predictive routing; agentic RL environments as a product category; on-prem sovereign/regulated deployments; MoE-specific serving (expert parallelism, expert-load balancing); hybrid attention/SSM models changing KV assumptions.


2. Phase-by-phase plan

Phase 0 — Foundations you can't skip (Weeks 1–2)

Not a course. A checklist. If you can already answer these cold, skip.

GPU & systems literacy - GPU memory hierarchy (HBM vs SRAM), why LLM decode is memory-bandwidth bound and prefill is compute bound - Arithmetic intensity, roofline model, MFU (model FLOPs utilization) - nvidia-smi, DCGM, NVLink vs PCIe, NCCL basics, CUDA/driver/toolkit version hell - Precision formats: fp32 / tf32 / bf16 / fp16 / fp8(e4m3,e5m2) / int8 / int4 / nvfp4 / mxfp4

Transformer internals (from the inference angle) - Where the parameters live: params ≈ 12·d²·L for dense; KV cache size = 2 · L · n_kv_heads · d_head · seq_len · batch · bytes_per_elem - MHA vs GQA vs MQA vs MLA — and why each is a KV-cache decision - Prefill vs decode phase; why they have opposite hardware profiles - RoPE, RMSNorm, SwiGLU, MoE routing (top-k experts, shared experts)

Do: 1. Write a spreadsheet: for a model you own, compute weights VRAM, KV cache per 1k tokens, max concurrent sequences at 80GB. Verify against reality later. 2. Implement attention + KV cache in ~100 lines of raw PyTorch, no libraries.

Resources: Karpathy "Let's build GPT" + nanoGPT; the vLLM PagedAttention paper (arXiv 2309.06180); the Anyscale continuous-batching post; GPU MODE lecture series.


Phase 1 — Inference & serving, deep (Weeks 3–8) ← the core investment

This is where you spend the most time and where expertise compounds.

1.1 Get fluent with vLLM (weeks 3–4) - Serve a 7–14B model with the OpenAI-compatible server. Then a 30B+ MoE. - Systematically learn each knob and measure its effect: --gpu-memory-utilization, --max-model-len, --max-num-seqs, --max-num-batched-tokens, --enable-prefix-caching, chunked prefill, --tensor-parallel-size, --pipeline-parallel-size, --data-parallel-size, CUDA graph / compile levels, --enable-lora + multi-LoRA - Structured output (xgrammar/guidance), tool-calling parsers, reasoning parsers - Read the design docs, not just the user guide: Architecture Overview, Paged Attention, Automatic Prefix Caching, Hybrid KV Cache Manager, Metrics, torch.compile integration.

1.2 Benchmark like an engineer (week 5) - Use vllm bench serve / GuideLLM to build a repeatable harness. - Learn the metric vocabulary and never say "it's fast" again: TTFT, TPOT/ITL, end-to-end latency, output tokens/s per user, total throughput, goodput (requests meeting SLO), queue time, KV cache hit rate, GPU utilization, MFU. - Sweep: input/output length distributions × concurrency × batch settings. Produce latency-vs-throughput Pareto curves. - Build a cost model: $/GPU-hr → $/1M input tokens and $/1M output tokens at a given SLO. This one artifact makes you credible with leadership.

1.3 Compare engines (week 6) - SGLang — RadixAttention prefix caching, structured output, strong on agentic/multi-turn reuse. - TensorRT-LLM — best raw NVIDIA perf, worst ergonomics; know when it's worth it. - llama.cpp / GGUF — CPU/edge/consumer, quant format ecosystem. - Hugging Face TGI, LMDeploy, MLC — know they exist and their niche. - Deliverable: a written comparison on your hardware with your workload shape. Opinionated, with numbers.

1.4 Advanced serving concepts (weeks 7–8) - Prefix caching / KV reuse — huge for RAG and agents with long shared system prompts. Measure the hit-rate → TTFT relationship. - Speculative decoding — n-gram, EAGLE, MTP/multi-token-prediction. Measure acceptance rate vs speedup; know when it hurts (high batch sizes). - Parallelism strategies — TP (intra-layer, needs NVLink), PP (cross-node, bubble cost), DP (replicas), EP (MoE experts), CP (long context). Know which to reach for and why. - Disaggregated prefill/decode — separate prefill and decode workers, transfer KV between them (NIXL/Mooncake/LMCache connectors). Understand why: prefill and decode have opposite hardware profiles and interfere with each other in a shared batch. This is becoming the default production architecture. - KV offloading — tiering KV cache to CPU RAM / SSD / remote store.

Capstone P1: A benchmarking repo + a written architecture decision record: "For workload X on hardware Y, use engine Z with config C, delivering TTFT p95 = …, throughput = …, at $/1M tokens = …"


Phase 2 — The from-scratch sprint (Weeks 9–11) ← timeboxed, do not overrun

Purpose: destroy the black box. Not to produce a useful model.

  • Run karpathy/nanochat end to end on your GPUs: tokenizer training → pretraining → SFT → RL → eval → inference. It's a single-node, minimal, readable codebase covering every stage. A GPT-2-capability model is ~2 hrs on 8×H100 (proportionally longer on smaller boxes; scale down --depth).
  • Then modify it: change the attention variant, swap the optimizer (Muon vs AdamW), change the data mixture, and observe val_bpb / CORE score / MFU move.
  • Read nanochat/engine.py (KV-cache inference) after having read vLLM's — the contrast is the lesson.
  • Skim scaling laws (Chinchilla, DCLM/FineWeb data papers) so you can reason about compute-optimal sizing.

Then look at pytorch/torchtitan to see what production-grade pretraining looks like: FSDP2 per-parameter sharding, TP/PP/CP/EP composition, async TP, activation checkpointing, distributed checkpointing, float8/MXFP8 training, torch.compile. You are reading and running, not mastering.

Exit criteria: you can draw the full lifecycle on a whiteboard and explain every box. Move on.


Phase 3 — Compression & optimization (Weeks 12–14)

The highest-ROI-per-hour skill after serving.

  • llm-compressor (vLLM project) — the standard open path from HF checkpoint → quantized, vLLM-deployable model.
  • Algorithms: RTN, GPTQ, AWQ, SmoothQuant, SpinQuant, QuIP, AutoRound
  • Schemes: W4A16, W8A8-INT8, W8A8-FP8 (Ada+), MXFP8 / NVFP4 (Blackwell), W4AFP8 (Hopper+), FP8 KV cache
  • Learn calibration-set design — this is where people silently destroy accuracy.
  • KV cache quantization — often the biggest win for long-context serving.
  • QAT (int8/int4/FP8/NVFP4) when PTQ loses too much — via Axolotl or torchao.
  • Distillation — GKD / MiniLLM (in TRL) to make a small model imitate a large one. Increasingly the right answer over "fine-tune the big one".
  • Rule you must internalize: never ship a quantized model without an accuracy delta report on task-relevant evals. Perplexity is not enough.

Capstone P3: Take one model, produce FP8 / W4A16 / NVFP4 variants; publish a table of accuracy delta × latency × throughput × VRAM × $/1M tokens. This artifact alone is portfolio-grade.


Phase 4 — Post-training / custom models (Weeks 15–19)

"Building our own custom LLM" in practice = continued pretraining + post-training on open weights, not pretraining from zero. Frame it that way.

4.1 Data first (week 15) — resist skipping this. - Dataset formats (chat/ShareGPT/instruct), templating, loss masking (assistant-only), packing/multipack - Dedup, decontamination against your eval sets, quality filtering - Synthetic data generation with a strong teacher model + rejection sampling - Build a small, excellent dataset rather than a large mediocre one

4.2 Supervised fine-tuning (weeks 16–17) - LoRA / QLoRA / DoRA — rank, alpha, target modules, merge-back, when full FT wins - Frameworks: Axolotl (YAML-driven, broadest features), HF TRL (SFTTrainer, most canonical), Unsloth (fastest single-GPU), torchtune - Efficiency: Flash Attention 2/3, Liger kernels, cut cross-entropy, gradient checkpointing, FSDP2/DeepSpeed ZeRO-3, sequence parallelism - Multi-LoRA: train several adapters, then serve them all from one vLLM instance with dynamic adapter loading. This is the killer cost pattern for multi-tenant.

4.3 Preference & RL (weeks 18–19) - Offline preference: DPO, KTO, ORPO, CPO, SimPO — know the tradeoffs, DPO/KTO are the workhorses - Reward modelling: outcome RM and process RM (PRM) - Online RL: GRPO (the DeepSeek-R1 method, now the default), RLOO, PPO, Online DPO - RLVR — RL with verifiable rewards (math, code, unit tests, tool-call correctness). This is the highest-signal current technique and where "custom model" work is heading. - Serving-in-the-loop: TRL/Axolotl co-locate vLLM for generation during RL — understand this architecture, it's where training and inference infra merge. - Scale-out: verl, OpenRLHF, NeMo-RL for multi-node RL. - Emerging: agentic RL with environments (OpenEnv/Harbor) — per-example environments with environment-owned rewards. Track this closely.

Capstone P4: Take a real task you care about. Build the dataset → SFT with LoRA → DPO pass → GRPO with a verifiable reward → quantize → serve with multi-LoRA → eval against base. Document every delta. This is a complete, hireable portfolio project.


Phase 5 — Platform: Kubernetes & production hosting (Weeks 20–23)

Only now does Kubernetes make sense — you now know what you're orchestrating.

5.1 GPU-on-Kubernetes fundamentals - NVIDIA GPU Operator, device plugin, DCGM exporter, node feature discovery - GPU sharing: MIG, time-slicing, MPS — and when each is a bad idea for LLMs - Topology-aware scheduling, node pools, taints/tolerations, huge model image/weight loading (this is a real bottleneck — look at model caching, run:ai streamer, CSI-backed weight volumes) - Gang scheduling for multi-node (Kueue, Volcano, LeaderWorkerSet)

5.2 Inference platforms — learn the layers - Gateway API Inference Extension (official Kubernetes, SIG-Network WG-Serving) — the emerging standard. Key concepts: InferencePool ("LLM-optimized Service"), Endpoint Picker (EPP) via Envoy ext-proc, model-aware routing, serving priority, model rollouts by traffic split, LoRA-adapter-aware routing. - llm-d (CNCF sandbox, Red Hat/Google/IBM) — the production reference architecture on top of it: Router (Envoy proxy + EPP), InferencePool with prefill/decode variants, prefix-cache-aware routing, KV-cache indexing, KV offloading tiers, disaggregated serving, latency-predictor-based routing, batch gateway, and autoscaling via HPA/KEDA or a Workload Variant Autoscaler. - NVIDIA Dynamo — the NVIDIA equivalent: KV-aware router, disaggregated prefill/decode, planner-based scaling, NIXL transfer library. - KServe — CRD-driven InferenceService, canary rollouts, scale-to-zero; broader ML scope. - Ray Serve LLM — Python-native, good for composite pipelines and mixed workloads. - AI gateways — LiteLLM / Envoy AI Gateway / kgateway for multi-provider routing, budgets, rate limits, virtual keys.

5.3 Autoscaling & reliability - Scale on the right signals: queue depth, KV cache utilization, TTFT — not CPU or GPU%. This is the single most common production mistake. - Cold start mitigation: preloaded weights, warm pools, scale-to-zero economics - Rollouts: canary by model name, shadow traffic, instant rollback - SkyPilot for burst-to-cloud / spot GPU orchestration when your own GPUs are saturated

Capstone P5: A working cluster serving 2+ models with an inference gateway, prefix-cache-aware routing, KEDA autoscaling on queue depth, canary rollout, and a Grafana dashboard. Bonus: disaggregated prefill/decode.


Phase 6 — Evals, observability, safety, governance (Weeks 24–26)

Run this concurrently from Phase 3 onward, then formalize here.

Evaluation - Harnesses: lm-evaluation-harness, HELM, evalchemy; task-specific suites - Build your own domain eval set — public benchmarks are contaminated and rarely predict your use case - LLM-as-judge: rubric design, position/verbosity bias, judge calibration against human labels, agreement metrics - Regression gating in CI: no model or config ships without passing the eval suite - Eval for agents/RAG: tool-call accuracy, trajectory-level scoring, retrieval quality separated from generation quality

Observability (two distinct layers — don't conflate them) 1. Infrastructure/serving layer — vLLM Prometheus metrics (TTFT, TPOT, running/waiting requests, KV cache usage, preemptions), DCGM GPU metrics, Grafana dashboards, alerting on SLO burn 2. Application/semantic layer — OpenTelemetry GenAI semantic conventions (now in the dedicated semantic-conventions-genai repo) as the vendor-neutral standard; Langfuse (open-source, self-hostable, OTel-based) for traces, sessions, prompt management/versioning, datasets, experiments, online LLM-as-judge scoring; OpenLLMetry/OpenLIT as alternatives

Safety & governance - Guardrails: input/output filtering, PII redaction, jailbreak & prompt-injection detection (Llama Guard, NeMo Guardrails, Granite Guardian) - OWASP Top 10 for LLM Applications — read it properly - Model provenance: licences (Apache-2.0 vs Llama Community vs research-only), model cards, SBOM, supply-chain (never trust_remote_code blindly; prefer safetensors) - EU AI Act obligations for deployers/providers if that applies to you

Capstone P6: An end-to-end golden path: request → gateway → engine → response, with correlated OTel traces, engine metrics on Grafana, automated eval regression gate in CI, and a guardrail layer. Deliberately break things (OOM, preemption storm, bad quant, cache thrash) and write the runbook.


Phase 7 — Specialize (Months 7–12)

Pick one or two. Generalists plateau here; specialists don't.

Branch What you learn Who needs it
Kernels & compilers Triton, CUDA, FlashAttention variants, fused MoE kernels, CUTLASS/CuTeDSL, torch.compile internals, custom vLLM attention backends Engine vendors, HFT-like latency shops. Highest pay ceiling.
Distributed training at scale torchtitan/Megatron/DeepSpeed, 5D parallelism, MFU tuning, fault-tolerant training (TorchFT), multi-node debugging Labs, national/sovereign AI programs
Agentic RL & environments GRPO at scale with verl, environment design, reward hacking mitigation, tool-use RL, long-horizon credit assignment Frontier product teams. Fastest-growing.
Serving platform engineering Contribute to vLLM / llm-d / Gateway API Inference Ext.; build KV connectors, EPP scorers, custom schedulers Platform teams; strongest public credential
Multi-tenant / cost engineering Fair-share scheduling, priority tiers, chargeback, capacity planning, MoE expert-load balancing, spot-instance strategies Enterprise platform / FinOps-for-GPU

Highest-leverage move in this phase: contribute upstream. A merged vLLM or llm-d PR is worth more than any certificate.


3. Sequencing summary

Weeks Phase Primary output
1–2 Foundations KV-cache/VRAM math sheet; hand-rolled attention
3–8 Inference & serving Benchmark harness + engine comparison + cost model
9–11 From-scratch sprint (timeboxed) Own GPT-2-class model, full lifecycle
12–14 Compression Quantization tradeoff report
15–19 Post-training SFT→DPO→GRPO custom model, multi-LoRA served
20–23 Kubernetes platform Inference gateway + autoscaling + canary
24–26 Evals & observability Golden path + CI eval gate + runbook
27–52 Specialization Upstream contributions / deep branch

4. Rules of engagement

  1. Measure everything. A claim without a p95 number is not an engineering claim. Every phase produces a chart.
  2. One model family, deep. Pick one (e.g. Qwen3.x or GLM or Llama) and know it intimately rather than sampling twenty.
  3. Keep a lab notebook in git. Every experiment: hypothesis, config diff, result, conclusion. This becomes your portfolio and your promotion case.
  4. Timebox the shiny things. The field ships weekly. Read release notes, not every paper. Fundamentals (memory bandwidth, batching, KV cache) don't churn — chase those.
  5. Reproduce before you innovate. Reproduce a published benchmark number on your hardware first. The gap you find is the learning.
  6. Build the boring layer. Config management, weight caching, image builds, eval CI. This is what separates a hobbyist from a platform engineer.
  7. Teach it. Given the workspace this lives in — turn each capstone into a training module. Teaching forces the gaps into the open.

5. Curated resource list

Inference / serving - vLLM docs — especially Design Documents: Architecture Overview, Paged Attention, Automatic Prefix Caching, Hybrid KV Cache Manager, Metrics, Optimization Levels - vLLM paper (SOSP 2023, arXiv 2309.06180); vLLM blog; vLLM Office Hours (recorded) - SGLang docs + RadixAttention paper - Anyscale: "How continuous batching enables 23x throughput…" - NVIDIA Dynamo design docs; TensorRT-LLM docs

Kubernetes / platform - Kubernetes Gateway API Inference Extension docs (gateway-api-inference-extension.sigs.k8s.io) — API overview, roles & personas, request flow - llm-d docs — Architecture (Router/EPP, InferencePool, Model Server) and Well-Lit Paths - KServe docs; Ray Serve LLM docs; SkyPilot docs - Kubernetes WG-Serving working group notes

Compression - llm-compressor docs (compression schemes, observers, memory requirements, deploying with vLLM) - Papers: GPTQ, AWQ, SmoothQuant, SpinQuant, EAGLE, Medusa/MTP

Training / post-training - karpathy/nanochat + its Discussions (the "Beating GPT-2 for <$100" write-up is excellent) - pytorch/torchtitan + the TorchTitan ICLR 2025 paper - HF TRL docs (trainer taxonomy: SFT/DPO/KTO/GRPO/RLOO/PRM/GKD) + HF smol-course - Axolotl docs (config reference, multipack, multi-GPU/multi-node, axolotl agent-docs) - Unsloth docs; verl docs; Open-R1 - Papers: LoRA, QLoRA, DPO, DeepSeek-R1 (GRPO), Tulu 3 (post-training recipe transparency)

Evals & observability - OpenTelemetry GenAI semantic conventions (open-telemetry/semantic-conventions-genai) - Langfuse docs (observability, prompt management, evaluation, datasets/experiments) - EleutherAI lm-evaluation-harness; HELM - OWASP Top 10 for LLM Applications

Communities worth being in - vLLM Slack/forum, llm-d Slack, GPU MODE Discord, EleutherAI Discord, Hugging Face forums, r/LocalLLaMA (noisy but fast signal)


6. What "expert" looks like at the end

You can walk into a room and: - Size hardware for a workload from first principles, and justify it with a $/1M-token model - Choose between vLLM / SGLang / TRT-LLM and defend it with your own measurements - Decide dense vs MoE, quant scheme, TP/PP/EP layout, and disaggregate-or-not — with reasons - Take a business task → dataset → SFT → preference → RL → quantized → multi-LoRA served → evaluated → monitored, unaided - Design the K8s serving topology with cache-aware routing, correct autoscaling signals, and safe rollouts - Explain every layer down to the KV cache block, because you built a model from scratch once