Skip to content

Skill 09 — Specialization

Generalists plateau here. Specialists don't. Skills 00–08 make you a strong, hireable LLM infrastructure engineer. This skill makes you the person people are told to go talk to.

Duration 6 months (months 7–12), ongoing
Impact varies by branch — see the table
Prerequisites Skills 00–08 complete, with capstones done
Hardware branch-dependent
Status planned

How this skill differs

It has no fixed topic list, because the frontier moves. It has a structure for going deep and five branch maps. Pick one branch, optionally a second later. Depth beats breadth from here on.

Choose your branch

Branch Market size Pay ceiling Difficulty Pick this if…
A. Kernels & Compilers Small Highest Very high You enjoy assembly-level reasoning and profiler output
B. Serving Platform Engineering Large High High You liked Skills 01 and 06 most, and want the strongest public credential
C. Agentic RL & Environments Growing fastest High High You liked Skill 05 M4, and want to be early to the next wave
D. Distributed Training at Scale Small High Very high You want lab / sovereign-AI / national-program roles
E. Cost & Multi-Tenancy Engineering Large Medium-high Medium You like the economics and want to own a P&L line

Recommendation given the path so far: B or C. B compounds directly on your strongest skills and produces public artifacts. C positions you ahead of a wave that hasn't crested.


The specialization method (applies to every branch)

This is the actual content of the skill — a repeatable loop, run monthly.

flowchart LR
    A[Pick a real problem<br/>in your branch] --> B[Reproduce the<br/>state of the art]
    B --> C[Find the gap:<br/>where does it fail?]
    C --> D[Build something<br/>that closes it]
    D --> E[Measure honestly]
    E --> F[Publish:<br/>PR, post, or talk]
    F --> A
Phase Duration Output
1. Map the territory Weeks 1–3 Read the top 20 papers/RFCs in the branch; write a landscape doc naming the 5 open problems
2. Reproduce Weeks 4–8 Reproduce one significant published result on your hardware. Document every discrepancy
3. Find your gap Weeks 9–12 Where does the SOTA fail on a workload you care about? Characterize it with data
4. Contribute Weeks 13–20 Fix something real. Merged upstream PR, or a tool people use
5. Publish Weeks 21–24 A technical post with reproducible numbers, or a conference talk

The output of this skill is public artifacts, not private knowledge. A merged vLLM PR is worth more than any certificate.


Branch A — Kernels & Compilers

Area What to learn Milestone
Triton Language model, autotuning, block pointers, when Triton beats hand-written CUDA Write a fused kernel beating the PyTorch eager baseline
CUDA Memory coalescing, shared memory, warp primitives, occupancy, tensor core intrinsics Implement a tiled GEMM within 2× of cuBLAS
Attention kernels FlashAttention 1/2 tiling and recomputation, paged variants, MLA kernels, decode-optimized kernels Implement a paged attention decode kernel
MoE kernels Fused MoE, grouped GEMM, all-to-all overlap, expert-load handling Contribute to a fused MoE path
Compilers torch.compile internals, Inductor, custom ops, graph passes, CUDA graphs Add a fusion pass that measurably helps
Profiling Nsight Compute/Systems, roofline analysis, occupancy vs ILP Diagnose and fix a real kernel bottleneck

Note on our hardware: SM 8.0 means no FP8 or FA3 kernel work. Plenty remains — INT8/INT4 kernels, paged decode, MoE, and compiler passes are all fully available on Ampere.

Milestone artifact: a merged kernel or compiler PR in vLLM, SGLang, or PyTorch.

Sources: GPU MODE lectures and Discord · Triton docs and tutorials · CUDA C++ Programming Guide · FlashAttention papers · CUTLASS/CuTe docs · vllm/kernels/ and csrc/ source


Branch B — Serving Platform Engineering

Area What to learn Milestone
Engine internals vLLM V1 scheduler, KV cache manager, worker/executor split, attention backend interface Add a scheduler policy or a config knob
KV connectors NIXL, LMCache, Mooncake; the connector interface; transfer protocols and lease semantics Implement or improve a KV connector
Routing EPP scorers and filters, prefix-cache indexing, latency prediction, SLO-aware endpoint selection Write a custom EPP scorer, prove it beats the default
Standards Gateway API Inference Extension conformance, the model-server protocol, WG-Serving proposals Contribute to a proposal or conformance test
Autoscaling Workload-variant autoscaling, predictive scaling, cost-optimal replica placement Build an autoscaler that beats KEDA on a real workload
Operations Multi-cluster, disaster recovery, upgrade choreography for stateful GPU services A public runbook or operator

Milestone artifact: a merged PR in vLLM, llm-d, or gateway-api-inference-extension.

Sources: vLLM design docs and source · llm-d architecture docs and repos · Gateway API Inference Extension enhancements/proposals · NVIDIA Dynamo design docs · Kubernetes WG-Serving meetings


Branch C — Agentic RL & Environments

Area What to learn Milestone
RL at scale verl / OpenRLHF architectures, async rollout, weight sync, colocated vs disaggregated Run multi-node-simulated GRPO efficiently
Environment design Task specification, reward shaping, partial credit, sandboxing, determinism and seeding Publish an environment others use
Verifiable rewards Verifier design, gaming resistance, verification cost, hybrid judge+verifier rewards A verifier suite resistant to a documented attack set
Multi-turn credit Long-horizon assignment, turn-level vs trajectory-level rewards, tool-use RL Train a tool-using agent beating a prompted baseline
Reward hacking Detection, monitoring, mitigation, KL control, adversarial evaluation of your own rewards A published taxonomy of hacks you found
Inference-time scaling Best-of-N, self-consistency, search, and the RL-vs-inference-compute tradeoff Quantify where each is cost-optimal

Milestone artifact: a published open environment + a trained policy with reproducible results.

Sources: DeepSeek-R1 and Open-R1 · verl docs · TRL GRPO + OpenEnv/Harbor docs · Tulu 3 · process-supervision and RLVR literature · agent benchmark suites (SWE-bench, τ-bench, WebArena)


Branch D — Distributed Training at Scale

Area What to learn Milestone
Parallelism composition FSDP2 + TP + PP + CP + EP together; when each dimension pays Compose 3+ dimensions on a real model
Communication NCCL internals, overlap strategies, async TP, all-to-all for MoE Diagnose and fix a comms bottleneck
Memory Activation checkpointing policies, offload, optimizer-state sharding, memory estimation Predict peak memory within 10% before running
Reliability Distributed and async checkpointing, TorchFT, flight recorder, straggler and hang diagnosis Recover a multi-rank failure with no loss
Efficiency MFU optimization, profiling training, dataloader bottlenecks Improve MFU by ≥15% on a reference config
Low precision bf16 today; float8/MXFP8 as a decision framework for future hardware A written hardware-purchase analysis

Reality check: this branch is genuinely constrained by 5–6 A100s. You can master the techniques and read the code, but you can't demonstrate scale. Pick it only if the concepts themselves are the goal.

Milestone artifact: a merged torchtitan PR, or a reproducible efficiency study.

Sources: torchtitan source + ICLR paper · PyTorch distributed blog series · Megatron-LM and DeepSpeed papers · NCCL docs · the Llama 3 and DeepSeek-V3 infrastructure sections


Branch E — Cost & Multi-Tenancy Engineering

Area What to learn Milestone
Unit economics $/1M tokens modelling, utilization math, headroom cost, marginal vs average cost A cost model adopted by a real team
Scheduling Fair share, priority tiers, preemption policy, batch backfill of interactive capacity A scheduler holding SLOs while raising utilization
Placement Bin-packing models across heterogeneous GPUs, MIG partition strategy, spot/on-demand mix An optimizer beating manual placement on cost
Capacity Forecasting, buy-vs-rent, burst-to-cloud, reservation strategy A defensible 12-month capacity plan
Chargeback Per-tenant attribution, showback dashboards, quota design, incentive effects A chargeback system in production
MoE economics Expert load balancing, why MoE changes the cost curve, memory-vs-compute arbitrage An MoE-vs-dense TCO analysis

Milestone artifact: a published cost-optimization case study with before/after numbers.

Sources: llm-d Workload Variant Autoscaler design · SkyPilot docs and papers · cloud GPU pricing analyses · Kubernetes scheduling and Kueue docs · FinOps Foundation materials


Ongoing practice (all branches)

Cadence Practice
Daily Read release notes, not every paper. vLLM, llm-d, TRL, and your branch's repos
Weekly One deep read: a paper, an RFC, or a subsystem of source code. Write 200 words on it
Monthly One experiment producing a number nobody has published. Post it
Quarterly One upstream contribution merged
Yearly One talk or substantial technical article

Anti-goal: collecting knowledge you never apply. If a month passes with no measurement produced, the loop has broken.


Capstone

Deliverable: a public technical identity.

  1. One landscape document for your branch, kept current
  2. One reproduced SOTA result with documented discrepancies
  3. At least one merged upstream contribution
  4. One original measurement or tool that others use
  5. One published article or talk
  6. A written point of view: where your branch is going over the next 24 months, and what you're betting on

Rubric: could a hiring manager in your branch find you? Could a peer cite your work? That's the bar.


Assessment

There is no quiz. The assessment is external:

  • An upstream maintainer merged your work
  • Someone you've never met used your tool or cited your measurement
  • You were asked to speak or consult on your branch
  • You can name five open problems in your branch and have an opinion on each

Asset inventory

Type Count (minimum)
Theory articles 8 (method + 5 branch maps + 2 on contributing and publishing)
Code artifacts branch-dependent
Mermaid diagrams ~12
Generated images ~25
Labs the loop itself
Capstone 1