Skill 09 — Specialization¶
Generalists plateau here. Specialists don't. Skills 00–08 make you a strong, hireable LLM infrastructure engineer. This skill makes you the person people are told to go talk to.
| Duration | 6 months (months 7–12), ongoing |
| Impact | varies by branch — see the table |
| Prerequisites | Skills 00–08 complete, with capstones done |
| Hardware | branch-dependent |
| Status | planned |
How this skill differs¶
It has no fixed topic list, because the frontier moves. It has a structure for going deep and five branch maps. Pick one branch, optionally a second later. Depth beats breadth from here on.
Choose your branch¶
| Branch | Market size | Pay ceiling | Difficulty | Pick this if… |
|---|---|---|---|---|
| A. Kernels & Compilers | Small | Highest | Very high | You enjoy assembly-level reasoning and profiler output |
| B. Serving Platform Engineering | Large | High | High | You liked Skills 01 and 06 most, and want the strongest public credential |
| C. Agentic RL & Environments | Growing fastest | High | High | You liked Skill 05 M4, and want to be early to the next wave |
| D. Distributed Training at Scale | Small | High | Very high | You want lab / sovereign-AI / national-program roles |
| E. Cost & Multi-Tenancy Engineering | Large | Medium-high | Medium | You like the economics and want to own a P&L line |
Recommendation given the path so far: B or C. B compounds directly on your strongest skills and produces public artifacts. C positions you ahead of a wave that hasn't crested.
The specialization method (applies to every branch)¶
This is the actual content of the skill — a repeatable loop, run monthly.
flowchart LR
A[Pick a real problem<br/>in your branch] --> B[Reproduce the<br/>state of the art]
B --> C[Find the gap:<br/>where does it fail?]
C --> D[Build something<br/>that closes it]
D --> E[Measure honestly]
E --> F[Publish:<br/>PR, post, or talk]
F --> A
| Phase | Duration | Output |
|---|---|---|
| 1. Map the territory | Weeks 1–3 | Read the top 20 papers/RFCs in the branch; write a landscape doc naming the 5 open problems |
| 2. Reproduce | Weeks 4–8 | Reproduce one significant published result on your hardware. Document every discrepancy |
| 3. Find your gap | Weeks 9–12 | Where does the SOTA fail on a workload you care about? Characterize it with data |
| 4. Contribute | Weeks 13–20 | Fix something real. Merged upstream PR, or a tool people use |
| 5. Publish | Weeks 21–24 | A technical post with reproducible numbers, or a conference talk |
The output of this skill is public artifacts, not private knowledge. A merged vLLM PR is worth more than any certificate.
Branch A — Kernels & Compilers¶
| Area | What to learn | Milestone |
|---|---|---|
| Triton | Language model, autotuning, block pointers, when Triton beats hand-written CUDA | Write a fused kernel beating the PyTorch eager baseline |
| CUDA | Memory coalescing, shared memory, warp primitives, occupancy, tensor core intrinsics | Implement a tiled GEMM within 2× of cuBLAS |
| Attention kernels | FlashAttention 1/2 tiling and recomputation, paged variants, MLA kernels, decode-optimized kernels | Implement a paged attention decode kernel |
| MoE kernels | Fused MoE, grouped GEMM, all-to-all overlap, expert-load handling | Contribute to a fused MoE path |
| Compilers | torch.compile internals, Inductor, custom ops, graph passes, CUDA graphs |
Add a fusion pass that measurably helps |
| Profiling | Nsight Compute/Systems, roofline analysis, occupancy vs ILP | Diagnose and fix a real kernel bottleneck |
Note on our hardware: SM 8.0 means no FP8 or FA3 kernel work. Plenty remains — INT8/INT4 kernels, paged decode, MoE, and compiler passes are all fully available on Ampere.
Milestone artifact: a merged kernel or compiler PR in vLLM, SGLang, or PyTorch.
Sources: GPU MODE lectures and Discord · Triton docs and tutorials · CUDA C++ Programming Guide ·
FlashAttention papers · CUTLASS/CuTe docs · vllm/kernels/ and csrc/ source
Branch B — Serving Platform Engineering¶
| Area | What to learn | Milestone |
|---|---|---|
| Engine internals | vLLM V1 scheduler, KV cache manager, worker/executor split, attention backend interface | Add a scheduler policy or a config knob |
| KV connectors | NIXL, LMCache, Mooncake; the connector interface; transfer protocols and lease semantics | Implement or improve a KV connector |
| Routing | EPP scorers and filters, prefix-cache indexing, latency prediction, SLO-aware endpoint selection | Write a custom EPP scorer, prove it beats the default |
| Standards | Gateway API Inference Extension conformance, the model-server protocol, WG-Serving proposals | Contribute to a proposal or conformance test |
| Autoscaling | Workload-variant autoscaling, predictive scaling, cost-optimal replica placement | Build an autoscaler that beats KEDA on a real workload |
| Operations | Multi-cluster, disaster recovery, upgrade choreography for stateful GPU services | A public runbook or operator |
Milestone artifact: a merged PR in vLLM, llm-d, or gateway-api-inference-extension.
Sources: vLLM design docs and source · llm-d architecture docs and repos · Gateway API Inference Extension enhancements/proposals · NVIDIA Dynamo design docs · Kubernetes WG-Serving meetings
Branch C — Agentic RL & Environments¶
| Area | What to learn | Milestone |
|---|---|---|
| RL at scale | verl / OpenRLHF architectures, async rollout, weight sync, colocated vs disaggregated | Run multi-node-simulated GRPO efficiently |
| Environment design | Task specification, reward shaping, partial credit, sandboxing, determinism and seeding | Publish an environment others use |
| Verifiable rewards | Verifier design, gaming resistance, verification cost, hybrid judge+verifier rewards | A verifier suite resistant to a documented attack set |
| Multi-turn credit | Long-horizon assignment, turn-level vs trajectory-level rewards, tool-use RL | Train a tool-using agent beating a prompted baseline |
| Reward hacking | Detection, monitoring, mitigation, KL control, adversarial evaluation of your own rewards | A published taxonomy of hacks you found |
| Inference-time scaling | Best-of-N, self-consistency, search, and the RL-vs-inference-compute tradeoff | Quantify where each is cost-optimal |
Milestone artifact: a published open environment + a trained policy with reproducible results.
Sources: DeepSeek-R1 and Open-R1 · verl docs · TRL GRPO + OpenEnv/Harbor docs · Tulu 3 · process-supervision and RLVR literature · agent benchmark suites (SWE-bench, τ-bench, WebArena)
Branch D — Distributed Training at Scale¶
| Area | What to learn | Milestone |
|---|---|---|
| Parallelism composition | FSDP2 + TP + PP + CP + EP together; when each dimension pays | Compose 3+ dimensions on a real model |
| Communication | NCCL internals, overlap strategies, async TP, all-to-all for MoE | Diagnose and fix a comms bottleneck |
| Memory | Activation checkpointing policies, offload, optimizer-state sharding, memory estimation | Predict peak memory within 10% before running |
| Reliability | Distributed and async checkpointing, TorchFT, flight recorder, straggler and hang diagnosis | Recover a multi-rank failure with no loss |
| Efficiency | MFU optimization, profiling training, dataloader bottlenecks | Improve MFU by ≥15% on a reference config |
| Low precision | bf16 today; float8/MXFP8 as a decision framework for future hardware | A written hardware-purchase analysis |
Reality check: this branch is genuinely constrained by 5–6 A100s. You can master the techniques and read the code, but you can't demonstrate scale. Pick it only if the concepts themselves are the goal.
Milestone artifact: a merged torchtitan PR, or a reproducible efficiency study.
Sources: torchtitan source + ICLR paper · PyTorch distributed blog series · Megatron-LM and DeepSpeed papers · NCCL docs · the Llama 3 and DeepSeek-V3 infrastructure sections
Branch E — Cost & Multi-Tenancy Engineering¶
| Area | What to learn | Milestone |
|---|---|---|
| Unit economics | $/1M tokens modelling, utilization math, headroom cost, marginal vs average cost | A cost model adopted by a real team |
| Scheduling | Fair share, priority tiers, preemption policy, batch backfill of interactive capacity | A scheduler holding SLOs while raising utilization |
| Placement | Bin-packing models across heterogeneous GPUs, MIG partition strategy, spot/on-demand mix | An optimizer beating manual placement on cost |
| Capacity | Forecasting, buy-vs-rent, burst-to-cloud, reservation strategy | A defensible 12-month capacity plan |
| Chargeback | Per-tenant attribution, showback dashboards, quota design, incentive effects | A chargeback system in production |
| MoE economics | Expert load balancing, why MoE changes the cost curve, memory-vs-compute arbitrage | An MoE-vs-dense TCO analysis |
Milestone artifact: a published cost-optimization case study with before/after numbers.
Sources: llm-d Workload Variant Autoscaler design · SkyPilot docs and papers · cloud GPU pricing analyses · Kubernetes scheduling and Kueue docs · FinOps Foundation materials
Ongoing practice (all branches)¶
| Cadence | Practice |
|---|---|
| Daily | Read release notes, not every paper. vLLM, llm-d, TRL, and your branch's repos |
| Weekly | One deep read: a paper, an RFC, or a subsystem of source code. Write 200 words on it |
| Monthly | One experiment producing a number nobody has published. Post it |
| Quarterly | One upstream contribution merged |
| Yearly | One talk or substantial technical article |
Anti-goal: collecting knowledge you never apply. If a month passes with no measurement produced, the loop has broken.
Capstone¶
Deliverable: a public technical identity.
- One landscape document for your branch, kept current
- One reproduced SOTA result with documented discrepancies
- At least one merged upstream contribution
- One original measurement or tool that others use
- One published article or talk
- A written point of view: where your branch is going over the next 24 months, and what you're betting on
Rubric: could a hiring manager in your branch find you? Could a peer cite your work? That's the bar.
Assessment¶
There is no quiz. The assessment is external:
- An upstream maintainer merged your work
- Someone you've never met used your tool or cited your measurement
- You were asked to speak or consult on your branch
- You can name five open problems in your branch and have an opinion on each
Asset inventory¶
| Type | Count (minimum) |
|---|---|
| Theory articles | 8 (method + 5 branch maps + 2 on contributing and publishing) |
| Code artifacts | branch-dependent |
| Mermaid diagrams | ~12 |
| Generated images | ~25 |
| Labs | the loop itself |
| Capstone | 1 |