Skill 07 — Evals & Observability
The skill that makes every other skill trustworthy. Nearly everyone is weak here, which is
exactly why being strong here is disproportionately valuable.
|
|
| Duration |
3 weeks (~40 hrs), but started in week 5 and never stopped |
| Impact |
★★★★☆ |
| Prerequisites |
Skill 01. Runs alongside Skills 03–06 rather than after them |
| Hardware |
1 GPU for eval runs; the rest is CPU and storage |
| Status |
planned |
Outcome
- You can build an eval suite that actually predicts production quality, not benchmark theatre
- You can prove a quantization, a fine-tune, or a config change helped — or admit it didn't
- You can diagnose a degraded production system from its telemetry in minutes
- You can calibrate an LLM judge against human labels and state its error bars
The two halves — don't conflate them
|
Evaluation |
Observability |
| Question |
Is this model good? |
Is this system healthy right now? |
| Timescale |
Pre-deployment, per-release |
Continuous, real-time |
| Layer |
Model behaviour |
Infrastructure + application |
| Tools |
lm-eval-harness, custom suites, judges |
Prometheus, Grafana, OTel, Langfuse |
| Failure |
Ship a worse model |
Don't notice you shipped a worse model |
You need both. Most teams have neither, and think a dashboard of token counts is observability.
Modules
| M |
Module |
Hrs |
Theme |
| M1 |
Evaluation Foundations |
12 |
Measuring model quality without lying to yourself |
| M2 |
LLM-as-Judge & Application Evals |
10 |
Grading things that have no right answer |
| M3 |
Infrastructure Observability |
10 |
Knowing your serving stack |
| M4 |
Application Observability & the Feedback Loop |
8 |
Tracing, and closing the loop back to data |
M1 — Evaluation Foundations
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M1T1 |
What benchmarks actually measure |
MMLU/ARC/GSM8K/HumanEval/IFEval, log-likelihood vs generative scoring, why scores aren't comparable across harnesses |
Same model, same benchmark, two harnesses |
Chart: score variance across harness configs |
The same model scoring differently, explained |
| M1T2 |
lm-evaluation-harness in depth |
Task definitions, few-shot handling, prompt formats, batching, reproducibility, chat-template effects |
Full eval run across a task suite |
Mermaid: harness execution flow |
A reproducible eval report with pinned config |
| M1T3 |
Contamination |
Benchmark leakage, why public leaderboards are compromised, detection methods, private held-out sets |
Contamination detector against your training data |
Chart: score on public vs private variants of a task |
Evidence of contamination in a real public model |
| M1T4 |
Statistical honesty |
Sample size, confidence intervals, variance across seeds, why a 1-point difference is usually noise |
Significance calculator for eval deltas |
Chart: score distribution across seeds, with CI bands |
The minimum detectable effect for your suite |
| M1T5 |
Building a domain eval set |
Task decomposition, item writing, difficulty calibration, held-out design, golden-set maintenance |
Your own domain eval suite |
Mermaid: eval set construction pipeline |
A suite that correlates with real task success |
| M1T6 |
Capability vs behaviour vs safety evals |
Three different questions needing three different instruments; refusal, format compliance, robustness |
Multi-axis eval runner |
Gen: three-axis evaluation model |
A model profiled on all three axes |
M1T5 is the module's real payload. Public benchmarks tell you almost nothing about whether your
model works for your users. Everything else here is scaffolding for building your own.
M2 — LLM-as-Judge & Application Evals
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M2T1 |
When judging is the only option |
Open-ended generation, summarization, style, helpfulness; reference-based vs reference-free |
Pairwise and pointwise judge implementations |
Mermaid: judging modes compared |
A judge running on real outputs |
| M2T2 |
Judge bias and calibration |
Position bias, verbosity bias, self-preference, sycophancy; mitigation via swapping and rubrics |
Bias measurement harness |
Chart: bias magnitude per type, before and after mitigation |
Quantified bias in your own judge |
| M2T3 |
Agreement with humans |
Cohen's kappa, correlation with human labels, when a judge is trustworthy enough to gate a release |
Human-vs-judge agreement study on 100 items |
Chart: agreement matrix |
Your judge's error bars, stated |
| M2T4 |
Evaluating RAG |
Separating retrieval quality from generation quality, faithfulness/groundedness, citation accuracy, answer relevance |
RAG eval pipeline with decomposed metrics |
Gen sequence: pipeline → per-stage measurement points |
Which stage is actually failing, identified |
| M2T5 |
Evaluating agents |
Trajectory-level scoring, tool-call accuracy, task completion, cost per task, partial credit |
Agent trajectory evaluator |
Mermaid: trajectory scoring flow |
Failure taxonomy across 50 real trajectories |
| M2T6 |
Datasets, experiments, regression gates |
Versioned eval datasets, experiment tracking, CI integration, blocking a release on a regression |
Eval gate as a CI job |
Mermaid: CI pipeline with the gate |
A regression caught automatically |
M3 — Infrastructure Observability
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M3T1 |
The engine's metrics |
vLLM Prometheus metrics: TTFT, TPOT, running/waiting requests, KV cache usage, preemptions, prefix hit rate — and what each means |
Metrics scraper + annotated reference |
Mermaid: where each metric originates in the engine |
A complete annotated metric catalogue |
| M3T2 |
GPU telemetry |
DCGM: utilization vs saturation, SM occupancy, memory, power, clocks, ECC, thermal throttling |
DCGM exporter + GPU dashboard |
Gen: utilization vs saturation contrast (Pattern B) |
Throttling detected under sustained load |
| M3T3 |
Dashboards that answer questions |
The four golden signals for inference, dashboard design, drill-down paths, what to put on the first screen |
Grafana dashboard set (engine / GPU / gateway) |
Mermaid: dashboard hierarchy |
Dashboards that answered a real incident |
| M3T4 |
Alerting on SLOs |
SLI/SLO/error budgets for inference, burn-rate alerts, alert fatigue, what should never page |
Alert rules with burn-rate windows |
Chart: error budget consumption over time |
An alert that fired correctly, and one that correctly didn't |
| M3T5 |
Diagnosing from telemetry alone |
Symptom→cause playbooks: rising TTFT flat throughput, preemption storms, cache thrash, memory creep, straggler GPU |
Fault-injection suite with six scenarios |
Mermaid: diagnostic decision tree |
All six diagnosed from dashboards only |
M4 — Application Observability & the Feedback Loop
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M4T1 |
OpenTelemetry GenAI semantic conventions |
The vendor-neutral standard: spans, events, metrics for LLM calls, agents, and MCP; why standardizing beats a proprietary SDK |
Instrument a stack to the GenAI semconv |
Mermaid: span hierarchy for an agent request |
Traces conforming to the spec, portable across backends |
| M4T2 |
Langfuse |
Self-hosted deployment, traces/sessions/users, cost and latency attribution, agent graph views |
Self-hosted Langfuse capturing a live app |
Gen: nested trace structure |
Full traces of a multi-step agent request |
| M4T3 |
Prompt management & versioning |
Prompts as versioned artifacts, deploy-by-label, linking prompts to traces, A/B by prompt version |
Prompt registry integrated into the app |
Mermaid: prompt lifecycle |
A prompt rollback executed without a deploy |
| M4T4 |
Online evaluation |
Scoring production traces with judges, sampling strategy, drift detection, user-feedback capture |
Online scoring pipeline on live traffic |
Chart: quality score over time, with a drift event |
Drift detected before users complained |
| M4T5 |
Closing the loop |
Production failures → annotation queue → eval dataset → training data → next model; the data flywheel |
Failure-to-dataset pipeline |
Gen sequence: the flywheel, one image per stage |
Real production failures converted into eval items |
M4T5 is the endpoint of the entire curriculum. Everything else feeds this loop.
Labs
| Lab |
Goal |
Success criterion |
| LAB-M1 |
Trust nothing |
Reproduce a published benchmark score; explain every point of discrepancy |
| LAB-M2 |
Calibrate a judge |
≥0.7 Cohen's kappa with human labels on 100 items, with documented bias mitigation |
| LAB-M3 |
Six incidents |
Diagnose all six injected faults from telemetry alone, no log access |
| LAB-M4 |
Full instrumentation |
End-to-end trace from gateway to token, correlated with engine metrics |
Capstone
Deliverable: quality-system — the measurement layer for everything you built in Skills 01–06.
- Domain eval suite with justification for every task, plus a contamination audit
- Calibrated judge with stated agreement and error bars
- Regression gate in CI blocking any model or config that degrades quality
- Full observability stack: engine + GPU + gateway metrics, dashboards, SLO alerts
- End-to-end tracing conforming to OTel GenAI semantic conventions
- Online evaluation scoring a sample of production traffic
- The flywheel: a working pipeline from production failure to eval item to training candidate
- Incident runbook derived from six failures you actually induced
- A retrospective: re-run your Skill 03 and Skill 05 conclusions through this system.
Were any of them wrong? Say so.
Item 9 is the point of the capstone. A quality system that never overturns a previous conclusion
isn't measuring anything.
Assessment
| Tier |
Count |
Example |
| Recall |
6 |
What is the difference between GPU utilization and GPU saturation? |
| Apply |
13 |
Your model gained 4 points on MMLU and users say it got worse. Give four explanations and how to test each |
| Design |
5 |
Design the eval and observability strategy for a customer-facing agent with 10k daily conversations |
Practical challenge: given a live degraded system, diagnose the root cause from telemetry within
30 minutes, and propose the alert that would have caught it earlier.
Asset inventory
| Type |
Count (minimum) |
| Theory articles |
22 |
| Code artifacts |
21 |
| Mermaid diagrams |
~28 |
| Generated images |
~80 |
| Charts |
~30 |
| Labs |
4 |
| Capstone |
1 |
Primary sources
- OpenTelemetry GenAI semantic conventions —
open-telemetry/semantic-conventions-genai (spans, events, metrics, agent and MCP conventions)
- Langfuse docs — observability, prompt management, evaluation, datasets & experiments, self-hosting
- EleutherAI
lm-evaluation-harness; HELM; evalchemy
- vLLM metrics design doc + the Prometheus endpoint reference
- NVIDIA DCGM exporter docs
- Papers: Judging LLM-as-a-Judge (MT-Bench/Chatbot Arena), G-Eval, RAGAS, position/verbosity bias studies
- Google SRE Workbook — SLOs and burn-rate alerting (the concepts transfer directly)