Skip to content

Skill 07 — Evals & Observability

The skill that makes every other skill trustworthy. Nearly everyone is weak here, which is exactly why being strong here is disproportionately valuable.

Duration 3 weeks (~40 hrs), but started in week 5 and never stopped
Impact ★★★★☆
Prerequisites Skill 01. Runs alongside Skills 03–06 rather than after them
Hardware 1 GPU for eval runs; the rest is CPU and storage
Status planned

Outcome

  • You can build an eval suite that actually predicts production quality, not benchmark theatre
  • You can prove a quantization, a fine-tune, or a config change helped — or admit it didn't
  • You can diagnose a degraded production system from its telemetry in minutes
  • You can calibrate an LLM judge against human labels and state its error bars

The two halves — don't conflate them

Evaluation Observability
Question Is this model good? Is this system healthy right now?
Timescale Pre-deployment, per-release Continuous, real-time
Layer Model behaviour Infrastructure + application
Tools lm-eval-harness, custom suites, judges Prometheus, Grafana, OTel, Langfuse
Failure Ship a worse model Don't notice you shipped a worse model

You need both. Most teams have neither, and think a dashboard of token counts is observability.


Modules

M Module Hrs Theme
M1 Evaluation Foundations 12 Measuring model quality without lying to yourself
M2 LLM-as-Judge & Application Evals 10 Grading things that have no right answer
M3 Infrastructure Observability 10 Knowing your serving stack
M4 Application Observability & the Feedback Loop 8 Tracing, and closing the loop back to data

M1 — Evaluation Foundations

ID Topic Key concepts Code artifact Visual Evidence
M1T1 What benchmarks actually measure MMLU/ARC/GSM8K/HumanEval/IFEval, log-likelihood vs generative scoring, why scores aren't comparable across harnesses Same model, same benchmark, two harnesses Chart: score variance across harness configs The same model scoring differently, explained
M1T2 lm-evaluation-harness in depth Task definitions, few-shot handling, prompt formats, batching, reproducibility, chat-template effects Full eval run across a task suite Mermaid: harness execution flow A reproducible eval report with pinned config
M1T3 Contamination Benchmark leakage, why public leaderboards are compromised, detection methods, private held-out sets Contamination detector against your training data Chart: score on public vs private variants of a task Evidence of contamination in a real public model
M1T4 Statistical honesty Sample size, confidence intervals, variance across seeds, why a 1-point difference is usually noise Significance calculator for eval deltas Chart: score distribution across seeds, with CI bands The minimum detectable effect for your suite
M1T5 Building a domain eval set Task decomposition, item writing, difficulty calibration, held-out design, golden-set maintenance Your own domain eval suite Mermaid: eval set construction pipeline A suite that correlates with real task success
M1T6 Capability vs behaviour vs safety evals Three different questions needing three different instruments; refusal, format compliance, robustness Multi-axis eval runner Gen: three-axis evaluation model A model profiled on all three axes

M1T5 is the module's real payload. Public benchmarks tell you almost nothing about whether your model works for your users. Everything else here is scaffolding for building your own.


M2 — LLM-as-Judge & Application Evals

ID Topic Key concepts Code artifact Visual Evidence
M2T1 When judging is the only option Open-ended generation, summarization, style, helpfulness; reference-based vs reference-free Pairwise and pointwise judge implementations Mermaid: judging modes compared A judge running on real outputs
M2T2 Judge bias and calibration Position bias, verbosity bias, self-preference, sycophancy; mitigation via swapping and rubrics Bias measurement harness Chart: bias magnitude per type, before and after mitigation Quantified bias in your own judge
M2T3 Agreement with humans Cohen's kappa, correlation with human labels, when a judge is trustworthy enough to gate a release Human-vs-judge agreement study on 100 items Chart: agreement matrix Your judge's error bars, stated
M2T4 Evaluating RAG Separating retrieval quality from generation quality, faithfulness/groundedness, citation accuracy, answer relevance RAG eval pipeline with decomposed metrics Gen sequence: pipeline → per-stage measurement points Which stage is actually failing, identified
M2T5 Evaluating agents Trajectory-level scoring, tool-call accuracy, task completion, cost per task, partial credit Agent trajectory evaluator Mermaid: trajectory scoring flow Failure taxonomy across 50 real trajectories
M2T6 Datasets, experiments, regression gates Versioned eval datasets, experiment tracking, CI integration, blocking a release on a regression Eval gate as a CI job Mermaid: CI pipeline with the gate A regression caught automatically

M3 — Infrastructure Observability

ID Topic Key concepts Code artifact Visual Evidence
M3T1 The engine's metrics vLLM Prometheus metrics: TTFT, TPOT, running/waiting requests, KV cache usage, preemptions, prefix hit rate — and what each means Metrics scraper + annotated reference Mermaid: where each metric originates in the engine A complete annotated metric catalogue
M3T2 GPU telemetry DCGM: utilization vs saturation, SM occupancy, memory, power, clocks, ECC, thermal throttling DCGM exporter + GPU dashboard Gen: utilization vs saturation contrast (Pattern B) Throttling detected under sustained load
M3T3 Dashboards that answer questions The four golden signals for inference, dashboard design, drill-down paths, what to put on the first screen Grafana dashboard set (engine / GPU / gateway) Mermaid: dashboard hierarchy Dashboards that answered a real incident
M3T4 Alerting on SLOs SLI/SLO/error budgets for inference, burn-rate alerts, alert fatigue, what should never page Alert rules with burn-rate windows Chart: error budget consumption over time An alert that fired correctly, and one that correctly didn't
M3T5 Diagnosing from telemetry alone Symptom→cause playbooks: rising TTFT flat throughput, preemption storms, cache thrash, memory creep, straggler GPU Fault-injection suite with six scenarios Mermaid: diagnostic decision tree All six diagnosed from dashboards only

M4 — Application Observability & the Feedback Loop

ID Topic Key concepts Code artifact Visual Evidence
M4T1 OpenTelemetry GenAI semantic conventions The vendor-neutral standard: spans, events, metrics for LLM calls, agents, and MCP; why standardizing beats a proprietary SDK Instrument a stack to the GenAI semconv Mermaid: span hierarchy for an agent request Traces conforming to the spec, portable across backends
M4T2 Langfuse Self-hosted deployment, traces/sessions/users, cost and latency attribution, agent graph views Self-hosted Langfuse capturing a live app Gen: nested trace structure Full traces of a multi-step agent request
M4T3 Prompt management & versioning Prompts as versioned artifacts, deploy-by-label, linking prompts to traces, A/B by prompt version Prompt registry integrated into the app Mermaid: prompt lifecycle A prompt rollback executed without a deploy
M4T4 Online evaluation Scoring production traces with judges, sampling strategy, drift detection, user-feedback capture Online scoring pipeline on live traffic Chart: quality score over time, with a drift event Drift detected before users complained
M4T5 Closing the loop Production failures → annotation queue → eval dataset → training data → next model; the data flywheel Failure-to-dataset pipeline Gen sequence: the flywheel, one image per stage Real production failures converted into eval items

M4T5 is the endpoint of the entire curriculum. Everything else feeds this loop.


Labs

Lab Goal Success criterion
LAB-M1 Trust nothing Reproduce a published benchmark score; explain every point of discrepancy
LAB-M2 Calibrate a judge ≥0.7 Cohen's kappa with human labels on 100 items, with documented bias mitigation
LAB-M3 Six incidents Diagnose all six injected faults from telemetry alone, no log access
LAB-M4 Full instrumentation End-to-end trace from gateway to token, correlated with engine metrics

Capstone

Deliverable: quality-system — the measurement layer for everything you built in Skills 01–06.

  1. Domain eval suite with justification for every task, plus a contamination audit
  2. Calibrated judge with stated agreement and error bars
  3. Regression gate in CI blocking any model or config that degrades quality
  4. Full observability stack: engine + GPU + gateway metrics, dashboards, SLO alerts
  5. End-to-end tracing conforming to OTel GenAI semantic conventions
  6. Online evaluation scoring a sample of production traffic
  7. The flywheel: a working pipeline from production failure to eval item to training candidate
  8. Incident runbook derived from six failures you actually induced
  9. A retrospective: re-run your Skill 03 and Skill 05 conclusions through this system. Were any of them wrong? Say so.

Item 9 is the point of the capstone. A quality system that never overturns a previous conclusion isn't measuring anything.


Assessment

Tier Count Example
Recall 6 What is the difference between GPU utilization and GPU saturation?
Apply 13 Your model gained 4 points on MMLU and users say it got worse. Give four explanations and how to test each
Design 5 Design the eval and observability strategy for a customer-facing agent with 10k daily conversations

Practical challenge: given a live degraded system, diagnose the root cause from telemetry within 30 minutes, and propose the alert that would have caught it earlier.


Asset inventory

Type Count (minimum)
Theory articles 22
Code artifacts 21
Mermaid diagrams ~28
Generated images ~80
Charts ~30
Labs 4
Capstone 1

Primary sources

  • OpenTelemetry GenAI semantic conventions — open-telemetry/semantic-conventions-genai (spans, events, metrics, agent and MCP conventions)
  • Langfuse docs — observability, prompt management, evaluation, datasets & experiments, self-hosting
  • EleutherAI lm-evaluation-harness; HELM; evalchemy
  • vLLM metrics design doc + the Prometheus endpoint reference
  • NVIDIA DCGM exporter docs
  • Papers: Judging LLM-as-a-Judge (MT-Bench/Chatbot Arena), G-Eval, RAGAS, position/verbosity bias studies
  • Google SRE Workbook — SLOs and burn-rate alerting (the concepts transfer directly)