Skip to content

Skill 08 — Security & Governance

The skill that decides whether any of this is allowed to reach production. Self-hosting removes your vendor's safety layer, legal team, and abuse detection. You now own all three.

Duration 2 weeks (~25 hrs)
Impact ★★★☆☆ for job titles, ★★★★★ for whether your project ships
Prerequisites Skill 06 (you need a platform to secure), Skill 07 (you need measurement to prove safety)
Hardware 1–2 GPUs (guard models are small)
Status planned

Outcome

  • You can threat-model an LLM system and name the attacks it's actually exposed to
  • You can deploy guardrails and state their false-positive/false-negative rates instead of assuming they work
  • You can audit a model's licence and provenance before it goes near production
  • You can answer "is this compliant?" with evidence rather than hope

What changes when you self-host

Hosted API Self-hosted
Provider filters inputs and outputs You build it, or there is none
Provider absorbs abuse and liability You absorb it
Model provenance is the provider's problem You must audit weights, licence, and supply chain
Rate limits and quotas built in You implement them
Compliance attestations provided You produce them

Every topic below exists because of a row in that table.


Modules

M Module Hrs Theme
M1 Threat Modelling & Attacks 9 What actually goes wrong
M2 Guardrails & Defence 9 Controls, and honest measurement of them
M3 Supply Chain, Licensing & Compliance 7 Provenance and the paperwork that unblocks launch

M1 — Threat Modelling & Attacks

ID Topic Key concepts Code artifact Visual Evidence
M1T1 The LLM threat model OWASP Top 10 for LLM Applications mapped to a self-hosted architecture; trust boundaries; what's in scope for you now Threat model of your Skill 06 platform Gen sequence: architecture → trust boundaries → attack surfaces A written threat model with ranked risks
M1T2 Prompt injection Direct vs indirect injection via retrieved content and tool output; why it's unsolved; why agents make it severe Working indirect injection through a RAG corpus Gen: trusted vs untrusted content flowing into one context (Pattern B) A successful injection against your own stack
M1T3 Jailbreaks & adversarial inputs Role-play, encoding, many-shot, gradient-based suffixes; why filters generalize poorly Jailbreak test suite run against your model Chart: success rate by technique Your model's actual jailbreak rate
M1T4 Data exfiltration & memorization Training-data extraction, PII leakage, system-prompt extraction, tool-mediated exfiltration Memorization probe against your fine-tuned model Gen: extraction attack path Whether your fine-tune memorized its training data
M1T5 Infrastructure-level attacks Model-serving DoS (long-prompt and expensive-generation attacks), tenant isolation, KV-cache side channels via shared prefixes, endpoint exposure Resource-exhaustion test + mitigation Mermaid: attack surface of the serving stack A DoS you executed and then blocked

M1T5 is the topic most security material misses entirely — and it's the one you're uniquely positioned to understand after Skills 01 and 06.


M2 — Guardrails & Defence

ID Topic Key concepts Code artifact Visual Evidence
M2T1 The defence architecture Input filter → model → output filter → tool authorization; defence in depth; where latency budget goes Guardrail pipeline wrapped around your gateway Mermaid: request path with control points Guardrails deployed with measured latency cost
M2T2 Guard models Llama Guard, Granite Guardian, NeMo Guardrails, ShieldGemma; taxonomy design; self-hosting a guard model Deploy a guard model alongside the main model Gen: parallel guard path Guard model serving with measured overhead
M2T3 Measuring a guardrail honestly False-positive vs false-negative rates, the usability cost of over-blocking, threshold tuning, per-category performance Guardrail eval on a labelled adversarial + benign set Chart: ROC per category, with the chosen threshold marked Your guardrail's real error rates, published
M2T4 PII and data handling Detection and redaction, data-residency, log hygiene (your traces contain user data), retention policy PII redaction in the trace pipeline Mermaid: data flow with redaction points Traces verified free of PII
M2T5 Structural defences Constrained decoding as a safety control, tool-call allowlists and authorization, sandboxed execution, least privilege for agents Sandboxed tool executor with an authorization layer Gen: privilege boundaries around the agent A tool-abuse attempt blocked by design, not by a filter
M2T6 Abuse prevention & rate limiting Per-tenant quotas, token budgets, anomaly detection, cost-based abuse (someone burning your GPUs) Quota + anomaly detection on the gateway Chart: cost per tenant with an anomaly flagged A simulated abuse pattern detected

M3 — Supply Chain, Licensing & Compliance

ID Topic Key concepts Code artifact Visual Evidence
M3T1 Model supply chain Weight provenance, safetensors vs pickle, trust_remote_code as arbitrary code execution, checksum verification, registry hygiene Model verification pipeline with a pickle-payload demo Gen: supply chain from upstream to serving A malicious checkpoint caught by your pipeline
M3T2 Licensing reality Apache-2.0 vs MIT vs Llama Community vs Gemma vs research-only; derivative works from fine-tuning; synthetic data from a teacher model's terms of service Licence auditor covering models, datasets, and outputs Mermaid: licence compatibility decision tree A clean licence audit for your capstone model
M3T3 Documentation as governance Model cards, datasheets, evaluation reports, SBOM, reproducibility manifests Auto-generated model card from your training run Mermaid: artifact lineage graph A complete, honest model card
M3T4 The regulatory layer EU AI Act obligations for providers vs deployers, risk classification, transparency duties, sector rules (health, finance), where self-hosting helps and where it doesn't Compliance checklist mapped to your artifacts Mermaid: obligation → evidence mapping A gap analysis with named owners
M3T5 Operational security Secrets, network isolation for GPU nodes, RBAC on the platform, audit logging, incident response for AI-specific incidents Hardened deployment config + audit trail Mermaid: incident response flow A tabletop incident exercise, written up

Labs

Lab Goal Success criterion
LAB-M1 Attack your own stack Successfully execute indirect prompt injection, a jailbreak, and a resource-exhaustion attack
LAB-M2 Defend it Block all three, with false-positive rate on benign traffic below a stated threshold
LAB-M3 Audit the chain Full licence and provenance audit of your capstone model, with one issue found and resolved

Capstone

Deliverable: security-and-governance-pack — the document set that lets your platform pass review.

  1. Threat model of the full platform, with ranked risks and named mitigations
  2. Red-team report: attacks attempted, which succeeded, evidence
  3. Guardrail architecture with measured FP/FN rates per category, and the accepted tradeoff
  4. PII handling policy, with traces verified clean
  5. Model supply-chain verification pipeline
  6. Licence and provenance audit covering models, datasets, and synthetic data
  7. Model card and datasheet for your Skill 05 model
  8. Regulatory gap analysis with owners and dates
  9. Incident response runbook for AI-specific incidents
  10. A one-page risk acceptance statement — what you're choosing not to mitigate, and why

Item 10 separates security theatre from security engineering. Every real system has accepted risks; maturity is naming them explicitly.


Assessment

Tier Count Example
Recall 6 What makes indirect prompt injection harder to defend than direct?
Apply 10 Your guardrail blocks 4% of legitimate traffic. Walk through how you'd decide whether that's acceptable
Design 4 Design the security architecture for a self-hosted agent with database and email tool access

Practical challenge: given an unfamiliar deployed LLM service, produce a threat model and a working exploit within 3 hours.


Asset inventory

Type Count (minimum)
Theory articles 16
Code artifacts 15
Mermaid diagrams ~20
Generated images ~55
Charts ~12
Labs 3
Capstone 1

Primary sources

  • OWASP Top 10 for LLM Applications + the OWASP Agentic AI threat guidance
  • NIST AI Risk Management Framework and the Generative AI Profile
  • Llama Guard / Granite Guardian / ShieldGemma model cards; NVIDIA NeMo Guardrails docs
  • Papers: Not what you've signed up for (indirect prompt injection), Universal adversarial suffixes (GCG), Extracting Training Data from LLMs
  • MITRE ATLAS — adversarial ML threat matrix
  • EU AI Act text (Articles on GPAI, transparency, and deployer obligations)
  • Model Cards for Model Reporting (Mitchell et al.); Datasheets for Datasets (Gebru et al.)
  • Hugging Face security docs — safetensors, malware scanning, trust_remote_code guidance