Skill 08 — Security & Governance
The skill that decides whether any of this is allowed to reach production.
Self-hosting removes your vendor's safety layer, legal team, and abuse detection. You now own all three.
|
|
| Duration |
2 weeks (~25 hrs) |
| Impact |
★★★☆☆ for job titles, ★★★★★ for whether your project ships |
| Prerequisites |
Skill 06 (you need a platform to secure), Skill 07 (you need measurement to prove safety) |
| Hardware |
1–2 GPUs (guard models are small) |
| Status |
planned |
Outcome
- You can threat-model an LLM system and name the attacks it's actually exposed to
- You can deploy guardrails and state their false-positive/false-negative rates instead of assuming they work
- You can audit a model's licence and provenance before it goes near production
- You can answer "is this compliant?" with evidence rather than hope
What changes when you self-host
| Hosted API |
Self-hosted |
| Provider filters inputs and outputs |
You build it, or there is none |
| Provider absorbs abuse and liability |
You absorb it |
| Model provenance is the provider's problem |
You must audit weights, licence, and supply chain |
| Rate limits and quotas built in |
You implement them |
| Compliance attestations provided |
You produce them |
Every topic below exists because of a row in that table.
Modules
| M |
Module |
Hrs |
Theme |
| M1 |
Threat Modelling & Attacks |
9 |
What actually goes wrong |
| M2 |
Guardrails & Defence |
9 |
Controls, and honest measurement of them |
| M3 |
Supply Chain, Licensing & Compliance |
7 |
Provenance and the paperwork that unblocks launch |
M1 — Threat Modelling & Attacks
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M1T1 |
The LLM threat model |
OWASP Top 10 for LLM Applications mapped to a self-hosted architecture; trust boundaries; what's in scope for you now |
Threat model of your Skill 06 platform |
Gen sequence: architecture → trust boundaries → attack surfaces |
A written threat model with ranked risks |
| M1T2 |
Prompt injection |
Direct vs indirect injection via retrieved content and tool output; why it's unsolved; why agents make it severe |
Working indirect injection through a RAG corpus |
Gen: trusted vs untrusted content flowing into one context (Pattern B) |
A successful injection against your own stack |
| M1T3 |
Jailbreaks & adversarial inputs |
Role-play, encoding, many-shot, gradient-based suffixes; why filters generalize poorly |
Jailbreak test suite run against your model |
Chart: success rate by technique |
Your model's actual jailbreak rate |
| M1T4 |
Data exfiltration & memorization |
Training-data extraction, PII leakage, system-prompt extraction, tool-mediated exfiltration |
Memorization probe against your fine-tuned model |
Gen: extraction attack path |
Whether your fine-tune memorized its training data |
| M1T5 |
Infrastructure-level attacks |
Model-serving DoS (long-prompt and expensive-generation attacks), tenant isolation, KV-cache side channels via shared prefixes, endpoint exposure |
Resource-exhaustion test + mitigation |
Mermaid: attack surface of the serving stack |
A DoS you executed and then blocked |
M1T5 is the topic most security material misses entirely — and it's the one you're uniquely
positioned to understand after Skills 01 and 06.
M2 — Guardrails & Defence
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M2T1 |
The defence architecture |
Input filter → model → output filter → tool authorization; defence in depth; where latency budget goes |
Guardrail pipeline wrapped around your gateway |
Mermaid: request path with control points |
Guardrails deployed with measured latency cost |
| M2T2 |
Guard models |
Llama Guard, Granite Guardian, NeMo Guardrails, ShieldGemma; taxonomy design; self-hosting a guard model |
Deploy a guard model alongside the main model |
Gen: parallel guard path |
Guard model serving with measured overhead |
| M2T3 |
Measuring a guardrail honestly |
False-positive vs false-negative rates, the usability cost of over-blocking, threshold tuning, per-category performance |
Guardrail eval on a labelled adversarial + benign set |
Chart: ROC per category, with the chosen threshold marked |
Your guardrail's real error rates, published |
| M2T4 |
PII and data handling |
Detection and redaction, data-residency, log hygiene (your traces contain user data), retention policy |
PII redaction in the trace pipeline |
Mermaid: data flow with redaction points |
Traces verified free of PII |
| M2T5 |
Structural defences |
Constrained decoding as a safety control, tool-call allowlists and authorization, sandboxed execution, least privilege for agents |
Sandboxed tool executor with an authorization layer |
Gen: privilege boundaries around the agent |
A tool-abuse attempt blocked by design, not by a filter |
| M2T6 |
Abuse prevention & rate limiting |
Per-tenant quotas, token budgets, anomaly detection, cost-based abuse (someone burning your GPUs) |
Quota + anomaly detection on the gateway |
Chart: cost per tenant with an anomaly flagged |
A simulated abuse pattern detected |
M3 — Supply Chain, Licensing & Compliance
| ID |
Topic |
Key concepts |
Code artifact |
Visual |
Evidence |
| M3T1 |
Model supply chain |
Weight provenance, safetensors vs pickle, trust_remote_code as arbitrary code execution, checksum verification, registry hygiene |
Model verification pipeline with a pickle-payload demo |
Gen: supply chain from upstream to serving |
A malicious checkpoint caught by your pipeline |
| M3T2 |
Licensing reality |
Apache-2.0 vs MIT vs Llama Community vs Gemma vs research-only; derivative works from fine-tuning; synthetic data from a teacher model's terms of service |
Licence auditor covering models, datasets, and outputs |
Mermaid: licence compatibility decision tree |
A clean licence audit for your capstone model |
| M3T3 |
Documentation as governance |
Model cards, datasheets, evaluation reports, SBOM, reproducibility manifests |
Auto-generated model card from your training run |
Mermaid: artifact lineage graph |
A complete, honest model card |
| M3T4 |
The regulatory layer |
EU AI Act obligations for providers vs deployers, risk classification, transparency duties, sector rules (health, finance), where self-hosting helps and where it doesn't |
Compliance checklist mapped to your artifacts |
Mermaid: obligation → evidence mapping |
A gap analysis with named owners |
| M3T5 |
Operational security |
Secrets, network isolation for GPU nodes, RBAC on the platform, audit logging, incident response for AI-specific incidents |
Hardened deployment config + audit trail |
Mermaid: incident response flow |
A tabletop incident exercise, written up |
Labs
| Lab |
Goal |
Success criterion |
| LAB-M1 |
Attack your own stack |
Successfully execute indirect prompt injection, a jailbreak, and a resource-exhaustion attack |
| LAB-M2 |
Defend it |
Block all three, with false-positive rate on benign traffic below a stated threshold |
| LAB-M3 |
Audit the chain |
Full licence and provenance audit of your capstone model, with one issue found and resolved |
Capstone
Deliverable: security-and-governance-pack — the document set that lets your platform pass review.
- Threat model of the full platform, with ranked risks and named mitigations
- Red-team report: attacks attempted, which succeeded, evidence
- Guardrail architecture with measured FP/FN rates per category, and the accepted tradeoff
- PII handling policy, with traces verified clean
- Model supply-chain verification pipeline
- Licence and provenance audit covering models, datasets, and synthetic data
- Model card and datasheet for your Skill 05 model
- Regulatory gap analysis with owners and dates
- Incident response runbook for AI-specific incidents
- A one-page risk acceptance statement — what you're choosing not to mitigate, and why
Item 10 separates security theatre from security engineering. Every real system has accepted risks;
maturity is naming them explicitly.
Assessment
| Tier |
Count |
Example |
| Recall |
6 |
What makes indirect prompt injection harder to defend than direct? |
| Apply |
10 |
Your guardrail blocks 4% of legitimate traffic. Walk through how you'd decide whether that's acceptable |
| Design |
4 |
Design the security architecture for a self-hosted agent with database and email tool access |
Practical challenge: given an unfamiliar deployed LLM service, produce a threat model and a
working exploit within 3 hours.
Asset inventory
| Type |
Count (minimum) |
| Theory articles |
16 |
| Code artifacts |
15 |
| Mermaid diagrams |
~20 |
| Generated images |
~55 |
| Charts |
~12 |
| Labs |
3 |
| Capstone |
1 |
Primary sources
- OWASP Top 10 for LLM Applications + the OWASP Agentic AI threat guidance
- NIST AI Risk Management Framework and the Generative AI Profile
- Llama Guard / Granite Guardian / ShieldGemma model cards; NVIDIA NeMo Guardrails docs
- Papers: Not what you've signed up for (indirect prompt injection), Universal adversarial suffixes (GCG), Extracting Training Data from LLMs
- MITRE ATLAS — adversarial ML threat matrix
- EU AI Act text (Articles on GPAI, transparency, and deployer obligations)
- Model Cards for Model Reporting (Mitchell et al.); Datasheets for Datasets (Gebru et al.)
- Hugging Face security docs — safetensors, malware scanning,
trust_remote_code guidance