Skill 04 — Data for Post-Training¶
The least glamorous skill with the highest real leverage. The quality of a fine-tune is ~80% data. Everyone wants to skip this module. Everyone who skips it produces mediocre models and can't explain why.
| Duration | 2 weeks (~25 hrs) |
| Impact | ★★★★☆ |
| Prerequisites | Skill 01 (you need a serving stack to generate synthetic data), Skill 02 M1 (tokenization) |
| Hardware | 1–2 GPUs for generation; CPU-heavy (our 24 cores are the constraint — pre-tokenize to disk) |
| Status | planned |
Outcome¶
- You can build a small, excellent dataset instead of a large, mediocre one — and know which you have
- You can generate synthetic data with a teacher model and filter it well enough to be worth training on
- You can prove your eval sets aren't contaminated by your training data
- You can look at a failed fine-tune and correctly attribute it to data rather than hyperparameters
Modules¶
| M | Module | Hrs | Theme |
|---|---|---|---|
| M1 | Dataset Anatomy & Formats | 7 | Getting the mechanics exactly right |
| M2 | Curation & Quality | 9 | Turning raw data into training data |
| M3 | Synthetic Data Generation | 9 | Manufacturing what you don't have |
M1 — Dataset Anatomy & Formats¶
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M1T1 | The formats and when each applies | Raw text (CPT), instruct pairs, ShareGPT/messages, preference triples, RLVR task+verifier, tool-call traces | Converter between all six formats, with validation | Mermaid: format → training method map | A validated converter with a test suite |
| M1T2 | Chat templating & loss masking | Applying the model's own template, assistant-only loss, multi-turn masking, system prompt handling, the tool-call mask | Mask visualizer: render a conversation with masked tokens highlighted | Gen: token stream with masked/unmasked regions (Pattern A) | A correct multi-turn mask, verified token by token |
| M1T3 | Packing & sequence length | Multipack/bin-packing, cross-contamination between packed samples, attention masking within packs, padding waste | Packer with correct intra-pack attention isolation | Chart: padding waste vs packing strategy | Measured throughput gain + proof of no cross-contamination |
| M1T4 | Storage & throughput | Pre-tokenization to disk, Arrow/Parquet, streaming vs in-memory, dataloader workers on a 24-core box | Pre-tokenization pipeline with a resumable streaming loader | Mermaid: data pipeline stages | Tokens/s sustained; GPU never starved |
M1T2 and M1T3 are where silent, hard-to-detect bugs live. Both topics must include a reproduction of the bug before the fix.
M2 — Curation & Quality¶
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M2T1 | Sourcing & licensing | Open datasets, licence compatibility, provenance tracking, what you can legally train on and ship | Dataset licence auditor | Mermaid: licence compatibility decision tree | A provenance manifest for your own dataset |
| M2T2 | Deduplication | Exact, MinHash/LSH near-dup, semantic dedup via embeddings; why dupes cause memorization and wasted compute | Dedup pipeline: exact → MinHash → embedding | Gen: clustering of near-duplicates (Pattern A) | Duplicate rate found in a real public dataset (it will surprise you) |
| M2T3 | Quality filtering | Heuristic filters, perplexity filtering, classifier-based filtering, LLM-as-judge scoring, length and format heuristics | Multi-stage filter with per-stage retention stats | Chart: retention funnel by stage | A funnel showing what each stage removed and why |
| M2T4 | Decontamination | N-gram overlap against eval sets, why benchmark contamination invalidates everything, how to audit and report it | Contamination checker against your eval suite | Chart: contamination rate by benchmark | A signed contamination report for your dataset |
| M2T5 | Mixture design | Domain ratios, replay to prevent forgetting, curriculum ordering, upsampling scarce high-value data | Mixture builder with target-ratio verification | Gen: proportional mixture composition | Two mixtures trained, capability differences measured |
| M2T6 | Small-and-excellent vs large-and-noisy | The LIMA argument, diminishing returns, how to know when to stop collecting | Ablation: 1k curated vs 50k raw samples, same budget | Chart: eval score vs dataset size, both tracks | Direct evidence for the quality-over-quantity claim |
M2T4 is non-negotiable. A model evaluated on contaminated benchmarks is worse than unevaluated, because it's confidently wrong.
M3 — Synthetic Data Generation¶
| ID | Topic | Key concepts | Code artifact | Visual | Evidence |
|---|---|---|---|---|---|
| M3T1 | Generation strategies | Self-instruct, Evol-Instruct, persona/seed-driven diversity, distillation from a stronger teacher, licence implications of teacher outputs | Generation pipeline running against your own vLLM server | Mermaid: seed → expand → generate → filter | 5k synthetic samples with measured diversity |
| M3T2 | Diversity & collapse | Mode collapse, n-gram and embedding diversity metrics, temperature/seeding strategies, why naive generation produces homogeneous slop | Diversity scorer + a deliberate collapse demonstration | Chart: diversity metrics vs generation strategy | Collapse reproduced, then fixed |
| M3T3 | Rejection sampling & verification | Generate-many-keep-best, verifiable rewards (code tests, math checkers, schema validators), LLM-judge filtering | Rejection sampling loop with a real verifier | Gen: many candidates, few survivors (Pattern C) | Acceptance rate + measured quality lift vs unfiltered |
| M3T4 | Preference data | Building chosen/rejected pairs, on-policy vs off-policy pairs, human vs AI feedback, annotation interfaces | Preference pair builder from generation + judge scores | Mermaid: pair construction flow | A preference dataset with judge-vs-human agreement measured |
| M3T5 | RLVR task data | Task + verifier + reward design; why verifiable tasks are the highest-signal data; reward hacking at the data layer | Task suite with programmatic verifiers | Mermaid: task → rollout → verify → reward | A verifier suite with a documented gaming attempt |
| M3T6 | Cost & throughput of generation | Batch offline inference, prefix caching for shared seeds, $/1k samples, when to buy vs generate | Generation cost model reusing Skill 01 M3 | Chart: $/1k accepted samples vs strategy | Your own generation cost curve |
Labs¶
| Lab | Goal | Success criterion |
|---|---|---|
| LAB-M1 | Mask forensics | Given three datasets with deliberately broken masks/templates, find and fix all three |
| LAB-M2 | Clean a public dataset | Dedup + filter + decontaminate a real dataset; report the funnel and the contamination found |
| LAB-M3 | Manufacture a dataset | 5k synthetic samples for a chosen domain, passing diversity and verification thresholds |
Capstone¶
Deliverable: dataset-with-a-datasheet — a training-ready dataset for one real task, plus the documentation that makes it trustworthy.
Required: 1. The dataset itself, in a validated format, correctly templated and masked 2. A datasheet: provenance, licence, construction method, size, domain composition 3. Curation funnel with retention statistics per stage 4. Diversity metrics 5. Contamination audit against your eval suite, signed and dated 6. A held-out eval split with justification for how it was chosen 7. An ablation proving your curation improved a downstream fine-tune vs the raw data
Item 7 is the whole point. Without it you have a file, not a dataset.
This dataset is consumed directly by Skill 05 — build it for real.
Assessment¶
| Tier | Count | Example |
|---|---|---|
| Recall | 5 | What is assistant-only loss masking and what happens without it? |
| Apply | 10 | This fine-tune degraded on general tasks while improving on the target. Give three data-side causes |
| Design | 4 | Design the data strategy for a domain assistant where you have 200 real examples and no budget for annotation |
Practical challenge: given a raw scraped corpus, produce a training-ready dataset with a full datasheet in 3 hours.
Asset inventory¶
| Type | Count |
|---|---|
| Theory files | 16 |
| Code artifacts | 16 |
| Mermaid diagrams | ~14 |
| Generated images | ~8 |
| Charts | ~14 |
| Labs | 3 |
| Capstone | 1 |
Primary sources¶
- Axolotl dataset-format docs + the multipack write-up
- HF TRL dataset formats conceptual guide;
smol-coursedata chapters - Papers: Self-Instruct, Evol-Instruct/WizardLM, LIMA, Tulu 3 (data curation sections), Phi technical reports (textbook-quality data argument)
- FineWeb / FineWeb-Edu and DCLM papers — the best public writeups of large-scale filtering
- Datasheets for Datasets (Gebru et al.)
- Contamination: the LM Contamination Index and the decontamination methodology in the DCLM paper