Skip to content

Skill 04 — Data for Post-Training

The least glamorous skill with the highest real leverage. The quality of a fine-tune is ~80% data. Everyone wants to skip this module. Everyone who skips it produces mediocre models and can't explain why.

Duration 2 weeks (~25 hrs)
Impact ★★★★☆
Prerequisites Skill 01 (you need a serving stack to generate synthetic data), Skill 02 M1 (tokenization)
Hardware 1–2 GPUs for generation; CPU-heavy (our 24 cores are the constraint — pre-tokenize to disk)
Status planned

Outcome

  • You can build a small, excellent dataset instead of a large, mediocre one — and know which you have
  • You can generate synthetic data with a teacher model and filter it well enough to be worth training on
  • You can prove your eval sets aren't contaminated by your training data
  • You can look at a failed fine-tune and correctly attribute it to data rather than hyperparameters

Modules

M Module Hrs Theme
M1 Dataset Anatomy & Formats 7 Getting the mechanics exactly right
M2 Curation & Quality 9 Turning raw data into training data
M3 Synthetic Data Generation 9 Manufacturing what you don't have

M1 — Dataset Anatomy & Formats

ID Topic Key concepts Code artifact Visual Evidence
M1T1 The formats and when each applies Raw text (CPT), instruct pairs, ShareGPT/messages, preference triples, RLVR task+verifier, tool-call traces Converter between all six formats, with validation Mermaid: format → training method map A validated converter with a test suite
M1T2 Chat templating & loss masking Applying the model's own template, assistant-only loss, multi-turn masking, system prompt handling, the tool-call mask Mask visualizer: render a conversation with masked tokens highlighted Gen: token stream with masked/unmasked regions (Pattern A) A correct multi-turn mask, verified token by token
M1T3 Packing & sequence length Multipack/bin-packing, cross-contamination between packed samples, attention masking within packs, padding waste Packer with correct intra-pack attention isolation Chart: padding waste vs packing strategy Measured throughput gain + proof of no cross-contamination
M1T4 Storage & throughput Pre-tokenization to disk, Arrow/Parquet, streaming vs in-memory, dataloader workers on a 24-core box Pre-tokenization pipeline with a resumable streaming loader Mermaid: data pipeline stages Tokens/s sustained; GPU never starved

M1T2 and M1T3 are where silent, hard-to-detect bugs live. Both topics must include a reproduction of the bug before the fix.


M2 — Curation & Quality

ID Topic Key concepts Code artifact Visual Evidence
M2T1 Sourcing & licensing Open datasets, licence compatibility, provenance tracking, what you can legally train on and ship Dataset licence auditor Mermaid: licence compatibility decision tree A provenance manifest for your own dataset
M2T2 Deduplication Exact, MinHash/LSH near-dup, semantic dedup via embeddings; why dupes cause memorization and wasted compute Dedup pipeline: exact → MinHash → embedding Gen: clustering of near-duplicates (Pattern A) Duplicate rate found in a real public dataset (it will surprise you)
M2T3 Quality filtering Heuristic filters, perplexity filtering, classifier-based filtering, LLM-as-judge scoring, length and format heuristics Multi-stage filter with per-stage retention stats Chart: retention funnel by stage A funnel showing what each stage removed and why
M2T4 Decontamination N-gram overlap against eval sets, why benchmark contamination invalidates everything, how to audit and report it Contamination checker against your eval suite Chart: contamination rate by benchmark A signed contamination report for your dataset
M2T5 Mixture design Domain ratios, replay to prevent forgetting, curriculum ordering, upsampling scarce high-value data Mixture builder with target-ratio verification Gen: proportional mixture composition Two mixtures trained, capability differences measured
M2T6 Small-and-excellent vs large-and-noisy The LIMA argument, diminishing returns, how to know when to stop collecting Ablation: 1k curated vs 50k raw samples, same budget Chart: eval score vs dataset size, both tracks Direct evidence for the quality-over-quantity claim

M2T4 is non-negotiable. A model evaluated on contaminated benchmarks is worse than unevaluated, because it's confidently wrong.


M3 — Synthetic Data Generation

ID Topic Key concepts Code artifact Visual Evidence
M3T1 Generation strategies Self-instruct, Evol-Instruct, persona/seed-driven diversity, distillation from a stronger teacher, licence implications of teacher outputs Generation pipeline running against your own vLLM server Mermaid: seed → expand → generate → filter 5k synthetic samples with measured diversity
M3T2 Diversity & collapse Mode collapse, n-gram and embedding diversity metrics, temperature/seeding strategies, why naive generation produces homogeneous slop Diversity scorer + a deliberate collapse demonstration Chart: diversity metrics vs generation strategy Collapse reproduced, then fixed
M3T3 Rejection sampling & verification Generate-many-keep-best, verifiable rewards (code tests, math checkers, schema validators), LLM-judge filtering Rejection sampling loop with a real verifier Gen: many candidates, few survivors (Pattern C) Acceptance rate + measured quality lift vs unfiltered
M3T4 Preference data Building chosen/rejected pairs, on-policy vs off-policy pairs, human vs AI feedback, annotation interfaces Preference pair builder from generation + judge scores Mermaid: pair construction flow A preference dataset with judge-vs-human agreement measured
M3T5 RLVR task data Task + verifier + reward design; why verifiable tasks are the highest-signal data; reward hacking at the data layer Task suite with programmatic verifiers Mermaid: task → rollout → verify → reward A verifier suite with a documented gaming attempt
M3T6 Cost & throughput of generation Batch offline inference, prefix caching for shared seeds, $/1k samples, when to buy vs generate Generation cost model reusing Skill 01 M3 Chart: $/1k accepted samples vs strategy Your own generation cost curve

Labs

Lab Goal Success criterion
LAB-M1 Mask forensics Given three datasets with deliberately broken masks/templates, find and fix all three
LAB-M2 Clean a public dataset Dedup + filter + decontaminate a real dataset; report the funnel and the contamination found
LAB-M3 Manufacture a dataset 5k synthetic samples for a chosen domain, passing diversity and verification thresholds

Capstone

Deliverable: dataset-with-a-datasheet — a training-ready dataset for one real task, plus the documentation that makes it trustworthy.

Required: 1. The dataset itself, in a validated format, correctly templated and masked 2. A datasheet: provenance, licence, construction method, size, domain composition 3. Curation funnel with retention statistics per stage 4. Diversity metrics 5. Contamination audit against your eval suite, signed and dated 6. A held-out eval split with justification for how it was chosen 7. An ablation proving your curation improved a downstream fine-tune vs the raw data

Item 7 is the whole point. Without it you have a file, not a dataset.

This dataset is consumed directly by Skill 05 — build it for real.


Assessment

Tier Count Example
Recall 5 What is assistant-only loss masking and what happens without it?
Apply 10 This fine-tune degraded on general tasks while improving on the target. Give three data-side causes
Design 4 Design the data strategy for a domain assistant where you have 200 real examples and no budget for annotation

Practical challenge: given a raw scraped corpus, produce a training-ready dataset with a full datasheet in 3 hours.


Asset inventory

Type Count
Theory files 16
Code artifacts 16
Mermaid diagrams ~14
Generated images ~8
Charts ~14
Labs 3
Capstone 1

Primary sources

  • Axolotl dataset-format docs + the multipack write-up
  • HF TRL dataset formats conceptual guide; smol-course data chapters
  • Papers: Self-Instruct, Evol-Instruct/WizardLM, LIMA, Tulu 3 (data curation sections), Phi technical reports (textbook-quality data argument)
  • FineWeb / FineWeb-Edu and DCLM papers — the best public writeups of large-scale filtering
  • Datasheets for Datasets (Gebru et al.)
  • Contamination: the LM Contamination Index and the decontamination methodology in the DCLM paper