Skip to content

T1.3 — Precision Formats and What They Cost

In one line: Precision is a resource allocation decision: spend bits on range, resolution, or throughput, and know which failure mode you bought.

Skill 00 — Foundations
Module M1 — GPU & Systems Literacy
Time 90 min
Prereqs Memory hierarchy, roofline arithmetic intensity
Status draft

Why this matters

Precision formats decide how many bytes you move, which tensor cores you can use, and how much numerical error the model must absorb. In serving, that decision is not cosmetic. A 7B model in bf16 needs roughly 14 GB for weights; the same weights in int8 need about 7 GB, and int4 gets close to 3.5 GB before scales and packing overhead. That can be the difference between one GPU and two, or between a low-batch decode path that is bandwidth-bound and one that has enough headroom to meet a streaming SLO.

The trap is treating precision as a ladder where fewer bits are always better. They are not. Lower precision can shift the roofline, unlock tensor-core paths, and reduce HBM traffic, but it can also add dequantization work, introduce scale tensors, damage outlier channels, or leave the actual bottleneck untouched. If decode is dominated by KV-cache reads, shrinking only the weights may not move the number you care about.

The mental model

A floating-point format splits a fixed number of bits between sign, exponent, and mantissa. The exponent buys dynamic range: how large and tiny values can be. The mantissa buys resolution: how finely values are spaced inside that range. Integer quantization moves the range decision out of the value itself and into a scale, usually per tensor, per channel, per group, or per block.

Bit budget moves between range, resolution, and storage.

The serving question is therefore not "what is the smallest format?" It is "where does this tensor need range, where does it need resolution, and what hardware path will execute it?" Weights, activations, KV cache, logits, optimizer state, and accumulated reductions all have different answers. Training is conservative because small errors compound through gradients. Inference can be more aggressive because weights are fixed and you can calibrate against real activations.

The mechanism

Start with the formats you will see in production:

  • fp32: 4 bytes/value. Excellent range and resolution, rarely appropriate for inference except accumulators, numerically sensitive reductions, or reference checks.
  • tf32: NVIDIA's tensor-core compromise for fp32-looking matmuls. It keeps fp32 range but uses fewer mantissa bits. Useful when code asks for fp32 but hardware can trade precision for speed.
  • bf16: 2 bytes/value with fp32-like exponent range and fewer mantissa bits. This is the default safe serving format because it rarely overflows.
  • fp16: 2 bytes/value with more mantissa than bf16 but much less exponent range. Fast, but easier to overflow.
  • fp8: 1 byte/value families such as E4M3 and E5M2. Useful on newer silicon, but the exact recipe is hardware- and calibration-dependent.
  • int8 / int4: integer values plus scale metadata. The value is small; the scaling policy is the model.

Range and resolution are separate tradeoffs, not one quality slider.

The critical difference between floating point and quantization is where the scale lives. In bf16, every value carries its own exponent. In int8, a group of values shares a scale. If the group contains one outlier, either the outlier clips or the normal values lose resolution. That is why per-channel and group-wise quantization usually beat one global scale: they spend a little metadata to avoid letting one channel set the range for everyone.

real_value ≈ scale[group] × integer_value

That approximation has two costs. First, storage is not exactly n_values × bits; scale tensors and packing overhead matter. Second, execution is not automatically faster. Some kernels dequantize weights into registers and then run a higher-precision matmul. Others use native low-precision tensor cores. The difference determines whether quantization is a bandwidth win, a compute win, or mostly a capacity win.

Scale granularity decides whether outliers poison the whole tensor or only a small group.

Roofline thinking gives you the first-pass prediction. If an operation is bandwidth-bound, reducing bytes can help almost linearly until you hit another bottleneck. If it is compute-bound, lower precision helps only if it maps to a faster compute path. If it adds dequantization around an already compute-bound kernel, it can lose.

Lower precision can shift the ridge point, but only if the kernel uses the matching hardware path.

In practice

For LLM serving, the default hierarchy is simple:

  1. Serve a baseline in bf16. It is the measurement reference.
  2. Try int8 weight-only or W8A8 when you need capacity and the hardware has mature kernels.
  3. Try W4A16 when low-batch latency or model fit is the problem.
  4. Treat fp8 as a newer-silicon optimization, not an Ampere success criterion.
  5. Quantize KV cache only after measuring quality and memory pressure; it directly affects every generated token.

Weight-only quantization is attractive because weights dominate model storage and are read repeatedly during decode. At batch size 1, reading fewer weight bytes can improve TPOT. At large batch, each weight read is amortized across many tokens, so the win shrinks. W8A8 can improve throughput if activations are also quantized and the matmul uses efficient int8 kernels, but activation calibration is harder because the runtime distribution depends on prompts.

KV-cache quantization is a different lever. The cache grows with batch × sequence length × layers × KV heads × head_dim. When context length or concurrency is the limiter, shrinking weights does not fix the capacity problem. INT8 KV cache may be the right trade, but the error sits inside attention scores for all future tokens, so it needs task-level evaluation, not just perplexity on a tiny sample.

Shrinking weights does not shrink the KV cache; the bottleneck can move.

Failure modes

  • Symptom: int4 looks great in memory but slower in throughput. Cause: the kernel dequantizes and runs a slower path, or batching made weights no longer the bottleneck. Fix: compare kernel traces and roofline position before keeping the format.
  • Symptom: model quality collapses on rare prompts. Cause: outlier channels clipped by coarse scales. Fix: use per-channel or group-wise scales, preserve sensitive layers in bf16, and calibrate on real traffic.
  • Symptom: fp16 inference occasionally emits NaNs while bf16 does not. Cause: exponent range, not mantissa resolution. Fix: use bf16 or keep numerically sensitive ops in fp32/bf16.
  • Symptom: KV-cache quantization saves memory but hurts long-context tasks. Cause: quantization noise accumulates in attention over many decode steps. Fix: test long prompts and consider per-head/per-channel scaling.

Do it

Build the numeric-range explorer for this topic: take one real weight tensor, encode it as bf16, fp16, int8, and int4 with group scales, then plot error and storage. The measurable success criterion is a table that reports bytes/value, max error, mean absolute error, and whether the format changes the top-1 output on a fixed prompt at temperature 0.

Check

  1. Why can bf16 be safer than fp16 even though both use 16 bits?
  2. When does W4A16 help latency but not high-batch throughput?
  3. Why is KV-cache quantization not equivalent to weight quantization?

Going deeper

  • NVIDIA — A100 Tensor Core GPU Architecture whitepaper.
  • NVIDIA Transformer Engine documentation for fp8 recipes on supported hardware.
  • Dettmers et al. — LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.
  • Frantar et al. — GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.
  • Lin et al. — AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.