Which objective should a post-training quantizer optimize? — evidence explorer

Every number on this page is recomputed in your browser from the per-window NLLs in the dataset. Read the article; code and log in the repository. Paired bootstrap over evaluation windows, 2 000 resamples; a negative Δ means the first arm is better.

Frozen protocol, 3.25 bits per weight, full wikitext-2 test (146 windows), one row per calibration draw. “Gap closed” is the fraction of the GPTQ-vs-fp16 loss the module post-pass recovers.

Qwen2.5-0.5B, calibration draw 0, the same frozen method at four exact bit widths (wikitext-2).

A 2.12-bpw per-block vector-quantized Qwen2.5-0.5B, plain vs with the post-pass. Without a prefix the worst windows all start with a digit-like token; with a fixed "\n\n" prefix the tail disappears and the post-pass wins on both sets.