Every issue in this series keeps returning to one argument: raw capacity never explains the outcome by itself, the system built around that capacity does. Model compression for edge deployment is where that argument gets tested in miniature, because two genuinely different techniques get called by the same casual phrase, "shrink the model," and teams that treat them as interchangeable end up solving the wrong problem for their hardware.
Quantization keeps a model's architecture exactly as trained and stores its numbers with less precision, the way rounding every measurement to fewer decimal places saves space without changing what's being measured. Distillation trains an entirely new, smaller model to imitate a larger one's behavior, including the parts of that behavior a simple right-or-wrong label never captures. One is a deployment-time transformation you can apply in an afternoon with no training pipeline at all. The other is closer in cost and complexity to a real fine-tuning project. Picking between them by vibes, rather than by which bottleneck you're actually up against, is how teams end up with a model that's technically smaller and still wrong for the hardware it has to run on.
~1 FLOP/Byte
Arithmetic Intensity
Batch-size-1 generation on an RTX 4090, far below its 165 TFLOPS peak
~87%
Size Reduction
Phi-3.5-mini, FP32 to INT4, no architecture change
97%
Capability Retained
DistilBERT vs. its BERT-base teacher, at 40% fewer parameters
Part 1
Why Edge Inference Is Memory-Bound, Not Compute-Bound
Start with the bottleneck itself, because the entire case for one quantization format over another rests on getting this right. A modern GPU's headline spec is its peak compute throughput, TFLOPS, floating-point operations per second. That number describes what the chip can do if it's never waiting on anything. Real-world autoregressive text generation at batch size one, one request being served at a time, which is the normal shape of on-device inference, spends most of its time waiting on something else entirely: getting the model's weights out of memory and onto the chip fast enough to use them.
The AWQ paper (Lin et al., 2023) profiled exactly this on an NVIDIA RTX 4090 running Llama-2-7B. The card offers 165 TFLOPS of FP16 compute. The arithmetic intensity actually achieved during generation, the ratio of math performed to bytes of memory moved, comes out to roughly one FLOP per byte. That is a roofline-model textbook case of a workload sitting deep in the memory-bound region: the chip could do vastly more math per byte than the workload asks it to, so the chip's cycles go to waiting, not computing. Breaking down where that memory traffic actually goes, the paper shows weight access dominates it overwhelmingly; activations are comparatively small.
Why this matters for which quantization format to pick: if weight traffic is the dominant cost and activations barely register, then compressing weights is where the real win lives. Compressing activations, the other half of a joint W8A8-style scheme, barely touches the actual bottleneck. This single observation is the entire technical justification for why weight-only quantization formats exist as a distinct category, not just a lazier version of full quantization.

Figure 1: Peak compute capacity vs. what batch-size-1 generation actually achieves. The gap is the memory wall, and it's why weight-only compression targets the real bottleneck.
Part 2
Quantization: Three Methods, One Format That Wins On-Device
Quantization notation describes bit-width for weights and activations separately: W8A8 means 8-bit weights and 8-bit activations, W4A16 means 4-bit weights kept alongside 16-bit (FP16) activations, weight-only compression. Given the memory-bound argument above, W8A8 is the right choice when the priority is maximizing throughput across many concurrent requests, the server-side, high-batch case where activation memory scales with batch size and compute genuinely becomes the constraint. W4A16 is the right choice when the priority is minimizing latency and memory footprint for one request at a time, the edge case.
Within W4A16, three distinct methods exist, and they are not interchangeable despite all producing "4-bit weights."
Round-to-nearest (RTN) is the simplest possible approach: pick a scale from a tensor's observed min and max, round every weight to the nearest representable level. It needs no calibration data at all, which sounds like an advantage until you notice why it's rarely used for large language models on its own. LLM weight matrices contain a small number of outlier values and a small fraction of channels that have an outsized effect on the model's output. Naive uniform rounding damages exactly those salient channels, producing accuracy loss disproportionate to the bits actually removed.
GPTQ (Frantar et al., 2022) was the first method to push GPT-scale models, up to 175 billion parameters, down to 3 to 4 bits per weight while holding accuracy close to the full-precision baseline; prior post-training methods could only hold accuracy at 8-bit. It works layer by layer, and within a layer it uses a greedy strategy informed by second-order (Hessian) information: at each step it picks the weight whose rounding will introduce the smallest error, quantizes it, then adjusts the remaining unquantized weights in that layer to compensate for the error just introduced. This is meaningfully more sophisticated than blind rounding, but it has a documented weakness: because it reconstructs against a specific calibration dataset, it can overfit to that dataset's domain. The AWQ paper's own cross-domain test found GPTQ's perplexity degraded by 2.3 to 4.9 points when evaluated outside its calibration domain, versus 0.5 to 0.6 points for AWQ under the same test.
AWQ (Lin et al., 2023) takes a different approach entirely: rather than reconstructing against calibration data, it identifies which weight channels are salient by looking at activation magnitude, not weight magnitude, a distinction the paper shows matters a great deal, since selecting by weight norm performs no better than random selection while selecting by activation magnitude does. The paper's finding is that only roughly 0.1 to 1 percent of channels are truly salient. Rather than keeping those channels in a separate, hardware-unfriendly higher precision, AWQ scales them up before quantizing and compensates by scaling the corresponding activations down at inference time, an equivalent reparameterization that keeps every weight uniformly 4-bit while protecting the channels that matter most. Because it needs no backpropagation and no reconstruction against a specific calibration set, it generalizes better across domains than GPTQ, which is exactly what that cross-domain perplexity gap demonstrates.
What this actually looks like in production, not just in a benchmark table: a Microsoft team fine-tuned Phi-3.5-mini, 3.8 billion parameters, for an in-car assistant's function-calling task, then quantized it with Microsoft's Olive toolchain. The model went from 15GB in PyTorch FP32, to 7.2GB in ONNX FP16, to roughly 2.0 to 2.1GB with GPTQ or AWQ INT4, a reduction the team describes as approximately 87 percent. On their actual downstream metric, exact-match accuracy on function calls rather than perplexity, GPTQ with a 256-sample calibration set scored 0.866 versus a 0.873 FP32 baseline, a real but modest cost. The team's own caveat is worth preserving: perplexity, the metric both GPTQ and AWQ optimize for directly, does not always predict downstream task performance, which is exactly why they measured exact match instead of trusting perplexity alone. They also hit real friction getting this onto actual NPU hardware: Qualcomm's NPU export support in Olive was still in beta at the time, with concrete device-mismatch and configuration errors, and didn't become officially supported until Olive 0.9.0 in May 2025, roughly two years after the underlying GPTQ and AWQ algorithms were published. Production tooling lagging the research by that long is itself a useful data point for anyone planning an edge deployment timeline.

Figure 2: Three quantization strategies compressing the same model, and the real size trajectory from a production deployment.
Part 3
Distillation: A Different Problem, Solved a Different Way
Distillation does not touch numeric precision at all. It trains a new, genuinely smaller model, the student, to imitate a larger, already-trained model, the teacher, and the reason it can preserve more capability than a size chart would predict comes down to what exactly the student is trained to imitate.
A trained model's raw output for a next-token or next-class prediction is a full probability distribution, not a single answer. Hinton, Vinyals, and Dean's original 2015 formulation gives the concrete example: a prediction might land at 92 percent "Paris," 5 percent "Lyon," 3 percent "France." A conventional hard label throws away everything except the top answer, "Paris," equals one, everything else equals zero. The 5 percent mass sitting on "Lyon" is not noise, it encodes real information about which alternatives the teacher considers plausible, and that signal is exactly what a student model learns from that a hard label never provides.
To make that signal usable, Hinton et al. introduce a temperature parameter in the softmax function used to produce the teacher's output distribution: raising the temperature flattens the distribution, making the relative sizes of the non-winning probabilities more visible to the student rather than getting rounded down toward invisibility. The student then trains against a combined objective, commonly framed as Total Loss equals alpha times a distillation loss (how well the student's own softened distribution matches the teacher's) plus one minus alpha times a standard loss against the real ground-truth label. Training on both the soft distribution and the real label outperforms training on either alone.
Two real examples show what this actually buys in practice. DistilBERT (Sanh et al., 2019) retains 97 percent of BERT-base's language understanding performance at 40 percent fewer parameters and roughly 60 percent faster inference, achieved by distilling during pre-training itself rather than only at task-specific fine-tuning time, and by combining the soft-label loss with a masked-language-modeling loss and a separate loss matching the teacher's and student's internal hidden-state geometry. TinyBERT-4 (Jiao et al., 2019) goes further on size, down to roughly 13.3 percent of BERT-base's parameter count, while still retaining more than 96.8 percent of teacher performance on the GLUE benchmark, by distilling not just the final output but intermediate transformer layers, mapping specific student layers to specific teacher layers during training.

Figure 3: What the student actually learns from the teacher, and what two real distillation efforts achieved.
Part 4
Which One You Actually Need, and Why the Order Matters
Quantization is the right first move when you already have a model whose capability you're satisfied with, you have no retraining budget or time, and the constraint is fitting an existing artifact onto memory-constrained hardware fast. It can typically be applied to a production model in hours: calibrate, quantize, export, done, with no training infrastructure beyond a small calibration set.
Distillation is the right move when the constraint isn't just memory, it's compute, or when you need a structurally smaller model rather than a numerically compressed version of the same one, and you can afford a real training pipeline: a frozen teacher, a student architecture decision, a data pipeline, and validation against both the teacher's ceiling and an un-distilled student's floor. That is a materially bigger commitment than a quantization pass, closer to a fine-tuning project than a deployment step.
In practice, the strongest edge deployments don't pick one, they sequence both: distill first to get a genuinely smaller architecture that preserves capability, then quantize that smaller model on top for the final memory and latency win. This combined ordering, sometimes framed as Pruning, then Knowledge Distillation, then Quantization, P-K-D-Q, shows up repeatedly in the compression literature, but it's worth being precise that the exact ordering is not settled consensus. At least one 2025 paper proposes distilling before pruning instead, and other published pipelines interleave the three techniques differently for specific edge hardware targets. Treat P-KD-Q as a commonly cited sequence, not the one correct answer, and validate the ordering against your own accuracy and latency numbers rather than assuming a paper's default order transfers cleanly to your model and hardware.
The Architect's Take
The same lesson this newsletter keeps returning to, now applied to model compression itself: the question is never "how do I make this smaller," it's "which resource am I actually out of."
Diagnose the bottleneck before picking the tool. Memory-bandwidth-bound and compute-bound are different problems with different fixes. Quantization fixes the first; it does very little for the second, since it doesn't touch the operation count, only the precision of the numbers involved.
"4-bit" is not one technique. RTN, GPTQ, and AWQ all produce 4-bit weights and behave differently under real-world distribution shift. The cross-domain perplexity gap between GPTQ and AWQ is not a rounding error, it's the direct consequence of one method reconstructing against a calibration set and the other identifying salience from activations instead.
Perplexity is not your actual metric. The Microsoft team's own lesson: the metric both major quantization methods optimize for doesn't reliably predict your downstream task's accuracy. Measure the thing you actually care about, on your actual task, before trusting a paper's perplexity table to tell you which method to ship.
Distillation buys you a different resource than quantization does. It costs real retraining time, and in exchange it can reduce both memory and compute, and preserve behavioral nuance a numeric rounding pass structurally cannot touch. Don't reach for it when a quantization pass would have solved the actual constraint faster and cheaper.
Production tooling lags research by years, plan for it. GPTQ and AWQ shipped in 2022 and 2023. Official NPU export support for one major production toolchain didn't land until 2025. Budget real integration time for edge-specific hardware paths, don't assume a paper's benchmark numbers imply your target chip is ready today.
The model was never just its parameter count. The deployment was never just the model. And it turns out which compression technique is "correct" depends entirely on which resource you were actually short on.
Sources & Further Reading
Memory-bandwidth-bound edge inference, RTX 4090/Llama-2-7B benchmark: Lin, Tang, Tang, Yang, Chen, Wang, Xiao, Dang, Gan, Han, "AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration" (arXiv:2306.00978, MLSys 2024 Best Paper)
GPTQ mechanism and cross-domain overfitting comparison: Frantar, Ashkboos, Hoefler, Alistarh, "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers" (arXiv:2210.17323, 2022)
Real production numbers, Phi-3.5-mini quantization, Olive toolchain, NPU deployment friction: Kosuke Fujimoto, "A practical guide to INT4 quantization for SLMs: GPTQ vs AWQ, Olive, and real-world results" (Data Science + AI at Microsoft, Medium, Feb 2026)
Distillation mechanics, temperature-softened softmax, combined loss: Hinton, Vinyals, Dean, "Distilling the Knowledge in a Neural Network" (arXiv:1503.02531, 2015)
DistilBERT results: Sanh, Debut, Chaumond, Wolf, "DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter" (arXiv:1910.01108, 2019)
TinyBERT results: Jiao et al., "TinyBERT: Distilling BERT for Natural Language Understanding" (arXiv:1909.10351, 2019)
Distillation workflow framing, decision heuristics: Redis Engineering, "Model distillation for LLMs: A practical guide to smaller, faster AI" (redis.io/blog, Feb 2026)
Nagarjun Rajendran · linkedin.com/in/nagarjunr

