Quantization is a technique that reduces the precision of the numbers used to store an AI model's parameters, so that it takes less memory and runs faster.
A model stores its weights as floating-point numbers, often on 16 or 32 bits. Quantization brings them down to 8 bits, 4 bits or even less. The model becomes lighter, fits on a more modest GPU or a laptop, and responds faster. The trade-off: a possible loss of quality, which has to be measured for each use case.
There are several methods (GPTQ, AWQ, GGUF formats for local execution, etc.), applied after training or taken into account during it. Quantization is widely used to serve open-weight models in production and for embedded AI.
Why it matters when hiring
Quantizing a model with an off-the-shelf tool takes a few minutes. Knowing whether the quality loss is acceptable for a specific use case requires a real evaluation protocol, and that is where skill shows. Ask: how did the candidate measure the impact of quantization, on which tasks, with what gain in cost or latency? This know-how is found among LLM engineers and inference engineers.
