vLLM, TGI (Text Generation Inference) and SGLang are open-source inference engines used to run a large language model quickly and efficiently on GPUs, in production.
Running an LLM naively wastes a lot of memory and compute time. These engines apply optimizations such as continuous batching of requests, fine-grained GPU memory management (vLLM popularized the PagedAttention technique), support for quantized models or splitting a model across several cards. They usually expose an API compatible with OpenAI's.
vLLM came out of research work at UC Berkeley, TGI is developed by Hugging Face, and SGLang also comes from academic research. They are core building blocks of inference platforms built on open-weight models.
Why it matters when hiring
Having contributed code to one of these projects is a very good signal: it shows a fine understanding of GPUs, memory and performance, publicly verifiable on GitHub. Having only launched a vLLM server with the default configuration is a good start, not expertise. Ask which parameters the candidate tuned, and for what measured gain.
