The article explains how vLLM can run GGUF‑formatted models via its official vllm‑gguf‑plugin, which currently supports only GPU execution on NVIDIA and AMD devices; CPU and Intel GPU paths are absent. To serve a model, the command uses the repo:quant_type syntax—for example, vllm serve unsloth/Qwen3-0.6B-GGUF:Q4_K_M --tokenizer Qwen/Qwen3-0.6B —where the --tokenizer flag must point to the original model repository, not the GGUF file, because vLLM loads the tokenizer separately from the weights. Supported quantizations include the commonly used K‑quants such as Q4_K_M, while more exotic schemes rely on llama.cpp as the reference implementation and may not yet be plugin‑compatible. The documentation labels GGUF support in vLLM as highly experimental and under‑optimized, emphasizing that the plugin is intended for reusing existing GGUF files rather than achieving speed gains; llama.cpp’s memory‑mapped loading remains faster for single‑user scenarios. For multi‑user workloads, vLLM’s continuous batching enables one GPU to serve concurrent requests from a single GGUF checkpoint without reloading the model. Failure modes typically stem from incorrect serve syntax, missing or wrong tokenizer specifications, unsupported hardware, or unsupported quantization types. The piece contrasts vLLM’s GPU‑focused, multi‑user approach with llama.cpp’s CPU‑first, flexible‑quantization design.
Read original
dev.to