Back to tags
#vllm
4 blog posts.
Four posts on running vLLM in production, all from the GPUs in production series. Inside one box, tensor parallelism only goes as fast as the interconnect underneath it. On the dashboard, the metrics that matter are queue time and KV-cache utilisation, not GPU utilisation. In front of the fleet, a router that reads the KV cache beats a round-robin load balancer. And at the pod level, a liveness probe that fires while weights are still loading will crash-loop a server that was never broken.
Blog posts
A model server that takes five minutes to load and a liveness probe that gives it ten seconds is a crash loop waiting to happen. Probes, drains, and safe…
Same GPUs, same model, same replica count. Swap round-robin for prefix-cache-aware routing and the fleet gets 2.3x faster. The router was throwing the…
The GPU dashboard says 92% busy and users are waiting eight seconds for the first token. Monitoring an LLM server means watching the queue, not the…
Eight GPUs in one server behave like a small network. NVLink vs PCIe, reading nvidia-smi topo -m, NCCL transports, the ACS trap, and fitting a 70B model.