Back to tags

#gpu

8 blog posts.

The GPUs in production series, eight parts, written while running LLM inference on Kubernetes. It starts at the dozen layers under a single GPU pod — driver, container toolkit, device plugin, scheduler — then works outward: NVLink and NCCL inside one box, InfiniBand and gang scheduling across many, what a green GPU dashboard hides, scaling to zero without a five-minute cold start, routing that beats a bigger card, the probe that crash-loops a healthy model server, and two tenants sharing one GPU with no wall between them.

Blog posts

A Kubernetes namespace isolates the API, not the silicon. Under time-slicing two teams share a physical GPU with no memory wall. Quotas, isolation, and…
A model server that takes five minutes to load and a liveness probe that gives it ten seconds is a crash loop waiting to happen. Probes, drains, and safe…
Same GPUs, same model, same replica count. Swap round-robin for prefix-cache-aware routing and the fleet gets 2.3x faster. The router was throwing the…
Idle GPUs at six dollars an hour are a bonfire. Scaling to zero saves the money, but the first user back waits minutes unless you kill the cold start.
The GPU dashboard says 92% busy and users are waiting eight seconds for the first token. Monitoring an LLM server means watching the queue, not the…
Past one node the network becomes the machine. InfiniBand vs RoCE, gang scheduling, FSDP and Megatron, and why a 16k-GPU cluster fails every three hours.
Eight GPUs in one server behave like a small network. NVLink vs PCIe, reading nvidia-smi topo -m, NCCL transports, the ACS trap, and fitting a 70B model.
A GPU pod sits on a dozen layers from silicon to scheduler, and each one fails its own way. Drivers, the container toolkit, MIG, DCGM, and the metrics…

Related tags

GPU infrastructure: 8 Kubernetes posts · Infra Magician