Back to tags

#inference

3 blog posts.

Three posts on serving LLM inference rather than training it. The observability part explains why GPU utilisation is a bad health metric and what to watch instead — time to first token, queue depth, KV-cache pressure. Scale-to-zero covers the money argument for idling a GPU fleet and the cold-start work that makes it survivable in front of real users. And the routing part makes the case that the cheapest speedup available is usually the load balancer, not a bigger card.

Blog posts

Same GPUs, same model, same replica count. Swap round-robin for prefix-cache-aware routing and the fleet gets 2.3x faster. The router was throwing the…
Idle GPUs at six dollars an hour are a bonfire. Scaling to zero saves the money, but the first user back waits minutes unless you kill the cold start.
The GPU dashboard says 92% busy and users are waiting eight seconds for the first token. Monitoring an LLM server means watching the queue, not the…

Related tags

LLM inference: 3 posts on serving at scale · Infra Magician