
Cloud Architecture
llm-d: The Kubernetes-Native Framework Solving Disaggregated LLM Inference Before Your GPU Budget Explodes
A principal architect's guide to llm-d, the CNCF sandbox project from IBM, Red Hat, and Google that disaggregates prefill and decode phases across GPU pools to fix the throughput wall that hits every high-concurrency LLM deployment.
