Beschrijving
llm-d is a Kubernetes-native, high-performance distributed LLM inference framework built on vLLM and the Kubernetes Gateway API Inference Extension, providing intelligent inference scheduling, prefix-cache-aware routing, prefill/decode disaggregation, hierarchical KV offloading, and traffic- and hardware-aware autoscaling across NVIDIA, AMD, Intel, and Google TPU accelerators.
Atlas-score
- CNCF landscape project (sandbox)
- Apache-2.0 (permissive)
- Last commit 0 days ago
- Released within the last 12 months
- Top 41% by stars in DevOps & Infrastructure::Container Orchestration