LLMKube
door defilantech · open-source AI & ML · Grote taalmodellen
58/100
Self-hostable
Beschrijving
Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.
Atlas-score
- Self-hostable
- Apache-2.0 (permissive)
- Last commit 0 days ago
- Released within the last 12 months
- Top 88% by stars in DevOps & Infrastructure::Container Orchestration