LLMKube

door defilantech · open-source AI & ML · Grote taalmodellen

58/100 Self-hostable

Open in Open Source Atlas

Beschrijving

Kubernetes operator for self-hosted LLM inference with pluggable runtimes (llama.cpp, vLLM, TGI, Ollama, vllm-swift), multi-GPU sharding, NVIDIA CUDA + Apple Silicon Metal support, and OpenAI-compatible API.

Atlas-score

Soevereiniteit 65
Goed onderhouden 100
Beveiligd onbekend
Populair 12
Open standaarden 0

Details

Links

Alternatieven in Grote taalmodellen

Open in Open Source Atlas