T
time-to-first-token
@patchy631
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
Apache-2.0 Commit 1 maand geleden ★ 960
Pijlers en weging
Waarom deze score
- Apache-2.0 (permissive)
- Last commit 43 days ago
- Top 61% by stars in AI & ML::AI Operations
Feiten
- Licentie
- Apache-2.0 (Permissief)
- Taal
- HTML
- GitHub-sterren
- 960
- Laatste commit
- 2026-08-14 (1 maand geleden)
- Laatste release
- Onbekend
- OpenSSF Scorecard
- Nog niet gemeten
- Zelf hosten
- Niet vastgesteld
- Maintainer-locatie
- Onbekend
- Bron
- github-search
Alternatieven in AI-operations
01 77 02 76 03 72 04 71
Q
qdrant
Qdrant - High-performance, massive-scale Vector Database and Vector Search Engine for the next generation of AI. Also available in the cloud https://cloud.qdrant.io/
V
vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
A
argo-workflows
Workflow Engine for Kubernetes
P
pipelines
Machine Learning Pipelines for Kubeflow