S
SageAttention
@thu-ml
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Apache-2.0 Commit 8 maanden geleden ★ 3.950
Pijlers en weging
Waarom deze score
- Apache-2.0 (permissief)
- Laatste commit 253 dagen geleden
- Top 35% naar sterren in AI & ML › Grote taalmodellen
Feiten
- Licentie
- Apache-2.0 (Permissief)
- Taal
- Cuda
- GitHub-sterren
- 3.950
- Laatste commit
- 2026-01-17 (8 maanden geleden)
- Laatste release
- Onbekend
- OpenSSF Scorecard
- Nog niet gemeten
- Zelf hosten
- Niet vastgesteld
- Maintainer-locatie
- Onbekend
- Bron
- github-crawl
Alternatieven in Grote taalmodellen
01 88 02 87 03 84 04 80
D
Dify.ai
Build, test and deploy LLM applications.
L
Langfuse
LLM engineering platform for model tracing, prompt management, and application evaluation. Langfuse helps teams collaboratively debug, analyze, and iterate on their LLM applications such as chatbots or AI agents.
T
textgen
Open-source desktop app for local LLMs. Text, vision, tool-calling, OpenAI/Anthropic-compatible API. 100% private.
F
Firecrawl
The web data API to search, scrape, and interact at scale. 🔥