Open source Self-hostable free
In these stacks
LicenseApache-2.0
PriceFree
UsageNo model cost
Model supportmulti
Languagespython
GitHub stars29.9k
Last activity2026-07-04
Verified2026-06-27
Verified byseed
A fast inference and structured-generation engine with RadixAttention caching, strong for complex prompting and constrained outputs.
Strengths
- Excellent for structured generation and multi-turn
- RadixAttention reuses shared prefixes efficiently
Tradeoffs
- Newer than vLLM, smaller ecosystem
- Requires GPU ops expertise