Open source Self-hostable free
In these stacks
LicenseApache-2.0
PriceFree
UsageNo model cost
Model supportmulti
Languagespython
GitHub stars85.3k
Last activity2026-07-04
Verified2026-06-27
Verified byseed
A high-throughput, memory-efficient inference engine (PagedAttention, continuous batching) for serving open-weight models on your own GPUs.
Strengths
- Industry-standard throughput and batching
- Broad model and hardware support
Tradeoffs
- Requires GPU ops expertise to run well
- Configuration surface is large
Covers
Context management