vLLM

High-throughput self-hosted LLM serving engine.

Visit site GitHub
Open source Self-hostable free
Inference engine
In these stacks
Write code with AIBuild my own agent
LicenseApache-2.0
PriceFree
UsageNo model cost
Model supportmulti
Languagespython
GitHub stars85.3k
Last activity2026-07-04
Verified2026-06-27
Verified byseed

A high-throughput, memory-efficient inference engine (PagedAttention, continuous batching) for serving open-weight models on your own GPUs.

Strengths

  • Industry-standard throughput and batching
  • Broad model and hardware support

Tradeoffs

  • Requires GPU ops expertise to run well
  • Configuration surface is large