LMDeploy

Efficient inference and serving for open models.

Visit site GitHub
Open source Self-hostable free
Inference engine
In these stacks
Write code with AIBuild my own agent
LicenseApache-2.0
PriceFree
UsageNo model cost
Model supportmulti
Languagespython, cpp
GitHub stars7.9k
Last activity2026-07-03
Verified2026-06-27
Verified byseed

An inference and serving toolkit (from the InternLM team) offering compressed, high-throughput deployment of open-weight models.

Strengths

  • Strong quantization and throughput optimizations
  • Good support for InternLM and common open models

Tradeoffs

  • Smaller community than vLLM
  • Less Western documentation