Text Generation Inference

Hugging Face's production LLM server.

Visit site GitHub
Open source Self-hostable free
Inference engine
In these stacks
Write code with AIBuild my own agent
LicenseApache-2.0
PriceFree
UsageNo model cost
Model supportmulti
Languagespython, rust
GitHub stars10.9k
Last activity2026-03-21
Verified2026-06-27
Verified byseed

Hugging Face's production-grade server for deploying large language models with optimized kernels and a tuned inference stack.

Strengths

  • Backed by Hugging Face's model ecosystem
  • Optimized, production-oriented serving

Tradeoffs

  • Heavier deployment than llama.cpp
  • Focus has shifted alongside the HF stack