Groq

Ultra-low-latency inference on custom LPU silicon.

Visit site
Proprietary usage
LLM provider
In these stacks
Write code with AIBuild my own agent
LicenseProprietary
PriceFree tier
UsageMetered
Model supportmulti
Languages
GitHub stars
Last activity
Verified2026-06-27
Verified byseed

An inference provider running open models on custom LPU hardware for exceptionally low token latency, with a generous free tier.

Strengths

  • Class-leading inference latency
  • Generous free tier for open models

Tradeoffs

  • Limited to the models Groq has ported
  • Throughput caps on free tier