Groq

The fastest LLM inference on the planet (LPU hardware)

groq.com · Free tier; per-token, very cheap

Groq serves open models (Llama, Qwen) at hundreds of tokens/sec on custom LPU chips — for AI features where response speed IS the feature.

Capabilities

Best for

  • Realtime AI UX: instant autocomplete, live transcription post-processing, voice agents
  • Cheap fast open-model inference with an OpenAI-compatible API

Not for

  • Frontier-quality reasoning (serve Claude/GPT for the hard thinking)