Groq provides cloud inference via its custom Language Processing Unit (LPU) chips, running open-source models (Llama, Mixtral, Gemma, DeepSeek R1, Qwen, Whisper) at 500+ tokens per second — typically 10-20x faster than GPU-based alternatives at significantly lower cost. Groq does not develop its own models; it operates as an inference provider. In December 2025, Nvidia announced a roughly $20 billion non-exclusive licensing deal for Groq's LPU architecture. Groq offers a free tier and pay-per-token API pricing with Batch API discounts. Key features: - 500+ tokens/second inference on LPU custom silicon (vs. GPU alternatives) - Support for Llama, Mixtral, Gemma, DeepSeek R1, Qwen, and Whisper models - Batch API with 50% cost reduction for async workloads - Automatic prompt caching halving repeated-prefix costs - Free tier available via GroqCloud - Linear, predictable per-token pricing with no idle infrastructure costs
Pay-per-token. Llama 3.1 8B: $0.05/$0.08 per 1M input/output tokens. Llama 3.3 70B: $0.59/$0.79 per 1M tokens. Free tier available.
