Together AI is an AI-native cloud platform specializing in fast, cost-efficient inference over open-source models including Llama, Mistral, Qwen, DeepSeek, and dozens of others. It offers serverless pay-per-token pricing, dedicated GPU endpoints, fine-tuning, and batch processing. New users receive $25 in free credits. In 2026, the catalog includes recent flagship models such as DeepSeek V3.1, Kimi K2.6, and Qwen 3.6 Plus. No subscription tiers; purely consumption-based. Key features: - Serverless inference over 200+ open-source models (Llama, Mistral, Qwen, DeepSeek) - Dedicated GPU endpoints (H100 from $6.49/hr) for latency-sensitive workloads - Fine-tuning for Llama, Mistral, and Qwen models including 405B scale - Batch processing API at 50% discount for non-real-time jobs - Rentable GPU clusters for training workloads
Pay-per-token: $0.05-$9.00 per million tokens depending on model. No subscription required. $25 free credits for new users. Dedicated H100 endpoints from $6.49/hr.
