A list of inference providers with good value for money token pricing and some inference provider which provide cheap models and cool models like dLLMs
| Name | Pricing / Tier | Details |
|---|---|---|
| Cerebras | Developer Tier $10 min deposit, Pay-As-You-Go (~$0.35 in / $0.75 out per 1M tokens on GPT-OSS-120B). Standard card billing. | Uses custom Wafer-Scale Engine hardware delivering ultra-high throughput (~1,800 to 3,000 tokens/sec) on hosted models like GPT-OSS and Gemma. OpenAI-compatible API. |
| Groq | Free Tier Free access with rate limits (up to 30 RPM, 6k–30k TPM depending on model). Pay-As-You-Go Standard Stripe token rates without extra surcharge. | Runs on proprietary LPU (Language Processing Unit) silicon optimized for low latency and high generation speeds (500–800 tokens/sec) across Llama, Qwen, and DeepSeek variants. |
| OpenCode Go | Subscription $10/month ($5 first month). Usage allowance caps at $12/5h, $30/week, and $60/month nominal usage value. | Managed subscription inside OpenCode providing access to 18 curated open-source models (DeepSeek, Qwen-Coder, Kimi, GLM, MiniMax) without requiring individual third-party API keys. |
| OpenCode Zen | Pay-As-You-Go $20 starting balance. Model tokens passed at zero markup + credit card processing fee (4.4% + $0.30 per transaction). | Managed PAYG gateway within OpenCode offering access to frontier proprietary models (Claude 3.5/3.7 Sonnet, GPT-4o, Gemini 1.5/2.0 Pro) and free community endpoints. |
| OpenRouter | Prepaid Pay-As-You-Go Model token pass-through pricing + 5.5% transaction fee ($0.80 minimum fee per deposit). Min deposit $5. | Universal API gateway aggregator providing access to hundreds of both proprietary (Anthropic Claude, OpenAI GPT, Google Gemini) and open-weights models through a single API key. |
| Together AI | Pay-As-You-Go Standard Stripe billing (no transaction surcharge). Includes $5 free starting credit. | Large decentralized AI cloud hosting 200+ open-source models, multimodal engines, and custom fine-tuning endpoints via a single unified API. |
| Fireworks AI | Pay-As-You-Go Standard Stripe billing (no transaction surcharge). Includes $1 free starting credit. | Low-latency inference platform utilizing custom compound AI acceleration (FireAttention) for fast Time-to-First-Token and optimized structured output/JSON streaming. |
| DeepInfra | Pay-As-You-Go Standard Stripe billing (no transaction surcharge). Offers free starter credits. | Serverless hosting platform featuring 100+ open-weight models (DeepSeek, Llama, Qwen, Mistral). Highly competitive wholesale per-token pricing with full OpenAI API compatibility. |
| Runware AI | Pay-As-You-Go Granular compute/token pricing with standard card billing (no deposit surcharge). $2 free starter credit. | Unified, high-throughput inference engine specialized in ultra-low-cost generative media (Flux, SD, Video, Audio, 3D) alongside hosted open-source LLMs (DeepSeek, GLM, Kimi, Llama). |
| Mistral AI (La Plateforme) | Free Tier Free experimentation tier with rate limits. Pay-As-You-Go Token consumption billed directly without deposit markups. | First-party API provider for the Mistral model family, including Codestral (specialized for code completion and fill-in-the-middle), Mistral Large, and Pixtral. |
| SambaNova Cloud | Developer Tier Pay-As-You-Go at competitive per-token rates with standard card billing. | Runs on custom Reconfigurable Dataflow Unit (RDU) chips. Known for fast inference on large models like Llama 3.1 405B/70B and DeepSeek. OpenAI-compatible endpoint. |
| Nebius AI Studio | Pay-As-You-Go Standard credit card billing with no platform surcharge. Frequent free starting credits for new accounts. | GPU infrastructure provider running dedicated Nvidia H100/H200 clusters. Offers near-cost token pricing for DeepSeek, Qwen 2.5 Coder, and Llama 3.3. OpenAI-compatible API. |
| Cloudflare Workers AI | Free Tier 10,000 free neurons/day (~100k+ tokens daily). Paid Plan $5/month Workers Paid plan for higher limits + Pay-As-You-Go compute. | Serverless inference executed directly on Cloudflare global edge data centers. Hosts Llama, Qwen, DeepSeek, and Mistral with very low network routing latency. |
| Novita AI | Pay-As-You-Go Standard balance top-ups without transaction penalties. | Serverless inference provider offering aggressive wholesale pricing for popular open-source LLMs (DeepSeek, Llama, Qwen) along with image and multimodal pipelines. |
| Hyperbolic | Pay-As-You-Go Standard balance deposits via card or crypto without hidden penalties. Low per-token rates. | Distributed open-access AI cloud hosting open-weights models (DeepSeek, Llama, Qwen) with a focus on high availability and low compute costs. |
| Inception Labs | Mercury 2 & Mercury Edit 2 $0.25/1M input, $0.75/1M output. Early Access (10x free tokens promo). | Diffusion-based LLMs (dLLMs) generating tokens in parallel vs. sequential auto-regressive. Sub-300ms TTFT, 5–7x higher throughput, up to 70% lower cost. OpenAI API-compatible. |