A list of inference providers with good value for money token pricing and some inference provider which provide cheap models and cool models like dLLMs

NamePricing / TierDetails
CerebrasDeveloper Tier $10 min deposit, Pay-As-You-Go (~$0.35 in / $0.75 out per 1M tokens on GPT-OSS-120B). Standard card billing.Uses custom Wafer-Scale Engine hardware delivering ultra-high throughput (~1,800 to 3,000 tokens/sec) on hosted models like GPT-OSS and Gemma. OpenAI-compatible API.
GroqFree Tier Free access with rate limits (up to 30 RPM, 6k–30k TPM depending on model). Pay-As-You-Go Standard Stripe token rates without extra surcharge.Runs on proprietary LPU (Language Processing Unit) silicon optimized for low latency and high generation speeds (500–800 tokens/sec) across Llama, Qwen, and DeepSeek variants.
OpenCode GoSubscription $10/month ($5 first month). Usage allowance caps at $12/5h, $30/week, and $60/month nominal usage value.Managed subscription inside OpenCode providing access to 18 curated open-source models (DeepSeek, Qwen-Coder, Kimi, GLM, MiniMax) without requiring individual third-party API keys.
OpenCode ZenPay-As-You-Go $20 starting balance. Model tokens passed at zero markup + credit card processing fee (4.4% + $0.30 per transaction).Managed PAYG gateway within OpenCode offering access to frontier proprietary models (Claude 3.5/3.7 Sonnet, GPT-4o, Gemini 1.5/2.0 Pro) and free community endpoints.
OpenRouterPrepaid Pay-As-You-Go Model token pass-through pricing + 5.5% transaction fee ($0.80 minimum fee per deposit). Min deposit $5.Universal API gateway aggregator providing access to hundreds of both proprietary (Anthropic Claude, OpenAI GPT, Google Gemini) and open-weights models through a single API key.
Together AIPay-As-You-Go Standard Stripe billing (no transaction surcharge). Includes $5 free starting credit.Large decentralized AI cloud hosting 200+ open-source models, multimodal engines, and custom fine-tuning endpoints via a single unified API.
Fireworks AIPay-As-You-Go Standard Stripe billing (no transaction surcharge). Includes $1 free starting credit.Low-latency inference platform utilizing custom compound AI acceleration (FireAttention) for fast Time-to-First-Token and optimized structured output/JSON streaming.
DeepInfraPay-As-You-Go Standard Stripe billing (no transaction surcharge). Offers free starter credits.Serverless hosting platform featuring 100+ open-weight models (DeepSeek, Llama, Qwen, Mistral). Highly competitive wholesale per-token pricing with full OpenAI API compatibility.
Runware AIPay-As-You-Go Granular compute/token pricing with standard card billing (no deposit surcharge). $2 free starter credit.Unified, high-throughput inference engine specialized in ultra-low-cost generative media (Flux, SD, Video, Audio, 3D) alongside hosted open-source LLMs (DeepSeek, GLM, Kimi, Llama).
Mistral AI (La Plateforme)Free Tier Free experimentation tier with rate limits. Pay-As-You-Go Token consumption billed directly without deposit markups.First-party API provider for the Mistral model family, including Codestral (specialized for code completion and fill-in-the-middle), Mistral Large, and Pixtral.
SambaNova CloudDeveloper Tier Pay-As-You-Go at competitive per-token rates with standard card billing.Runs on custom Reconfigurable Dataflow Unit (RDU) chips. Known for fast inference on large models like Llama 3.1 405B/70B and DeepSeek. OpenAI-compatible endpoint.
Nebius AI StudioPay-As-You-Go Standard credit card billing with no platform surcharge. Frequent free starting credits for new accounts.GPU infrastructure provider running dedicated Nvidia H100/H200 clusters. Offers near-cost token pricing for DeepSeek, Qwen 2.5 Coder, and Llama 3.3. OpenAI-compatible API.
Cloudflare Workers AIFree Tier 10,000 free neurons/day (~100k+ tokens daily). Paid Plan $5/month Workers Paid plan for higher limits + Pay-As-You-Go compute.Serverless inference executed directly on Cloudflare global edge data centers. Hosts Llama, Qwen, DeepSeek, and Mistral with very low network routing latency.
Novita AIPay-As-You-Go Standard balance top-ups without transaction penalties.Serverless inference provider offering aggressive wholesale pricing for popular open-source LLMs (DeepSeek, Llama, Qwen) along with image and multimodal pipelines.
HyperbolicPay-As-You-Go Standard balance deposits via card or crypto without hidden penalties. Low per-token rates.Distributed open-access AI cloud hosting open-weights models (DeepSeek, Llama, Qwen) with a focus on high availability and low compute costs.
Inception LabsMercury 2 & Mercury Edit 2 $0.25/1M input, $0.75/1M output. Early Access (10x free tokens promo).Diffusion-based LLMs (dLLMs) generating tokens in parallel vs. sequential auto-regressive. Sub-300ms TTFT, 5–7x higher throughput, up to 70% lower cost. OpenAI API-compatible.