Understanding AI API pricing is essential for controlling costs. This guide covers current 2026 pricing across every major provider, plus the discounts and strategies that can cut your bill by half or more.
Current API Pricing Table
| Provider / Model | Input | Output | Context |
|---|---|---|---|
| Google Gemini 3.5 Flash | $0.15/M | $0.60/M | 1M |
| Google Gemini 3.1 Pro | $2/M | $12/M | 2M |
| OpenAI GPT-5.6 Luna | $1/M | $6/M | 1M |
| OpenAI GPT-5.6 Terra | $2.5/M | $15/M | 1M |
| OpenAI GPT-5.6 Sol | $5/M | $30/M | 1M |
| Anthropic Claude Opus 5 | $5/M | $25/M | 1M |
| Anthropic Claude Fable 5 | $10/M | $50/M | 1M |
| Anthropic Claude Sonnet 5 | $3/M | $15/M | 200K |
| DeepSeek V3.2 | $0.5/M | $2/M | 128K |
| Thinking Machines Inkling | $1.87/M | $4.68/M | 1M |
| Moonshot Kimi K3 | $3/M | $15/M | 1M |
| Alibaba Qwen 3 | $0.6/M | $2/M | 131K |
| Mistral Large 3 | $2/M | $6/M | 128K |
| xAI Grok 4.5 | — | $6/M | 1M |
Discounts You Should Use
Cached Input Discounts
Most major providers offer 50-90% discounts on cached input tokens. If your prompts share static prefixes (system instructions, document chunks), caching can slash your biggest cost.
Batch APIs
Non-real-time workloads via batch endpoints typically get 50% off. Use for indexing, classification, and nightly processing.
Free Tiers
Google AI Studio (~1,500 req/day), OpenRouter (25+ free models), and DeepSeek (5M free tokens) can cover prototyping and light production. See our Best Free AI APIs guide.
Cost Optimization Strategies
- Model routing: budget models for simple tasks, frontier for hard ones
- Context trimming: every context token is billed — keep prompts lean
- Measure cost-per-task, not cost-per-token
- Negotiate volume discounts at scale (typical: 10-30%)
Related: AI Model Comparison 2026 · Best Free AI Models