August 2026’s price leaders are dramatically cheaper than six months ago. OpenAI’s Luna cut to $0.20/$1.20, DeepSeek holds at $0.14/$0.28, and Alibaba’s Qwen 3.7 Flash now lists at $0.03/$0.13 — the cheapest paid API of any mainstream provider. Here is the updated ranking.
The August 2026 Cheap Tier
| Model | Input $/M | Output $/M | Notes |
|---|---|---|---|
| Qwen 3.7 Flash | $0.03 | $0.13 | Cheapest listed paid API; vision capable |
| Muse Spark 1.2 Contributor | $0.10 | $0.20 | Meta’s contributor tier |
| DeepSeek V4 Flash | $0.14 | $0.28 | 1M context, MIT weights |
| GPT-5.6 Luna | $0.20 | $1.20 | Cheapest flagship-tier |
| DeepSeek V4 Pro | $0.435 | $0.87 | Best open-weight reasoning value |
Reading the Table Correctly
List price is not effective cost. Output-heavy workloads should sort by output price; retry-prone workloads by quality per dollar. A model that produces more tokens or needs more retries can cost more per completed task than its token rate suggests.
Quality Reality Check
Only DeepSeek V4 Flash and GPT-5.6 Luna are production-grade for coding and agentic work in the cheap tier. Qwen 3.7 Flash and Muse Spark Contributor are fine for high-volume simple tasks where budget dominates. For frontier quality you still pay frontier prices (Opus 5: $5/$25, Sol: $5/$30).
Free Tiers Still Exist
DeepSeek continues to offer a 5M-token free allowance; Google keeps generous Free quotas on Flash models. Our /best-free-ai-apis/ guide covers the full free tier landscape.
Bottom Line
The floor for AI API pricing keeps dropping — verify current rates before committing architecture, and re-check monthly. Our AI Model Comparison 2026 includes the methodology.


