Kimi K3
The World's Largest Open-Source Model — 2.8 Trillion Parameters
Benchmarks & Evaluation Results
GDPval-AA v2
1,687 (3rd overall, behind only Fable 5 Max & GPT-5.6 Sol Max)
AA-Briefcase
1,527 (2nd place, beats GPT-5.6 Sol Max)
BrowseComp
91.2 / 100 (State-of-the-art, long-horizon information seeking)
Arena.ai Front-End Coding
#1 overall — beats Anthropic Fable 5
Full Analysis of Kimi K3
Moonshot AI, the Beijing-based startup backed by Alibaba and Tencent, has delivered the most significant open-source AI release in history. Kimi K3's 2.8 trillion parameters make it not just the largest open model ever, but a genuine competitor to proprietary frontier systems. The story behind K3 is as remarkable as the technology: Moonshot AI lost significant market position after DeepSeek's rise in early 2025, sliding from third to seventh in monthly active users. The company's strategic pivot
to open-source models — beginning with Kimi K2 in July 2025 — was a bet-the-company move. K3 is the culmination of that bet. On GDPval-AA v2, a benchmark that measures real-world task performance across 44 occupations and 9 major industries, K3 scored 1,687 — placing it ahead of Claude Opus 4.8 (1,600) and behind only Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8). On BrowseComp, a benchmark designed to test long-horizon, high-difficulty information seeking, K3 achieved a state-
of-the-art score of 91.2 out of 100. Perhaps most striking is K3's performance on Arena.ai's front-end coding leaderboard, where it ranks first overall — a result that led Arena CEO Anastasios Angelopoulos to call it potentially 'the single biggest release of the year.' The model features automatic context caching — no cache ID, TTL, or extra parameter required — a meaningful developer experience advantage over competitors that require explicit cache management.
Strengths & Considerations
Strengths
- + Largest open-weight model ever released at 2.8T parameters
- + Frontier-level benchmark scores — third globally, #1 on front-end coding
- + 1M-token context window with automatic caching (no cache ID required)
- + Native vision eliminates need for separate multimodal adapter
- + OpenAI SDK compatible API for zero-friction migration
- + Priced aggressively at $3/$15 vs $10/$50 for comparable proprietary models
- + Open weights under permissive modified MIT license
Considerations
- − Weights not yet independently verifiable until July 27 release
- − 2.8T sparse model requires substantial hardware for self-hosting
- − Max thinking effort locked as default at launch; low/high-effort modes pending
- − Benchmark claims await independent reproduction outside Moonshot's test suite
- − Smaller Western developer ecosystem compared to Llama or Qwen
Best Use Cases for Kimi K3
Where this model excels and the types of workloads it is best suited for.
Use Case 1
Long-horizon autonomous coding and software development
Use Case 2
Large codebase analysis and multi-file refactoring
Use Case 3
Vision-in-the-loop development (UI design, game dev, CAD)
Use Case 4
Enterprise document analysis at 1M-token context
Use Case 5
Cross-lingual knowledge work (native Chinese-English bilingual)
Use Case 6
Cost-sensitive production deployments needing frontier intelligence
How to Use Kimi K3
Quick Start Guide
Visit kimi.com and sign up with a Google account or phone number — no credit card required. For API access, use the OpenAI SDK compatible endpoint at api.moonshot.ai/v1 with model ID 'kimi-k3'. Pricing is straightforward: $3/MTok input (uncached), $0.30/MTok cached, $15/MTok output. A promotional top-up rebate running through August 12 offers up to 30% back in vouchers for API credits of $1,000 or more. For self-hosting, the full weights in BF16 and NVFP4 formats will be available on Hugging Face from July 27.
Other Models from July 2026
Compare Kimi K3 with other recent releases.