Inkling: Thinking Machines’ 975B MoE Model Review

Thinking Machines’ Inkling, released July 15, 2026, represents a watershed moment for open-source AI. At 975B total parameters with 41B active per token, it uses Mixture of Experts architecture for efficient inference. The Apache 2.0 license makes it the most permissive model at this scale.

Why Inkling Matters

Inkling proves open-source can compete at the frontier. At $1.87/$4.68 per million tokens, it undercuts proprietary alternatives by 50-80%. More importantly, Apache 2.0 means any organization can use, modify, and redistribute it without restrictions.

Performance

On key benchmarks, Inkling performs competitively with GPT-5.6 Terra and Claude Opus 4.8. Its MoE architecture delivers fast inference despite the massive total parameter count. The 41B active parameters per token mean efficient operation.

Who Should Use It

Organizations prioritizing AI sovereignty and transparency. Teams that want to inspect, fine-tune, and deploy without vendor lock-in. Developers building open-source AI projects that need a strong foundation.

Related: AI Model Comparison 2026

AI Models HQ Team

Independent AI model comparison experts benchmarking every major language model: OpenAI, Anthropic, Google, xAI, Meta and more. Real pricing, real benchmarks, zero hype.

Leave a Comment