Results: How Top AI Models Compared in the July 2026 Benchmark Blowout

July 2026 will be remembered as the month the AI industry held 35 simultaneous benchmark shootouts. With Kimi K3 (2.8T, Moonshot AI), Claude Fable 5 (Anthropic), GPT-5.6 Sol (OpenAI), Grok 4.5 (xAI), Inkling (Thinking Machines Lab), and Hy3 (Tencent) all launching within the same window, developers finally got the head-to-head data needed to make informed decisions. Here is how these AI models compared across the benchmarks that matter most.

The Aggregate Picture: AI Models Compared on the Intelligence Index

The Artificial Analysis Intelligence Index, which combines multiple benchmarks into a single composite score, ranks the models as follows: Claude Fable 5 holds the #1 position at 60, GPT-5.6 Sol Ultra is #2 at 59, GPT-5.6 Sol is #3 at 57, Kimi K3 is #4 at 57 (tied with Sol on raw score but behind on the composite), and Claude Opus 4.8 rounds out the top 5 at 56. Hy3 and Inkling, while competitive on individual benchmarks, have not yet been independently ranked on this composite index. When you compare AI models on aggregate intelligence, Kimi K3’s position at #4 — ahead of Opus 4.8 and every GPT-5.5 variant — makes it the highest-ranked open-weight model in history. Use our AI model comparison tool to see how these models align on pricing and features.

Coding Benchmarks: How AI Models Compared for Software Engineering

When comparing AI models for coding, the July 2026 data reveals a fragmented leadership landscape. On SWE Marathon (sustained multi-hour engineering), Kimi K3 leads at 42.0, ahead of GPT-5.6 Sol (39.0) and Fable 5 (35.0). On FrontierSWE (hardest frontier coding problems), Claude Fable 5 retakes the lead at 86.6, with K3 at 81.2 and Sol at 71.3. On TerminalBench 2.1 (agentic system operations), GPT-5.6 Sol leads at 91.9%, followed by Kimi K3 at 88.3% and Fable 5 at 84.6%.

On Program Bench (binary-to-source code reconstruction), K3 edges the field at 77.8, just ahead of Sol at 77.6 and Fable 5 at 76.8. The consistent pattern across all coding benchmarks is that no single model dominates — the right choice depends entirely on whether your task profile matches sustained autonomous sessions (K3), frontier problem-solving (Fable 5), or fast interactive coding (Sol).

Agentic and Knowledge Work Benchmarks

K3 achieved a state-of-the-art 91.2 on BrowseComp — the highest score ever published on this long-horizon web agent benchmark. On OfficeQA Pro (agentic document analysis), K3 scored 87.3, and on SpreadsheetBench 2, it scored 34.8. Claude Fable 5 leads the GDPval-AA v2 composite at 1,815 (Fable 5 Max), with GPT-5.6 Sol Max at 1,747.8 and Kimi K3 at 1,687.

The full benchmark data is available on each model’s detail page on our What’s Hot section, and you can compare AI models side by side using our free AI model comparison tool — add up to 8 models and see all specs, pricing, and features aligned in one table.

Key Takeaways from July 2026

Three conclusions emerge from this month’s data. First, open-weight models have arrived at the frontier — Kimi K3 demonstrates that self-hostable models can compete with $10-$50/MTok proprietary APIs. Second, model selection is now use-case specific — the age of asking “which model is best” is over; the right question is “best for what workload.” Third, the AI models compared across different benchmarks show remarkably tight clustering at the top — differences of 2-5% on most metrics mean that ecosystem, pricing, and integration quality may be more important differentiators than raw benchmark scores for most teams.

Hamza Shehzad

AI industry analyst and researcher at AI Models HQ. Covering the latest developments in artificial intelligence, machine learning, and language models.

Leave a Comment