Trustworthy comparisons require transparent methodology. This page explains exactly how AI Models HQ evaluates and compares AI models — the data sources, benchmark weights, and limitations of our analysis.
Data Sources
- Official provider documentation and pricing pages
- Vendor-published benchmark results (treated as ceiling, marked as vendor-reported)
- Independent leaderboards: Artificial Analysis, BenchLM, LMArena
- Standardized evaluation harnesses (Scale Morph, standardized SWE-bench)
- Public API measurements for latency and reliability
- Community arenas and blind preference votes
Pricing Data
All prices are listed per 1M tokens for input and output, as published by providers. Prices change frequently — we mark verification dates on every comparison. Where providers offer free tiers, caching discounts, or batch pricing, we note them separately.
Benchmark Weights
When we summarize overall capability, we weight benchmarks by relevance: SWE-bench for coding (heaviest for developer use cases), GPQA Diamond for reasoning, MMLU for knowledge, and Arena ELO for conversational quality. No single number captures a model — we publish dimension-by-dimension scores.
The Harness Problem
Vendor-reported scores run 17-21 points above standardized harness scores for the same models. We treat vendor numbers as ceilings and standardized leaderboards as floors, and we say which is which. If a figure cannot be verified, we say so rather than invent it.
Editorial Independence
We accept no paid placements. No provider can purchase a higher ranking. Our comparisons are funded by site operations, and we disclose any affiliate relationships. The ranking is our judgment based on the evidence above — not a leaderboard scrape.
Limitations
- Benchmarks are proxies; your tasks may differ
- Model behavior changes silently with provider updates
- Regional availability varies
- Data verified as of the date shown on each page
Corrections
Found an error? Contact us and we will verify and correct within 48 hours. Corrections are logged publicly.
Related: AI Model Comparison 2026 · Best Free AI Models