AgentLens launched in July 2026 as the first serious benchmark for software engineering agents — testing Claude Code, Codex CLI, Gemini CLI, and Aider on real GitHub issues and pull requests. The results mark the beginning of the agent-evaluation era.
Results
Claude Code 5.8 leads with 55.4%, followed by Codex CLI (37.6%), Gemini CLI (33.8%), and Aider (18.6%). The field is defined by real-world repo solving — not synthetic questions.
Why Agent Benchmarks Matter
2026’s agentic shift means the question is no longer “which model is smartest” but “which agent ships code.” Benchmarks like AgentLens measure the full workflow: understanding issues, navigating repositories, running tests, and landing pull requests.
Limitations
Results depend heavily on test setup, repository selection, and difficulty calibration — and they age quickly as tools update. Use AgentLens as a directional signal, not gospel.
Takeaway
Claude Code’s lead is real and durable — it powered Anthropic’s revenue overtake. But the spread between Claude Code 5.8 and Codex CLI is narrowing as OpenAI invests heavily in agentic tooling.
Related: AI Model Comparison 2026 · AI Models in 2026: Complete Guide


