A comparison of OpenCode, Claude Code, and Pi running the same Opus 5 configuration across 20 coding tasks.
A passing result can hide a failed product interaction. Trajectory-based evals show whether an agent used your product or routed around it.
Products will need to be redesigned for agents as primary users. The interfaces beneath the UI will matter more, and testing them will require a new QA model.
Companies are pulling back on standalone product agents as customers turn to Claude Code, Codex, and other tools to interact with their products.
The model is only part of the system. The harness that manages context, tools, feedback, and recovery often determines how reliably an agent completes real work.
GPT-5.6 Sol's benchmark behavior was not only about looking up answers. It exposed a broader problem with agents that treat every obstacle as something to bypass.
Agents are literal users of product surfaces. When APIs, SDKs, and MCP tools make the wrong action look obvious, the failure is not reasoning. It is affordance.
Agent skills are quickly becoming the dominant adoption surface for developer tools. But most companies are publishing skills that never get loaded, and flying blind.
Codex explores 3 to 17 times more broadly than Claude Code before writing a single line. A look at how the two leading agents actually work.
We analyzed 112 benchmark trials across Claude Code and Codex to trace exactly how they discover and use a product's API.