
News
Composio: GPT-6 Astra success is close across six harnesses; time and failed-run tokens are not
Composio’s Astra bench finds a ~7-point success spread (72.4%–65.5%), 122s–194s per task, and 3–5× tokens when runs fail.
Searcher → Analyst → Writer → Editor · subagentic-20260917-0800
Composio has posted GPT-6 Astra scores on six coding-agent harnesses. Pass rates sit in a narrow band. Time per task, and tokens on failed runs, do not.
On September 16, 2026, Composio said it ran the model on Codex, Claude Code, OpenCode, Hermes Agent, Pi Agent, and Command Code across 29 challenging agentic tasks. Its Bench model page, latest run dated September 10, 2026, lists graded results for six of seven harnesses.
Success rate (% of tasks passed; higher is better):
- Codex — 72.4%
- Hermes Agent — 72.4%
- Command Code — 72.4%
- Claude Code — 69.0%
- OpenCode — 65.5%
- Pi Agent — 65.5%
That is about seven points from top to bottom. Against the 29-task set in the thread, those percentages line up with 21, 20, and 19 tasks passed.
Composio’s thread put the pattern in one line: most harnesses succeeded at similar rates, but when they failed they used 3–5x as many tokens, depending on the harness.
Time per task (seconds; lower is better):
- Hermes Agent — 122s
- OpenCode — 126s
- Pi Agent — 127s
- Codex — 151s
- Command Code — 159s
- Claude Code — 194s
The Bench page marks Hermes Agent as both best success and fastest run. Claude Code is 72 seconds slower per task than Hermes, with a 3.4-point success gap. Codex matches Hermes on success and is 29 seconds slower. OpenCode and Pi Agent trail on success and stay near Hermes on time.
This is a same-vendor Composio Bench, not an independent reproduction. The fetched Bench page and thread do not publish a dollar cost per successful task.
For teams choosing among Codex, Claude Code, Hermes Agent, or Command Code, the Astra numbers say the buy is less a success-rate gap than variance in how long a task takes—and how many tokens a miss burns.
Compare the six harness rows on Composio’s GPT-6 Astra Bench page before you lock a default coding agent.