Two synthetic hosts, Wren and Ash, on the tracker's current state. The script was checked claim by claim against the tracker page and the cited sources below; every claim carries a verbatim quote.
Transcript
Wren Two AI agents together do worse than one working alone.
Ash CooperBench now has its own website. GPT-5 and Claude Sonnet 4.5 hit only 25% success as a pair, roughly 50% below one agent doing both tasks.
Wren Wait, we'd logged that 50% already, but secondhand, via Zylos Research on Sep 14. Basically a press summary.
Ash Right, and our planner verdict rested on a boss agent handing out tasks, plus the Hugging Face swarm postmortem. No direct team benchmark.
Wren So "more cheap workers, more output" now meets a primary source. Treat parallel delegation as a cost to justify, and keep the verification gate first.
Ash Full tracker's at DreamLab Research.