Software Factories · 2026-10-04 · 0:47

Two Agents Do Worse Than One

Two synthetic hosts, Wren and Ash, on the tracker's current state. The script was checked claim by claim against the tracker page and the cited sources below; every claim carries a verbatim quote.

Download MP4
Transcript

Wren Two AI agents together do worse than one working alone.

Ash CooperBench now has its own website. GPT-5 and Claude Sonnet 4.5 hit only 25% success as a pair, roughly 50% below one agent doing both tasks.

Wren Wait, we'd logged that 50% already, but secondhand, via Zylos Research on Sep 14. Basically a press summary.

Ash Right, and our planner verdict rested on a boss agent handing out tasks, plus the Hugging Face swarm postmortem. No direct team benchmark.

Wren So "more cheap workers, more output" now meets a primary source. Treat parallel delegation as a cost to justify, and keep the verification gate first.

Ash Full tracker's at DreamLab Research.

Sources
  1. The tracker page
  2. CooperBench: Benchmarking Agent Teams | Why Coding Agents Cannot be Your Teammates Yet
  3. Leaderboard · CooperBench
  4. The Curious Case of Miscoordination

More shorts

0:30 TTS ModelsBreeze TTS 2 Leads Downloadable Models 0:42 HBOT & Red LightCustoms Can Hold Your Chamber 0:48 Hackable IP CamerasSupported List, Unsupported Camera 0:47 ESP32 & ESPHome EcosystemESPHome 2026.9.0 Released 0:45 Consumer Wearables You Can Still Own The Data FromGarmin's API Door Is Closed 0:45 Open & Hackable AI PendantsMeta Buys, Halts, Then Announces