AI Programming · 2026-10-04 · 0:57

Scale's Cleaner Coding Leaderboard

Two synthetic hosts, Wren and Ash, on the tracker's current state. The script was checked claim by claim against the tracker page and the cited sources below; every claim carries a verbatim quote.

Download MP4
Transcript

Wren Coding agent scores are inflated, and Scale just changed the rulebook.

Ash Yeah, SWE-Bench Pro V2. It drops 89 invalid tasks, locks the network so agents can't look up fixes, and re-grades every answer on a clean copy.

Wren Compared to Verified, though? That one's already contaminated, so models probably saw those tests in training.

Ash Claude Opus 4.5 fell from 80.9% on Verified to 45.89% on the original Pro. V2 tries to keep that honesty.

Wren And Bito's 60.8% on Pro? That's just the vendor talking, unverified, and from the old version.

Ash So check which version a number comes from. Scale says V2's numbers can now be trusted.

Wren Which means weigh the private set most, since Scale calls it the only clean measurement of memorisation.

Ash Full tracker's at DreamLab Research.

Sources
  1. The tracker page
  2. SWE-Bench Pro V2
  3. SWE-Bench Pro V2: A Cleaner, Harder-to-Game Leaderboard | Scale Labs

More shorts

0:30 TTS ModelsBreeze TTS 2 Leads Downloadable Models 0:42 HBOT & Red LightCustoms Can Hold Your Chamber 0:48 Hackable IP CamerasSupported List, Unsupported Camera 0:47 Software FactoriesTwo Agents Do Worse Than One 0:47 ESP32 & ESPHome EcosystemESPHome 2026.9.0 Released 0:45 Consumer Wearables You Can Still Own The Data FromGarmin's API Door Is Closed