Last updated 2026-10-03 · updated daily by an automated research job · run #7.
The verdicts
This topic judges "best" along 4 axes, each with its own standing verdict and history.
Autonomous
Judged on: Best evidenced agent for multi-step work handed off with a goal rather than a diff — plans, edits many files, runs its own verification. Judged on independent task benchmarks and reported real-world completion, not vendor demos.
The current state of the art for Autonomous, as of 2026-08-18 (confidence primary-source):
Based on the ProjDevBench benchmark evaluating six coding agents, Codex+GPT-5 achieves the highest overall performance (77.85%) on end-to-end project tasks, though performance gaps widen on from-scratch construction tasks. However, independent empirical evidence from a METR randomized trial suggests AI assistance may actually slow down experienced developers in some contexts.
What triggered or contributed to this call:
- 2026-08-18 — ProjDevBench benchmark evaluates six coding agents on end-to-end project tasks. (
study,primary-source) - 2026-08-18 — METR randomized trial finds AI assistance made experienced open-source developers 19% slower. (
study,primary-source)
Contenders:
- Codex — Top performer in the ProjDevBench benchmark when combined with GPT-5.
- Claude Code — Shows strong performance in multi-file tasks and refactoring according to comparisons and receives frequent updates.
Pairing
Judged on: Best evidenced tool for real-time human-and-agent work in one session — interruptibility, shared context, how the human steers mid-run.
The current state of the art for Pairing, as of 2026-08-18 (confidence primary-source):
The open-source tool 'The Pair' introduces a multi-agent architecture with Mentor+Executor cross-checking to catch hallucinations, providing a concrete implementation for AI pair programming. VS Code's integration of Claude and Codex agents offers a widely accessible platform for real-time collaborative work.
What triggered or contributed to this call:
- 2026-08-18 — Open-source tool 'The Pair' provides multi-agent coding with cross-checking. (
discovery,primary-source) - 2026-08-18 — VS Code adds support for Claude and Codex agents under GitHub Copilot subscription. (
update,primary-source)
Contenders:
- The Pair — Open-source, cross-model support with built-in review agent for safety.
- Visual Studio Code — Broad IDE integration with popular agents under a Copilot subscription.
Orchestration
Judged on: Best evidenced way to run several agents on one codebase at once — coordination, isolation, conflict and drift handling.
The current state of the art for Orchestration, as of 2026-08-24 (confidence vendor-claim):
The launch of Moderne's 'Moddy' agent provides a vendor-claimed, concrete implementation for multi-repo, large-scale codebase orchestration, moving beyond single-repo coordination patterns. The previously noted guide from AugmentCode remains the primary independent documentation of failure modes for multi-agent work on a single repository.
What triggered or contributed to this call:
- 2026-08-24 — Moderne launches 'Moddy', a multi-repo AI agent for transforming enterprise codebases at scale. (
launch,vendor-claim)
Contenders:
- The Pair — Open-source implementation with a multi-agent, cross-checking architecture for single-repo work.
- Moderne — Claims to orchestrate transformations across multiple repositories at scale for enterprise codebases.
Dethronement history — Orchestration
Every verdict this facet has ever retired, and the receipts for retiring it.
Dethroned 2026-08-24 — held since 2026-08-18
A guide from AugmentCode documents specific failure modes when running multiple AI coding agents on one repository without coordination, highlighting the need for deliberate orchestration. The open-source tool 'The Pair' demonstrates one architectural approach using multiple agents with defined roles.
How it was beaten, and how we know the new one is better: The previous verdict highlighted a guide documenting failure modes and The Pair's single-repo architecture. The new item, Moderne's Moddy, is claimed to operate at a multi-repo, enterprise-scale level, representing a different and more complex orchestration target, though its effectiveness is not yet independently verified.
Triggering entries:
- 2026-08-18 — Guide documents failure modes of running multiple AI coding agents on one repository. (
discovery,secondary) - 2026-08-18 — Open-source tool 'The Pair' provides multi-agent coding with cross-checking. (
discovery,primary-source)
Local
Judged on: Best coding setup that runs entirely on hardware you own, judged on usable quality at a stated VRAM budget, not on leaderboard position.
The current state of the art for Local, as of 2026-09-07 (confidence secondary):
Effective autonomous coding agents require 20B+ parameter models, with 32B+ offering significantly better performance, creating a fundamental tension with VRAM constraints on local hardware. Guides discuss hardware requirements and model options, but no single leading local setup or 'best' model is identified from current evidence that balances high coding capability with typical consumer VRAM budgets.
What triggered or contributed to this call:
- 2026-09-07 — Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension. (
discovery,secondary)
Dethronement history — Local
Every verdict this facet has ever retired, and the receipts for retiring it.
Dethroned 2026-09-07 — held since 2026-08-18
Several guides discuss hardware requirements and model options for self-hosted AI coding, noting that effective autonomous agents typically require 20B+ parameter models, creating a tension with VRAM constraints. No single leading local setup is identified from the current results.
How it was beaten, and how we know the new one is better: The previous verdict stated that 'No single leading local setup is identified from the current results.' New evidence from a Medium article based on practical lessons provides a more concrete technical threshold, specifying that true autonomous agents require 20B parameter models minimum, which clarifies the VRAM-performance tension but does not yet identify a leading setup that resolves it.
What this is, and how it works
This page is generated, not written. A scheduled job runs daily on a machine in a homelab. Each run it:
- Searches the open web for both the products already tracked here and for category-level terms designed to turn up ones we've never heard of.
- Feeds those results to a language model along with everything already on this page, and asks it what is genuinely new. Finding nothing is an acceptable answer, and most runs should find little.
- Writes the result into a versioned JSON dataset (
data/live-research/ai-programming.json) and commits it. This page is re-rendered from that file.
So: a machine wrote the prose here. Every changelog entry carries at least one source link, and every item and entry carries a confidence label:
primary-source— vendor documentation, a published paper, or a regulator.secondary— reputable press.vendor-claim— marketing or an unreplicated vendor statement.unverified— a single low-quality source, usually auto-discovered.
Any label weaker than primary-source is printed beside the finding, so an unflagged row is a primary source. A tracker that hides its own uncertainty is worse than no tracker.
What this is not. Not medical advice. Not a review site — nothing here has been tested in our hands. No affiliate relationships, no sponsored placements, nothing bought. Vendor claims are attributed to the vendor rather than restated as findings. Where a price, a date, or a regulatory status is unknown, it is left blank instead of guessed.
What's new
44 developments in the last 30 days, newest first.
| Date | Subject | Finding |
|---|---|---|
| Oct 3 | Codex GitHub release 0.162.0-alpha.9 |
OpenAI releases Codex 0.162.0-alpha.9. |
| Oct 3 | SWE-bench, Scale AI Scale Labs blog post |
Scale AI releases SWE-Bench Pro V2 with a cleaner, harder-to-game leaderboard. |
| Oct 3 | discovery VentureBeat article |
Autoheal launches, claiming to manage the work AI coding agents leave behind. (secondary) |
| Oct 3 | autonomous VentureBeat article |
MIT and Sakana AI propose SIFT framework using an LLM judge to cut evaluation costs for self-improving coding agents. (secondary) |
| Oct 3 | autonomous Shattered.io article |
Argo-Bench AI benchmark released, top model clears 34.8%. (secondary) |
| Oct 3 | METR Ingenire blog post |
METR 2026 study summarizes AI productivity data, reiterating 19% slowdown finding. (secondary) |
| Oct 3 | orchestration Cloudflare Blog |
Cloudflare blog post calls for building the next Git platform to handle multi-agent coordination. |
| Oct 3 | negative SecNews article |
AI coding agents expose thousands of corporate images to public GitHub repositories. (secondary) |
| Oct 3 | discovery The Fintech Times |
Xsolla launches AI Toolkit embedding game commerce integration into AI coding tools. (secondary) |
| Oct 3 | Claude Code The Next Web article |
Claude services experience outage, Anthropic investigates high error rates. (secondary) |
| Oct 3 | local YouTube video description |
miii, a free offline AI coding agent for the terminal, launched as a Claude Code alternative. (secondary) |
| Oct 3 | autonomous Thoughtworks blog post |
Thoughtworks publishes a practical pattern for reliable coding agents with analyze-blueprint-red-green workflow. |
| Oct 2 | Claude Code GitHub release v2.1.288 |
Claude Code v2.1.288 adds UI selection API, built-in GitHub API, and mods system. |
| Oct 1 | Claude Code GitHub release v2.1.287 |
Claude Code v2.1.287 adds Claude Mods and 'You should know' built-in mod. |
| Sep 30 | Claude Code GitHub release v2.1.286 |
Claude Code v2.1.286 fixes API errors, remote control, and cloud session issues. |
| Sep 21 | Kotlin Benchmark Kotlin Benchmark page |
Kotlin Benchmark leaderboard updated with new agent results. |
| Sep 21 | study Anthropic research |
Anthropic publishes RCT on AI assistance's impact on coding skill formation. |
| Sep 21 | Cursor Hacker News comment |
Hacker News discussion questions if anyone is still using Cursor in 2026. (secondary) |
| Sep 21 | Aider, Claude Code Developers Digest blog |
Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. (secondary) |
| Sep 14 | discovery ProdCodeBench paper on arXiv |
ProdCodeBench benchmark introduced to evaluate AI coding agents on production-derived tasks. |
| Sep 14 | discovery medium.com |
Coding Agent Index 2026 claims to be the first independent benchmark of full agent stacks. (secondary) |
| Sep 14 | METR cerbos.dev |
Blog post cites METR randomized trial finding AI-assisted developers 19% slower. (secondary) |
| Sep 14 | Visual Studio Code, Claude Code, Codex helpnetsecurity.com |
GitHub enables multi-agent AI coding inside repository workflows. (secondary) |
| Sep 14 | discovery raffertyuy.com |
Repo-of-Repos pattern described for multi-repo workspace with AI coding agents. (secondary) |
| Sep 14 | SWE-bench labs.scale.com |
Scale AI publishes public dataset and leaderboard for SWE-bench Pro. |
| Sep 14 | Moderne Moderne Docs for Moddy |
Moderne publishes documentation for its multi-repo AI agent 'Moddy'. (vendor-claim) |
| Sep 7 | launch JetBrains Blog |
JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks. (vendor-claim) |
| Sep 7 | update OpenHands Blog |
OpenHands Enterprise platform enables governed AI agent benchmarking in private cloud. (vendor-claim) |
| Sep 7 | study arXiv |
Google RCT estimates impact of three AI features on developer time for complex tasks. |
| Sep 7 | study alphaXiv |
Randomized controlled experiment measures GitHub Copilot's effect on developer productivity. |
| Sep 7 | orchestration Medium |
Submodule-based multi-repo workspace strategy proposed to avoid file conflicts with concurrent AI agents. (secondary) |
| Sep 7 | orchestration GitHub Discussion |
GitHub acknowledges multi-repo agent capabilities are on the roadmap, recommends task decomposition for now. |
| Sep 7 | Aider, Claude Code, Cursor MorphLLM |
Independent comparison benchmarks Claude Code, Cursor, and Aider on different task types. (secondary) |
| Sep 7 | Claude Code, Cursor, Codex, Aider Requesty AI Blog |
Analysis highlights high operational costs of autonomous agents and need for LLM gateway routing. (secondary) |
| Sep 7 | SWE-bench Paddo.dev |
Analysis illustrates large performance gap between SWE-bench Verified and Pro, highlighting contamination. (secondary) |
| Sep 7 | local Medium |
Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension. (secondary) |
| Sep 7 | pairing Graphite Guides |
Guide outlines best practices for pair programming with AI assistants. (vendor-claim) |
| Sep 7 | pairing ForgeCode Blog |
ForgeCode publishes 12 practical lessons from six months of daily AI pair programming. (vendor-claim) |
| Sep 7 | discovery Martin Fowler |
Multiple sources document concepts and architectures for 'agent harness engineering'. (secondary) |
| Sep 7 | Claude Code Claude Code Docs |
Claude Code docs detail updates for week 20 of 2026, including default fast model change and agent view. |
| Sep 7 | Codex OpenAI Blog |
OpenAI announces Codex integration in ChatGPT mobile app for working from anywhere. (vendor-claim) |
| May 6 | OpenHands Enterprise OpenHands blog |
OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. (vendor-claim) |
| Feb 3 | SWE-bench Yahoo Finance (PRNewswire) |
Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. (vendor-claim) |
| Feb 2 | Codex CNBC article |
OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center. (secondary) |
The last 30 days
The evidence layer got louder this month. METR’s 2025 randomized trial, which found experienced open-source developers were ~19% slower with AI assistance, was cited repeatedly in summaries and blog posts. This negative finding is now a central reference point. On the benchmark front, Scale AI launched SWE-Bench Pro v2, a cleaned-up version designed to be harder to game, while new entrants like Argo-Bench and the vendor-claim-heavy Coding Agent Index arrived. The performance gap remains stark: Bito claims its AI Architect context engine pushed Claude Sonnet 4.5 to 60.8% on SWE-bench Pro, but an analysis noted a 35-point collapse for Claude Opus 4.5 between the contaminated SWE-bench Verified and the cleaner Pro, underscoring how much benchmark design matters.
Tool development was active but pointed toward consolidation and infrastructure. Claude Code had a flurry of releases (v2.1.286 through 288), adding a mods system, a built-in GitHub API, and UI controls, while also suffering a service outage. OpenAI shipped a Codex alpha update. The chatter shifted from which agent to use to how to run them. New platforms launched claiming to manage the fallout: Autoheal (vendor claim) says it handles post-agent work to cut costs, OpenHands Enterprise (vendor claim) offers an on-premise control plane, and miii emerged as a free, local terminal-based alternative. A secondary report noted Aider’s version has been stable while Claude Code advanced, narrowing the model gap.
The multi-agent coordination problem is moving from theory to demanded infrastructure. A Cloudflare blog post explicitly called for building “the next Git platform” to handle coordination when hundreds of agents work on one codebase. Practical patterns were published: Thoughtworks detailed an analyze-blueprint-red-green harness, and the “Repo-of-Repos” workspace pattern was documented to avoid file conflicts. A security report also highlighted a real cost: AI agents have exposed thousands of corporate images to public GitHub repos.
For someone deciding where to spend attention: ignore new vendor benchmarks unless they are contamination-free like SWE-bench Pro. The METR slowdown data is the most important finding to internalize; any productivity claim must now contend with it. The tool race is cooling into an infrastructure build-out phase—watch for platforms that manage agent sprawl, cost, and collision, not just new chat interfaces. Finally, if you run multiple agents, implement a workspace strategy like submodules or Repo-of-Repos now; the collision failure modes are real and the platforms to fix them are just being imagined.
Written 2026-10-03 from the changelog below, not from a fresh search.
Drawn from:
- Claude Code v2.1.288 adds UI selection API, built-in GitHub API, and mods system. — 2026-10-03, confidence
primary-source - OpenAI releases Codex 0.162.0-alpha.9. — 2026-10-03, confidence
primary-source - Scale AI releases SWE-Bench Pro V2 with a cleaner, harder-to-game leaderboard. — 2026-10-03, confidence
primary-source - Autoheal launches, claiming to manage the work AI coding agents leave behind. — 2026-10-03, confidence
secondary - MIT and Sakana AI propose SIFT framework using an LLM judge to cut evaluation costs for self-improving coding agents. — 2026-10-03, confidence
secondary - Argo-Bench AI benchmark released, top model clears 34.8%. — 2026-10-03, confidence
secondary - METR 2026 study summarizes AI productivity data, reiterating 19% slowdown finding. — 2026-10-03, confidence
secondary - Cloudflare blog post calls for building the next Git platform to handle multi-agent coordination. — 2026-10-03, confidence
primary-source - AI coding agents expose thousands of corporate images to public GitHub repositories. — 2026-10-03, confidence
secondary - Xsolla launches AI Toolkit embedding game commerce integration into AI coding tools. — 2026-10-03, confidence
secondary - Claude services experience outage, Anthropic investigates high error rates. — 2026-10-03, confidence
secondary - miii, a free offline AI coding agent for the terminal, launched as a Claude Code alternative. — 2026-10-03, confidence
secondary - Thoughtworks publishes a practical pattern for reliable coding agents with analyze-blueprint-red-green workflow. — 2026-10-03, confidence
primary-source - Claude Code v2.1.287 adds Claude Mods and 'You should know' built-in mod. — 2026-10-03, confidence
primary-source - Claude Code v2.1.286 fixes API errors, remote control, and cloud session issues. — 2026-10-03, confidence
primary-source - Kotlin Benchmark leaderboard updated with new agent results. — 2026-09-21, confidence
primary-source - Anthropic publishes RCT on AI assistance's impact on coding skill formation. — 2026-09-21, confidence
primary-source - Hacker News discussion questions if anyone is still using Cursor in 2026. — 2026-09-21, confidence
secondary - Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. — 2026-09-21, confidence
secondary - OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. — 2026-09-21, confidence
vendor-claim - Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. — 2026-09-21, confidence
vendor-claim - OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center. — 2026-09-21, confidence
secondary - ProdCodeBench benchmark introduced to evaluate AI coding agents on production-derived tasks. — 2026-09-14, confidence
primary-source - Coding Agent Index 2026 claims to be the first independent benchmark of full agent stacks. — 2026-09-14, confidence
secondary - Blog post cites METR randomized trial finding AI-assisted developers 19% slower. — 2026-09-14, confidence
secondary - GitHub enables multi-agent AI coding inside repository workflows. — 2026-09-14, confidence
secondary - Repo-of-Repos pattern described for multi-repo workspace with AI coding agents. — 2026-09-14, confidence
secondary - Scale AI publishes public dataset and leaderboard for SWE-bench Pro. — 2026-09-14, confidence
primary-source - Moderne publishes documentation for its multi-repo AI agent 'Moddy'. — 2026-09-14, confidence
vendor-claim - JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks. — 2026-09-07, confidence
vendor-claim - OpenHands Enterprise platform enables governed AI agent benchmarking in private cloud. — 2026-09-07, confidence
vendor-claim - Google RCT estimates impact of three AI features on developer time for complex tasks. — 2026-09-07, confidence
primary-source - Randomized controlled experiment measures GitHub Copilot's effect on developer productivity. — 2026-09-07, confidence
primary-source - Submodule-based multi-repo workspace strategy proposed to avoid file conflicts with concurrent AI agents. — 2026-09-07, confidence
secondary - GitHub acknowledges multi-repo agent capabilities are on the roadmap, recommends task decomposition for now. — 2026-09-07, confidence
primary-source - Independent comparison benchmarks Claude Code, Cursor, and Aider on different task types. — 2026-09-07, confidence
secondary - Analysis highlights high operational costs of autonomous agents and need for LLM gateway routing. — 2026-09-07, confidence
secondary - Analysis illustrates large performance gap between SWE-bench Verified and Pro, highlighting contamination. — 2026-09-07, confidence
secondary - Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension. — 2026-09-07, confidence
secondary - Guide outlines best practices for pair programming with AI assistants. — 2026-09-07, confidence
vendor-claim - ForgeCode publishes 12 practical lessons from six months of daily AI pair programming. — 2026-09-07, confidence
vendor-claim - Multiple sources document concepts and architectures for 'agent harness engineering'. — 2026-09-07, confidence
secondary - Claude Code docs detail updates for week 20 of 2026, including default fast model change and agent view. — 2026-09-07, confidence
primary-source - OpenAI announces Codex integration in ChatGPT mobile app for working from anywhere. — 2026-09-07, confidence
vendor-claim
The last year
The evidence layer delivered the most important shift this year: the METR randomized trial’s finding that experienced open-source developers were on average 19% slower with AI assistance. This result, reiterated in multiple summaries through 2026, directly contradicts widespread vendor productivity claims and establishes a critical negative baseline. The Google RCT and other studies added further empirical weight, moving the conversation from marketing to measurement. Meanwhile, benchmark contamination became a recognized failure mode, with OpenAI ceasing evaluation on SWE-bench Verified and Scale AI launching SWE-bench Pro V2 specifically to close gaming channels. The gap between vendor-friendly Pass@k metrics and the stricter Pass^k metric, reported as 15-25 percentage points, means most published scores are inflated. For anyone spending money, the takeaway is that independent, contamination-free benchmarks and RCTs are the only trustworthy signals; everything else is noise.
On the tool front, Claude Code was the most active, with a rapid release cadence adding mods, UI APIs, and GitHub integration. In contrast, chatter on Hacker News questioned if anyone still uses Cursor in 2026, and Aider remained at a stable v0.86.2. The launch of miii as a free, offline terminal agent and the standalone Codex app for Apple and Windows show the field diversifying into local and dedicated command-center models. However, a consistent secondary finding is the high operational cost of autonomous agents, with sessions making 50-200+ LLM calls, creating real tension for deployment. Vendor claims like Bito’s AI Architect achieving 60.8% on SWE-bench Pro or Autoheal promising 30% cost reductions remain just that—unverified claims.
Multi-agent coordination moved from theory to a documented problem. Guides from AugmentCode and others explicitly outlined failure modes, like agents independently building the same pipeline. The Repo-of-Repos pattern and submodule strategies were proposed as structural workarounds, while GitHub’s own community discussion admitted multi-repo capabilities are still on the roadmap. Moderne launched ‘Moddy’ as a vendor-claim solution for multi-repo transformation, and Thoughtworks published a detailed harness pattern for reliable multi-agent workflows. The direct experience from this estate—concurrent sessions causing collisions—was mirrored in the broader discussion, confirming that orchestration is unsolved and a major blocker for scaling.
The conceptual framing solidified around “harness engineering.” Articles from OpenAI, Martin Fowler, and Microsoft detailed the scaffolding needed around agents for reliable operation, marking a shift from focusing solely on the model to the surrounding system. This aligns with the launch of platforms like OpenHands Enterprise, offering a governed control plane for on-premise agent management. Security entered the narrative with reports of AI agents accidentally exposing corporate images to public GitHub, a tangible new risk. Finally, the push for specialized agent skills emerged, with Xsolla’s AI Toolkit embedding game commerce API paths—a sign of vertical integration attempts.
In short, the last twelve months were defined by empirical skepticism, harness over hype, and the hard reality of multi-agent chaos. The tools are proliferating, but the evidence says they often slow you down, cost more than advertised, and create coordination messes. The state of the art is now about engineering the system around the agent, not just buying the agent.
Written 2026-10-03 from the changelog below, not from a fresh search.
| Month | Entries |
|---|---|
| October 2026 | 15 |
| September 2026 | 29 |
| August 2026 | 16 |
The ones that mattered:
- 2026-10-03 — METR 2026 study summarizes AI productivity data, reiterating 19% slowdown finding. (
study,secondary) - 2026-10-03 — AI coding agents expose thousands of corporate images to public GitHub repositories. (
negative,secondary) - 2026-10-03 — Claude services experience outage, Anthropic investigates high error rates. (
negative,secondary) - 2026-09-21 — Anthropic publishes RCT on AI assistance's impact on coding skill formation. (
study,primary-source) - 2026-09-21 — Hacker News discussion questions if anyone is still using Cursor in 2026. (
negative,secondary) - 2026-09-21 — OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. (
launch,vendor-claim) - 2026-09-21 — Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. (
study,vendor-claim) - 2026-09-21 — OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center. (
launch,secondary)
All time
The field is now defined by its contradictions. On one side, a flurry of new tools, benchmarks, and frameworks promises autonomous coding. On the other, a single, stubborn piece of evidence from METR’s 2025 randomized controlled trial shows experienced open-source developers were, on average, 19% slower when using AI assistance, despite believing they were faster. This finding, reiterated in a summary as recently as October 2026, is the anchor point. Every vendor claim and benchmark score must be weighed against it. The central question is no longer whether agents can write code—they can—but whether they make the human developer faster, and at what operational and cognitive cost.
The taxonomy splits into three layers: the agents themselves, the benchmarks trying to measure them, and the emerging infrastructure to manage the chaos they create.
Agents and the IDE Wars. The interactive/pairing mode is where daily work happens. Claude Code is the current pace-setter, with a rapid release cadence (v2.1.288 just added a UI selection API and built-in GitHub API for cloud sessions) and deep integration into workflows like VS Code via GitHub Copilot subscriptions. Its main competition is OpenAI’s Codex, which has a standalone desktop app positioned as a multi-agent command center. The older open-source contender, Aider, remains stable at v0.86.2, while Cursor’ relevance is being questioned—a Hacker News comment from September 2026 asked if anyone is still using it, noting a shift to Claude Code and Codex. A new entrant, miii, offers a free, offline, terminal-based alternative, highlighting a demand for local, private operation. The consensus from comparisons is that Claude Code excels at complex, multi-file refactoring, Cursor (where still used) wins on quick bug fixes, and Aider lands in between. All share a critical weakness: high operational cost. Autonomous sessions can make 50-200+ LLM calls, making LLM gateway routing for cost and performance a necessity, not an optimization.
The Benchmark Crisis. Evaluating these tools is a mess of competing claims and methodological landmines. The old standard, SWE-bench Verified, is now widely considered contaminated; OpenAI itself stopped evaluating on it. Its successor, SWE-bench Pro (public dataset hosted by Scale AI), uses GPL-licensed and private repos to combat contamination, causing a dramatic performance collapse—Claude Opus 4.5 dropped from 80.9% on Verified to 45.89% on Pro. Scale AI just released a V2 with a “cleaner, harder-to-game leaderboard.” New benchmarks proliferate: ProjDevBench for end-to-end project tasks, ProdCodeBench focusing on production code changes, Kotlin Benchmark from JetBrains for language-specific evaluation, and the Coding Agent Index claiming to test full agent stacks. The vendor-claim-driven Argo-Bench reports a top model score of 34.8%. The fundamental problem is the gap between metrics like Pass@k (which rewards lucky attempts) and Pass^k (which requires reliability), inflating vendor scores by 15-25 percentage points. Independent analysis rightly criticizes vendor-run benchmarks as biased, akin to pharmaceutical companies grading their own drugs.
What Does Not Work. Multi-agent coordination on a single repository is a proven failure mode without explicit guardrails. This estate experienced it directly on 2026-08-18: two concurrent sessions independently built the same ingestion pipeline, each reporting success. Guides from AugmentCode and others document these coordination challenges. The tools are not built for this. GitHub acknowledges multi-repo capabilities are “on the roadmap,” recommending breaking work into repo-specific sub-tasks for now. Architectural workarounds like the “Repo-of-Repos” pattern or using Git submodules are being proposed to structurally prevent file conflicts. The problem is significant enough that a Cloudflare blog post recently called for building an entirely new Git platform to handle coordination when hundreds of agents work on a codebase.
The Infrastructure Layer Emerges. This is where the most meaningful change is happening, driven by the failures above. The concept of “harness engineering”—the scaffolding around an agent—is now central, with articles from Thoughtworks, Martin Fowler, and OpenAI detailing patterns. Thoughtworks’ “analyze-blueprint-red-green” workflow is a practical example. Companies are building products to manage the aftermath: Autoheal launched claiming to manage the work agents leave behind for up to 30% cost reduction. OpenHands Enterprise offers an on-premise Agent Control Plane for governed benchmarking and scaling. Moderne’s “Moddy” agent focuses on multi-repo transformations at scale. Bito sells a context engine (“AI Architect”) claiming a 60.8% success rate on SWE-bench Pro. Even Xsolla has an AI Toolkit embedding game commerce APIs into agents. This layer exists because raw agent output is unreliable and expensive to clean up.
The Evidence Gap Remains. Beyond the METR slowdown, other studies are emerging. A Google RCT measured the impact of AI features on developer time for complex tasks. Anthropic ran an RCT on whether AI assistance prevents skill formation when learning a new library. The results are nuanced and task-dependent. The loudest vendor benchmarks are the least trustworthy; the quietest academic studies are the most cautionary.
Open Gaps. We lack a reliable, contamination-free benchmark that correlates with real developer productivity gains (or losses). We lack built-in, low-friction coordination protocols for multi-agent work. We lack cost-effective, high-performance local models—analysis suggests effective autonomous agents require 20B+ parameter models, creating a VRAM tension. And we lack a clear answer to the core question: when does an AI coding agent actually make the whole process faster and better, rather than just more automated and more costly?
The field is maturing past the hype of “autonomous coding.” The focus is shifting from the agent to the harness, from raw capability to managed workflow, and from vendor metrics to independent, often negative, evidence. The biggest change since tracking began is the dawning realization that the tool is not the product; the product is the entire system you build to make the tool usable without breaking your project or your budget.
Written 2026-10-03 from the changelog below, not from a fresh search.
How the field breaks down, by what we're actually tracking:
- uncategorised — 22 items: ProjDevBench, METR, Visual Studio Code, Claude Code, Codex, SWE-bench, The Pair, Moderne, Cursor, Aider, Kotlin Benchmark, OpenHands Enterprise, ProdCodeBench, Coding Agent Index, AugmentCode, Scale AI, BenchLM, Bito, Autoheal, Argo-Bench, miii, Xsolla AI Toolkit.
Tracked products
Nothing on this beat is buyable right now, at least nothing we've verified. Everything currently tracked sits in the forward-looking sections below.
Upcoming
Announced, no date.
Nothing in this bucket right now.
Coming Soon
Announced with a date, or an open pre-order.
Nothing in this bucket right now.
What We're Watching
Exists, unproven, or newly discovered. This is where auto-discovered items land.
ProjDevBench
Benchmark for evaluating AI coding agents on end-to-end project development tasks.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: ProjDevBench benchmark evaluates six coding agents on end-to-end project tasks. (2026-08-18) — The benchmark, detailed in an arXiv paper, assesses agents on tasks like from-scratch construction, finding performance varies significantly across models and tasks, with Codex+GPT-5 achieving the best overall score of 77.85%.
Sources: arXiv paper
METR
Organization conducting empirical research on AI safety and impact, including developer productivity studies.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: METR 2026 study summarizes AI productivity data, reiterating 19% slowdown finding. (2026-10-03) — A blog post summarizes METR's 2025 randomized controlled trial which found experienced open-source developers using AI tools were on average 19% slower.
Sources: METR blog
Visual Studio Code
Code editor with integrated AI coding agent support via extensions and GitHub Copilot.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: GitHub enables multi-agent AI coding inside repository workflows. (2026-09-14) — A news article reports that GitHub has enabled Claude and Codex agents within repository workflows, allowing contributors to submit requests and select multiple agents to execute tasks in parallel, with context attached to the work.
Sources: VS Code blog
Claude Code
AI coding agent by Anthropic, available as a CLI and integrated into various IDEs.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: Claude Code v2.1.288 adds UI selection API, built-in GitHub API, and mods system. (2026-10-03) — Anthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh API for cloud sessions, recovery for prompts cleared with Ctrl+C, and --max-findings flag for code review.
Sources: OpenHands comparison blog
Codex
AI coding agent by OpenAI, often integrated into tools like GitHub Copilot and VS Code.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: OpenAI releases Codex 0.162.0-alpha.9. (2026-10-03) — OpenAI published release 0.162.0-alpha.9 of the Codex CLI/desktop app.
Sources: OpenAI harness engineering post
SWE-bench
Benchmark for evaluating AI models on real-world software engineering issues from GitHub.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: Scale AI releases SWE-Bench Pro V2 with a cleaner, harder-to-game leaderboard. (2026-10-03) — Scale AI launched SWE-Bench Pro v2, a refreshed version with the same repositories and tasks but improved data quality and evaluation process to close channels that were open.
Sources: CodeSOTA guide
The Pair
Open-source desktop app for AI pair programming using a multi-agent cross-checking architecture.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: Open-source tool 'The Pair' provides multi-agent coding with cross-checking. (2026-08-18) — A GitHub repository describes The Pair, an open-source desktop app that uses a Mentor+Executor agent architecture to cross-check code and catch hallucinations, compatible with various AI models.
Sources: GitHub repository
Moderne
Company offering a multi-repo AI agent ('Moddy') for large-scale codebase transformation and maintenance.
First seen 2026-08-24 · confidence unverified · uncategorised
Why it's here: Moderne publishes documentation for its multi-repo AI agent 'Moddy'. (2026-09-14) — Moderne has published user documentation for 'Moddy', describing it as an AI agent that modernizes multi-repository codebases by combining LLMs with structured LST code data.
Sources: Moderne blog
Cursor
AI-powered IDE with integrated coding agents, competing with VS Code and Windsurf.
First seen 2026-08-24 · confidence unverified · uncategorised
Why it's here: Hacker News discussion questions if anyone is still using Cursor in 2026. (2026-09-21) — A Hacker News comment from September 2026 asks if anyone is still actually using Cursor, noting that people they've spoken with are now using Claude Code, Codex, or Copilot.
Sources: Cursor vs. Claude Code comparison
Aider
Open-source, terminal-first AI pair programming agent with deep git integration and multi-model support.
First seen 2026-08-24 · confidence unverified · uncategorised
Why it's here: Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. (2026-09-21) — A 2026 update notes Aider remains at version 0.86.2 with a stable maintenance cadence, while Claude Code has advanced past v2.1.200 with auto-mode support and Opus 5 model integration, narrowing the model gap.
Sources: Aider skill page on LobeHub
Kotlin Benchmark
JetBrains' official benchmark for evaluating AI coding agents on real-world Kotlin software engineering tasks.
First seen 2026-09-07 · confidence unverified · uncategorised
Why it's here: Kotlin Benchmark leaderboard updated with new agent results. (2026-09-21) — The official Kotlin Benchmark leaderboard shows updated results for various AI coding agents, with Claude Code + Opus 4.6 medium achieving 71.43% resolution and Codex + GPT 5.3 Codex medium achieving 68.57%.
Sources: JetBrains Blog Announcement
OpenHands Enterprise
Enterprise platform for running governed AI coding agent benchmarks in a private cloud with audit logs and cost attribution.
First seen 2026-09-07 · confidence unverified · uncategorised
Why it's here: OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. (2026-09-21) — OpenHands Enterprise introduces an Agent Control Plane, providing a fully self-hosted, on-premise system for running, controlling, observing, and scaling AI agents across an organization with enforced policies and audit logs.
Sources: OpenHands Blog
ProdCodeBench
Benchmark for evaluating AI coding agents on tasks derived from production code changes.
First seen 2026-09-14 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: ProdCodeBench paper on arXiv
Coding Agent Index
Independent benchmark claiming to evaluate full agent stacks (model + harness).
First seen 2026-09-14 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: Coding Agent Index 2026 article on Medium
AugmentCode
Provider of guides and resources on AI coding practices, including multi-agent coordination.
First seen 2026-09-14 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: AugmentCode guide on multi-agent coding workspace
Scale AI
Company providing AI data and evaluation platforms, including the SWE-bench Pro public dataset.
First seen 2026-09-14 · confidence unverified · uncategorised
Why it's here: Scale AI releases SWE-Bench Pro V2 with a cleaner, harder-to-game leaderboard. (2026-10-03) — Scale AI launched SWE-Bench Pro v2, a refreshed version with the same repositories and tasks but improved data quality and evaluation process to close channels that were open.
Sources: Scale AI SWE-bench Pro leaderboard
BenchLM
Platform providing AI model benchmark leaderboards and evaluations, including for SWE-bench Pro.
First seen 2026-09-14 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: BenchLM SWE-bench Pro leaderboard
Bito
Company building deep context graphs for coding agents, with an AI Architect context engine.
First seen 2026-09-21 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: Company mention (PR)
Autoheal
Platform claiming to manage the work AI coding agents leave behind, aiming to reduce costs.
First seen 2026-10-03 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: VentureBeat article
Argo-Bench
AI agent benchmark evaluating models on agentic tasks.
First seen 2026-10-03 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: Shattered.io article
miii
Local, terminal-based AI coding agent that runs offline as an alternative to Claude Code.
First seen 2026-10-03 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: YouTube video description
Xsolla AI Toolkit
Set of agent skills and plugins embedding game commerce API integration paths into AI coding assistants.
First seen 2026-10-03 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: The Fintech Times
Promising
Early-stage, but the evidence or the approach is genuinely interesting. The only editorial bucket on this page — an item only lands here with a reason recorded in the changelog.
Nothing in this bucket right now.
Open questions
Publishing what we don't know is the point. These are things the job is actively watching for; when one gets answered it becomes a changelog entry and moves down here to the answered list.
No open questions recorded. Either everything we wanted to know got answered, or nobody has written down what we don't know — the second is more likely on a young tracker.
Full changelog
Everything, newest first, grouped by the month we found it. Long by design — it's the receipts.
October 2026
| Date | Subject | Finding |
|---|---|---|
| Oct 3 | Codex GitHub release 0.162.0-alpha.9 |
OpenAI releases Codex 0.162.0-alpha.9. |
| Oct 3 | SWE-bench, Scale AI Scale Labs blog post |
Scale AI releases SWE-Bench Pro V2 with a cleaner, harder-to-game leaderboard. |
| Oct 3 | discovery VentureBeat article |
Autoheal launches, claiming to manage the work AI coding agents leave behind. (secondary) |
| Oct 3 | autonomous VentureBeat article |
MIT and Sakana AI propose SIFT framework using an LLM judge to cut evaluation costs for self-improving coding agents. (secondary) |
| Oct 3 | autonomous Shattered.io article |
Argo-Bench AI benchmark released, top model clears 34.8%. (secondary) |
| Oct 3 | METR Ingenire blog post |
METR 2026 study summarizes AI productivity data, reiterating 19% slowdown finding. (secondary) |
| Oct 3 | orchestration Cloudflare Blog |
Cloudflare blog post calls for building the next Git platform to handle multi-agent coordination. |
| Oct 3 | negative SecNews article |
AI coding agents expose thousands of corporate images to public GitHub repositories. (secondary) |
| Oct 3 | discovery The Fintech Times |
Xsolla launches AI Toolkit embedding game commerce integration into AI coding tools. (secondary) |
| Oct 3 | Claude Code The Next Web article |
Claude services experience outage, Anthropic investigates high error rates. (secondary) |
| Oct 3 | local YouTube video description |
miii, a free offline AI coding agent for the terminal, launched as a Claude Code alternative. (secondary) |
| Oct 3 | autonomous Thoughtworks blog post |
Thoughtworks publishes a practical pattern for reliable coding agents with analyze-blueprint-red-green workflow. |
| Oct 2 | Claude Code GitHub release v2.1.288 |
Claude Code v2.1.288 adds UI selection API, built-in GitHub API, and mods system. |
| Oct 1 | Claude Code GitHub release v2.1.287 |
Claude Code v2.1.287 adds Claude Mods and 'You should know' built-in mod. |
| Sep 30 | Claude Code GitHub release v2.1.286 |
Claude Code v2.1.286 fixes API errors, remote control, and cloud session issues. |
September 2026
| Date | Subject | Finding |
|---|---|---|
| Sep 21 | Kotlin Benchmark Kotlin Benchmark page |
Kotlin Benchmark leaderboard updated with new agent results. |
| Sep 21 | study Anthropic research |
Anthropic publishes RCT on AI assistance's impact on coding skill formation. |
| Sep 21 | Cursor Hacker News comment |
Hacker News discussion questions if anyone is still using Cursor in 2026. (secondary) |
| Sep 21 | Aider, Claude Code Developers Digest blog |
Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. (secondary) |
| Sep 14 | discovery ProdCodeBench paper on arXiv |
ProdCodeBench benchmark introduced to evaluate AI coding agents on production-derived tasks. |
| Sep 14 | discovery medium.com |
Coding Agent Index 2026 claims to be the first independent benchmark of full agent stacks. (secondary) |
| Sep 14 | METR cerbos.dev |
Blog post cites METR randomized trial finding AI-assisted developers 19% slower. (secondary) |
| Sep 14 | Visual Studio Code, Claude Code, Codex helpnetsecurity.com |
GitHub enables multi-agent AI coding inside repository workflows. (secondary) |
| Sep 14 | discovery raffertyuy.com |
Repo-of-Repos pattern described for multi-repo workspace with AI coding agents. (secondary) |
| Sep 14 | SWE-bench labs.scale.com |
Scale AI publishes public dataset and leaderboard for SWE-bench Pro. |
| Sep 14 | Moderne Moderne Docs for Moddy |
Moderne publishes documentation for its multi-repo AI agent 'Moddy'. (vendor-claim) |
| Sep 7 | launch JetBrains Blog |
JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks. (vendor-claim) |
| Sep 7 | update OpenHands Blog |
OpenHands Enterprise platform enables governed AI agent benchmarking in private cloud. (vendor-claim) |
| Sep 7 | study arXiv |
Google RCT estimates impact of three AI features on developer time for complex tasks. |
| Sep 7 | study alphaXiv |
Randomized controlled experiment measures GitHub Copilot's effect on developer productivity. |
| Sep 7 | orchestration Medium |
Submodule-based multi-repo workspace strategy proposed to avoid file conflicts with concurrent AI agents. (secondary) |
| Sep 7 | orchestration GitHub Discussion |
GitHub acknowledges multi-repo agent capabilities are on the roadmap, recommends task decomposition for now. |
| Sep 7 | Aider, Claude Code, Cursor MorphLLM |
Independent comparison benchmarks Claude Code, Cursor, and Aider on different task types. (secondary) |
| Sep 7 | Claude Code, Cursor, Codex, Aider Requesty AI Blog |
Analysis highlights high operational costs of autonomous agents and need for LLM gateway routing. (secondary) |
| Sep 7 | SWE-bench Paddo.dev |
Analysis illustrates large performance gap between SWE-bench Verified and Pro, highlighting contamination. (secondary) |
| Sep 7 | local Medium |
Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension. (secondary) |
| Sep 7 | pairing Graphite Guides |
Guide outlines best practices for pair programming with AI assistants. (vendor-claim) |
| Sep 7 | pairing ForgeCode Blog |
ForgeCode publishes 12 practical lessons from six months of daily AI pair programming. (vendor-claim) |
| Sep 7 | discovery Martin Fowler |
Multiple sources document concepts and architectures for 'agent harness engineering'. (secondary) |
| Sep 7 | Claude Code Claude Code Docs |
Claude Code docs detail updates for week 20 of 2026, including default fast model change and agent view. |
| Sep 7 | Codex OpenAI Blog |
OpenAI announces Codex integration in ChatGPT mobile app for working from anywhere. (vendor-claim) |
| May 6 | OpenHands Enterprise OpenHands blog |
OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. (vendor-claim) |
| Feb 3 | SWE-bench Yahoo Finance (PRNewswire) |
Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. (vendor-claim) |
| Feb 2 | Codex CNBC article |
OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center. (secondary) |
August 2026
| Date | Subject | Finding |
|---|---|---|
| Aug 31 | autonomous SoftwareSeni |
Article highlights gap between Pass@k and Pass^k metrics, inflating vendor-reported scores. (secondary) |
| Aug 31 | SWE-bench OpenAI blog |
OpenAI announces it no longer evaluates on SWE-bench Verified, citing contamination issues. |
| Aug 24 | Claude Code news.ycombinator.com |
Claude Code promotional weekly usage limits are ending, reverting to previous levels. (secondary) |
| Aug 24 | study deepsource.com |
Independent analysis criticizes vendor-run AI code review benchmarks as biased and lacking trustworthy methodology. (secondary) |
| Aug 24 | study arxiv.org |
Google randomized controlled trial estimates the impact of three AI features on developer time for complex tasks. |
| Aug 24 | Moderne moderne.ai |
Moderne launches 'Moddy', a multi-repo AI agent for transforming enterprise codebases at scale. (vendor-claim) |
| Aug 24 | SWE-bench benchlm.ai |
SWE-bench Pro leaderboard for August 2026 shows Claude Mythos 5 leading with 80.3%. (secondary) |
| Aug 19 | Visual Studio Code code.visualstudio.com |
VS Code 1.134 adds features for organizing chats across windows and navigating long conversations. |
| Aug 18 | ProjDevBench arXiv ProjDevBench paper |
ProjDevBench benchmark evaluates six coding agents on end-to-end project tasks. |
| Aug 18 | study METR blog post |
METR randomized trial finds AI assistance made experienced open-source developers 19% slower. |
| Aug 18 | discovery AugmentCode guide |
Guide documents failure modes of running multiple AI coding agents on one repository. (secondary) |
| Aug 18 | Visual Studio Code, Claude Code, Codex VS Code blog |
VS Code adds support for Claude and Codex agents under GitHub Copilot subscription. |
| Aug 18 | SWE-bench arxiv.org |
Criticism highlights contamination concerns in SWE-bench benchmark. (secondary) |
| Aug 18 | The Pair GitHub repository for The Pair |
Open-source tool 'The Pair' provides multi-agent coding with cross-checking. |
| Aug 18 | discovery OpenAI blog post |
OpenAI publishes article on 'harness engineering' for coding agents. |
| Mar 4 | Codex openai.com |
OpenAI's Codex app, initially for macOS, is now available on Windows. |
Dreamlab Live Research uses autonomous systems to track the state of the art in fields of interest. New trackers appear as the interests do.
Generated from a versioned dataset in a git repository: every state this page has ever been in is a commit, and a bad run is revertible. The dataset is the source; this page is build output and is not itself edited.
- Dataset:
data/live-research/ai-programming.json(schema version 2) - Runs completed: 7 · cadence: daily
- Items tracked: 22 · changelog entries: 60
- Distinct sources seen: 286
- Last run recorded: 2026-10-03T05:42:39Z (status:
ok)