Live tracker · machine-maintained

AI Programming

A live tracker of how software actually gets written with machines — coding agents, interactive pairing, multi-agent orchestration, and whether any of it measurably helps.

Last updated 2026-10-03 · updated daily by an automated research job · run #7.

The verdicts

This topic judges "best" along 4 axes, each with its own standing verdict and history.

Autonomous

Judged on: Best evidenced agent for multi-step work handed off with a goal rather than a diff — plans, edits many files, runs its own verification. Judged on independent task benchmarks and reported real-world completion, not vendor demos.

The current state of the art for Autonomous, as of 2026-08-18 (confidence primary-source):

Based on the ProjDevBench benchmark evaluating six coding agents, Codex+GPT-5 achieves the highest overall performance (77.85%) on end-to-end project tasks, though performance gaps widen on from-scratch construction tasks. However, independent empirical evidence from a METR randomized trial suggests AI assistance may actually slow down experienced developers in some contexts.

What triggered or contributed to this call:

Contenders:

Pairing

Judged on: Best evidenced tool for real-time human-and-agent work in one session — interruptibility, shared context, how the human steers mid-run.

The current state of the art for Pairing, as of 2026-08-18 (confidence primary-source):

The open-source tool 'The Pair' introduces a multi-agent architecture with Mentor+Executor cross-checking to catch hallucinations, providing a concrete implementation for AI pair programming. VS Code's integration of Claude and Codex agents offers a widely accessible platform for real-time collaborative work.

What triggered or contributed to this call:

Contenders:

Orchestration

Judged on: Best evidenced way to run several agents on one codebase at once — coordination, isolation, conflict and drift handling.

The current state of the art for Orchestration, as of 2026-08-24 (confidence vendor-claim):

The launch of Moderne's 'Moddy' agent provides a vendor-claimed, concrete implementation for multi-repo, large-scale codebase orchestration, moving beyond single-repo coordination patterns. The previously noted guide from AugmentCode remains the primary independent documentation of failure modes for multi-agent work on a single repository.

What triggered or contributed to this call:

Contenders:

Dethronement history — Orchestration

Every verdict this facet has ever retired, and the receipts for retiring it.

Dethroned 2026-08-24 — held since 2026-08-18

A guide from AugmentCode documents specific failure modes when running multiple AI coding agents on one repository without coordination, highlighting the need for deliberate orchestration. The open-source tool 'The Pair' demonstrates one architectural approach using multiple agents with defined roles.

How it was beaten, and how we know the new one is better: The previous verdict highlighted a guide documenting failure modes and The Pair's single-repo architecture. The new item, Moderne's Moddy, is claimed to operate at a multi-repo, enterprise-scale level, representing a different and more complex orchestration target, though its effectiveness is not yet independently verified.

Triggering entries:

Local

Judged on: Best coding setup that runs entirely on hardware you own, judged on usable quality at a stated VRAM budget, not on leaderboard position.

The current state of the art for Local, as of 2026-09-07 (confidence secondary):

Effective autonomous coding agents require 20B+ parameter models, with 32B+ offering significantly better performance, creating a fundamental tension with VRAM constraints on local hardware. Guides discuss hardware requirements and model options, but no single leading local setup or 'best' model is identified from current evidence that balances high coding capability with typical consumer VRAM budgets.

What triggered or contributed to this call:

Dethronement history — Local

Every verdict this facet has ever retired, and the receipts for retiring it.

Dethroned 2026-09-07 — held since 2026-08-18

Several guides discuss hardware requirements and model options for self-hosted AI coding, noting that effective autonomous agents typically require 20B+ parameter models, creating a tension with VRAM constraints. No single leading local setup is identified from the current results.

How it was beaten, and how we know the new one is better: The previous verdict stated that 'No single leading local setup is identified from the current results.' New evidence from a Medium article based on practical lessons provides a more concrete technical threshold, specifying that true autonomous agents require 20B parameter models minimum, which clarifies the VRAM-performance tension but does not yet identify a leading setup that resolves it.

What this is, and how it works

This page is generated, not written. A scheduled job runs daily on a machine in a homelab. Each run it:

  1. Searches the open web for both the products already tracked here and for category-level terms designed to turn up ones we've never heard of.
  2. Feeds those results to a language model along with everything already on this page, and asks it what is genuinely new. Finding nothing is an acceptable answer, and most runs should find little.
  3. Writes the result into a versioned JSON dataset (data/live-research/ai-programming.json) and commits it. This page is re-rendered from that file.

So: a machine wrote the prose here. Every changelog entry carries at least one source link, and every item and entry carries a confidence label:

Any label weaker than primary-source is printed beside the finding, so an unflagged row is a primary source. A tracker that hides its own uncertainty is worse than no tracker.

What this is not. Not medical advice. Not a review site — nothing here has been tested in our hands. No affiliate relationships, no sponsored placements, nothing bought. Vendor claims are attributed to the vendor rather than restated as findings. Where a price, a date, or a regulatory status is unknown, it is left blank instead of guessed.

What's new

44 developments in the last 30 days, newest first.

Date Subject Finding
Oct 3 Codex
GitHub release 0.162.0-alpha.9
OpenAI releases Codex 0.162.0-alpha.9.
Oct 3 SWE-bench, Scale AI
Scale Labs blog post
Scale AI releases SWE-Bench Pro V2 with a cleaner, harder-to-game leaderboard.
Oct 3 discovery
VentureBeat article
Autoheal launches, claiming to manage the work AI coding agents leave behind. (secondary)
Oct 3 autonomous
VentureBeat article
MIT and Sakana AI propose SIFT framework using an LLM judge to cut evaluation costs for self-improving coding agents. (secondary)
Oct 3 autonomous
Shattered.io article
Argo-Bench AI benchmark released, top model clears 34.8%. (secondary)
Oct 3 METR
Ingenire blog post
METR 2026 study summarizes AI productivity data, reiterating 19% slowdown finding. (secondary)
Oct 3 orchestration
Cloudflare Blog
Cloudflare blog post calls for building the next Git platform to handle multi-agent coordination.
Oct 3 negative
SecNews article
AI coding agents expose thousands of corporate images to public GitHub repositories. (secondary)
Oct 3 discovery
The Fintech Times
Xsolla launches AI Toolkit embedding game commerce integration into AI coding tools. (secondary)
Oct 3 Claude Code
The Next Web article
Claude services experience outage, Anthropic investigates high error rates. (secondary)
Oct 3 local
YouTube video description
miii, a free offline AI coding agent for the terminal, launched as a Claude Code alternative. (secondary)
Oct 3 autonomous
Thoughtworks blog post
Thoughtworks publishes a practical pattern for reliable coding agents with analyze-blueprint-red-green workflow.
Oct 2 Claude Code
GitHub release v2.1.288
Claude Code v2.1.288 adds UI selection API, built-in GitHub API, and mods system.
Oct 1 Claude Code
GitHub release v2.1.287
Claude Code v2.1.287 adds Claude Mods and 'You should know' built-in mod.
Sep 30 Claude Code
GitHub release v2.1.286
Claude Code v2.1.286 fixes API errors, remote control, and cloud session issues.
Sep 21 Kotlin Benchmark
Kotlin Benchmark page
Kotlin Benchmark leaderboard updated with new agent results.
Sep 21 study
Anthropic research
Anthropic publishes RCT on AI assistance's impact on coding skill formation.
Sep 21 Cursor
Hacker News comment
Hacker News discussion questions if anyone is still using Cursor in 2026. (secondary)
Sep 21 Aider, Claude Code
Developers Digest blog
Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. (secondary)
Sep 14 discovery
ProdCodeBench paper on arXiv
ProdCodeBench benchmark introduced to evaluate AI coding agents on production-derived tasks.
Sep 14 discovery
medium.com
Coding Agent Index 2026 claims to be the first independent benchmark of full agent stacks. (secondary)
Sep 14 METR
cerbos.dev
Blog post cites METR randomized trial finding AI-assisted developers 19% slower. (secondary)
Sep 14 Visual Studio Code, Claude Code, Codex
helpnetsecurity.com
GitHub enables multi-agent AI coding inside repository workflows. (secondary)
Sep 14 discovery
raffertyuy.com
Repo-of-Repos pattern described for multi-repo workspace with AI coding agents. (secondary)
Sep 14 SWE-bench
labs.scale.com
Scale AI publishes public dataset and leaderboard for SWE-bench Pro.
Sep 14 Moderne
Moderne Docs for Moddy
Moderne publishes documentation for its multi-repo AI agent 'Moddy'. (vendor-claim)
Sep 7 launch
JetBrains Blog
JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks. (vendor-claim)
Sep 7 update
OpenHands Blog
OpenHands Enterprise platform enables governed AI agent benchmarking in private cloud. (vendor-claim)
Sep 7 study
arXiv
Google RCT estimates impact of three AI features on developer time for complex tasks.
Sep 7 study
alphaXiv
Randomized controlled experiment measures GitHub Copilot's effect on developer productivity.
Sep 7 orchestration
Medium
Submodule-based multi-repo workspace strategy proposed to avoid file conflicts with concurrent AI agents. (secondary)
Sep 7 orchestration
GitHub Discussion
GitHub acknowledges multi-repo agent capabilities are on the roadmap, recommends task decomposition for now.
Sep 7 Aider, Claude Code, Cursor
MorphLLM
Independent comparison benchmarks Claude Code, Cursor, and Aider on different task types. (secondary)
Sep 7 Claude Code, Cursor, Codex, Aider
Requesty AI Blog
Analysis highlights high operational costs of autonomous agents and need for LLM gateway routing. (secondary)
Sep 7 SWE-bench
Paddo.dev
Analysis illustrates large performance gap between SWE-bench Verified and Pro, highlighting contamination. (secondary)
Sep 7 local
Medium
Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension. (secondary)
Sep 7 pairing
Graphite Guides
Guide outlines best practices for pair programming with AI assistants. (vendor-claim)
Sep 7 pairing
ForgeCode Blog
ForgeCode publishes 12 practical lessons from six months of daily AI pair programming. (vendor-claim)
Sep 7 discovery
Martin Fowler
Multiple sources document concepts and architectures for 'agent harness engineering'. (secondary)
Sep 7 Claude Code
Claude Code Docs
Claude Code docs detail updates for week 20 of 2026, including default fast model change and agent view.
Sep 7 Codex
OpenAI Blog
OpenAI announces Codex integration in ChatGPT mobile app for working from anywhere. (vendor-claim)
May 6 OpenHands Enterprise
OpenHands blog
OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. (vendor-claim)
Feb 3 SWE-bench
Yahoo Finance (PRNewswire)
Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. (vendor-claim)
Feb 2 Codex
CNBC article
OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center. (secondary)

The last 30 days

The evidence layer got louder this month. METR’s 2025 randomized trial, which found experienced open-source developers were ~19% slower with AI assistance, was cited repeatedly in summaries and blog posts. This negative finding is now a central reference point. On the benchmark front, Scale AI launched SWE-Bench Pro v2, a cleaned-up version designed to be harder to game, while new entrants like Argo-Bench and the vendor-claim-heavy Coding Agent Index arrived. The performance gap remains stark: Bito claims its AI Architect context engine pushed Claude Sonnet 4.5 to 60.8% on SWE-bench Pro, but an analysis noted a 35-point collapse for Claude Opus 4.5 between the contaminated SWE-bench Verified and the cleaner Pro, underscoring how much benchmark design matters.

Tool development was active but pointed toward consolidation and infrastructure. Claude Code had a flurry of releases (v2.1.286 through 288), adding a mods system, a built-in GitHub API, and UI controls, while also suffering a service outage. OpenAI shipped a Codex alpha update. The chatter shifted from which agent to use to how to run them. New platforms launched claiming to manage the fallout: Autoheal (vendor claim) says it handles post-agent work to cut costs, OpenHands Enterprise (vendor claim) offers an on-premise control plane, and miii emerged as a free, local terminal-based alternative. A secondary report noted Aider’s version has been stable while Claude Code advanced, narrowing the model gap.

The multi-agent coordination problem is moving from theory to demanded infrastructure. A Cloudflare blog post explicitly called for building “the next Git platform” to handle coordination when hundreds of agents work on one codebase. Practical patterns were published: Thoughtworks detailed an analyze-blueprint-red-green harness, and the “Repo-of-Repos” workspace pattern was documented to avoid file conflicts. A security report also highlighted a real cost: AI agents have exposed thousands of corporate images to public GitHub repos.

For someone deciding where to spend attention: ignore new vendor benchmarks unless they are contamination-free like SWE-bench Pro. The METR slowdown data is the most important finding to internalize; any productivity claim must now contend with it. The tool race is cooling into an infrastructure build-out phase—watch for platforms that manage agent sprawl, cost, and collision, not just new chat interfaces. Finally, if you run multiple agents, implement a workspace strategy like submodules or Repo-of-Repos now; the collision failure modes are real and the platforms to fix them are just being imagined.

Written 2026-10-03 from the changelog below, not from a fresh search.

Drawn from:

The last year

The evidence layer delivered the most important shift this year: the METR randomized trial’s finding that experienced open-source developers were on average 19% slower with AI assistance. This result, reiterated in multiple summaries through 2026, directly contradicts widespread vendor productivity claims and establishes a critical negative baseline. The Google RCT and other studies added further empirical weight, moving the conversation from marketing to measurement. Meanwhile, benchmark contamination became a recognized failure mode, with OpenAI ceasing evaluation on SWE-bench Verified and Scale AI launching SWE-bench Pro V2 specifically to close gaming channels. The gap between vendor-friendly Pass@k metrics and the stricter Pass^k metric, reported as 15-25 percentage points, means most published scores are inflated. For anyone spending money, the takeaway is that independent, contamination-free benchmarks and RCTs are the only trustworthy signals; everything else is noise.

On the tool front, Claude Code was the most active, with a rapid release cadence adding mods, UI APIs, and GitHub integration. In contrast, chatter on Hacker News questioned if anyone still uses Cursor in 2026, and Aider remained at a stable v0.86.2. The launch of miii as a free, offline terminal agent and the standalone Codex app for Apple and Windows show the field diversifying into local and dedicated command-center models. However, a consistent secondary finding is the high operational cost of autonomous agents, with sessions making 50-200+ LLM calls, creating real tension for deployment. Vendor claims like Bito’s AI Architect achieving 60.8% on SWE-bench Pro or Autoheal promising 30% cost reductions remain just that—unverified claims.

Multi-agent coordination moved from theory to a documented problem. Guides from AugmentCode and others explicitly outlined failure modes, like agents independently building the same pipeline. The Repo-of-Repos pattern and submodule strategies were proposed as structural workarounds, while GitHub’s own community discussion admitted multi-repo capabilities are still on the roadmap. Moderne launched ‘Moddy’ as a vendor-claim solution for multi-repo transformation, and Thoughtworks published a detailed harness pattern for reliable multi-agent workflows. The direct experience from this estate—concurrent sessions causing collisions—was mirrored in the broader discussion, confirming that orchestration is unsolved and a major blocker for scaling.

The conceptual framing solidified around “harness engineering.” Articles from OpenAI, Martin Fowler, and Microsoft detailed the scaffolding needed around agents for reliable operation, marking a shift from focusing solely on the model to the surrounding system. This aligns with the launch of platforms like OpenHands Enterprise, offering a governed control plane for on-premise agent management. Security entered the narrative with reports of AI agents accidentally exposing corporate images to public GitHub, a tangible new risk. Finally, the push for specialized agent skills emerged, with Xsolla’s AI Toolkit embedding game commerce API paths—a sign of vertical integration attempts.

In short, the last twelve months were defined by empirical skepticism, harness over hype, and the hard reality of multi-agent chaos. The tools are proliferating, but the evidence says they often slow you down, cost more than advertised, and create coordination messes. The state of the art is now about engineering the system around the agent, not just buying the agent.

Written 2026-10-03 from the changelog below, not from a fresh search.

Month Entries
October 2026 15
September 2026 29
August 2026 16

The ones that mattered:

All time

The field is now defined by its contradictions. On one side, a flurry of new tools, benchmarks, and frameworks promises autonomous coding. On the other, a single, stubborn piece of evidence from METR’s 2025 randomized controlled trial shows experienced open-source developers were, on average, 19% slower when using AI assistance, despite believing they were faster. This finding, reiterated in a summary as recently as October 2026, is the anchor point. Every vendor claim and benchmark score must be weighed against it. The central question is no longer whether agents can write code—they can—but whether they make the human developer faster, and at what operational and cognitive cost.

The taxonomy splits into three layers: the agents themselves, the benchmarks trying to measure them, and the emerging infrastructure to manage the chaos they create.

Agents and the IDE Wars. The interactive/pairing mode is where daily work happens. Claude Code is the current pace-setter, with a rapid release cadence (v2.1.288 just added a UI selection API and built-in GitHub API for cloud sessions) and deep integration into workflows like VS Code via GitHub Copilot subscriptions. Its main competition is OpenAI’s Codex, which has a standalone desktop app positioned as a multi-agent command center. The older open-source contender, Aider, remains stable at v0.86.2, while Cursor’ relevance is being questioned—a Hacker News comment from September 2026 asked if anyone is still using it, noting a shift to Claude Code and Codex. A new entrant, miii, offers a free, offline, terminal-based alternative, highlighting a demand for local, private operation. The consensus from comparisons is that Claude Code excels at complex, multi-file refactoring, Cursor (where still used) wins on quick bug fixes, and Aider lands in between. All share a critical weakness: high operational cost. Autonomous sessions can make 50-200+ LLM calls, making LLM gateway routing for cost and performance a necessity, not an optimization.

The Benchmark Crisis. Evaluating these tools is a mess of competing claims and methodological landmines. The old standard, SWE-bench Verified, is now widely considered contaminated; OpenAI itself stopped evaluating on it. Its successor, SWE-bench Pro (public dataset hosted by Scale AI), uses GPL-licensed and private repos to combat contamination, causing a dramatic performance collapse—Claude Opus 4.5 dropped from 80.9% on Verified to 45.89% on Pro. Scale AI just released a V2 with a “cleaner, harder-to-game leaderboard.” New benchmarks proliferate: ProjDevBench for end-to-end project tasks, ProdCodeBench focusing on production code changes, Kotlin Benchmark from JetBrains for language-specific evaluation, and the Coding Agent Index claiming to test full agent stacks. The vendor-claim-driven Argo-Bench reports a top model score of 34.8%. The fundamental problem is the gap between metrics like Pass@k (which rewards lucky attempts) and Pass^k (which requires reliability), inflating vendor scores by 15-25 percentage points. Independent analysis rightly criticizes vendor-run benchmarks as biased, akin to pharmaceutical companies grading their own drugs.

What Does Not Work. Multi-agent coordination on a single repository is a proven failure mode without explicit guardrails. This estate experienced it directly on 2026-08-18: two concurrent sessions independently built the same ingestion pipeline, each reporting success. Guides from AugmentCode and others document these coordination challenges. The tools are not built for this. GitHub acknowledges multi-repo capabilities are “on the roadmap,” recommending breaking work into repo-specific sub-tasks for now. Architectural workarounds like the “Repo-of-Repos” pattern or using Git submodules are being proposed to structurally prevent file conflicts. The problem is significant enough that a Cloudflare blog post recently called for building an entirely new Git platform to handle coordination when hundreds of agents work on a codebase.

The Infrastructure Layer Emerges. This is where the most meaningful change is happening, driven by the failures above. The concept of “harness engineering”—the scaffolding around an agent—is now central, with articles from Thoughtworks, Martin Fowler, and OpenAI detailing patterns. Thoughtworks’ “analyze-blueprint-red-green” workflow is a practical example. Companies are building products to manage the aftermath: Autoheal launched claiming to manage the work agents leave behind for up to 30% cost reduction. OpenHands Enterprise offers an on-premise Agent Control Plane for governed benchmarking and scaling. Moderne’s “Moddy” agent focuses on multi-repo transformations at scale. Bito sells a context engine (“AI Architect”) claiming a 60.8% success rate on SWE-bench Pro. Even Xsolla has an AI Toolkit embedding game commerce APIs into agents. This layer exists because raw agent output is unreliable and expensive to clean up.

The Evidence Gap Remains. Beyond the METR slowdown, other studies are emerging. A Google RCT measured the impact of AI features on developer time for complex tasks. Anthropic ran an RCT on whether AI assistance prevents skill formation when learning a new library. The results are nuanced and task-dependent. The loudest vendor benchmarks are the least trustworthy; the quietest academic studies are the most cautionary.

Open Gaps. We lack a reliable, contamination-free benchmark that correlates with real developer productivity gains (or losses). We lack built-in, low-friction coordination protocols for multi-agent work. We lack cost-effective, high-performance local models—analysis suggests effective autonomous agents require 20B+ parameter models, creating a VRAM tension. And we lack a clear answer to the core question: when does an AI coding agent actually make the whole process faster and better, rather than just more automated and more costly?

The field is maturing past the hype of “autonomous coding.” The focus is shifting from the agent to the harness, from raw capability to managed workflow, and from vendor metrics to independent, often negative, evidence. The biggest change since tracking began is the dawning realization that the tool is not the product; the product is the entire system you build to make the tool usable without breaking your project or your budget.

Written 2026-10-03 from the changelog below, not from a fresh search.

How the field breaks down, by what we're actually tracking:

Tracked products

Nothing on this beat is buyable right now, at least nothing we've verified. Everything currently tracked sits in the forward-looking sections below.

Upcoming

Announced, no date.

Nothing in this bucket right now.

Coming Soon

Announced with a date, or an open pre-order.

Nothing in this bucket right now.

What We're Watching

Exists, unproven, or newly discovered. This is where auto-discovered items land.

ProjDevBench

Benchmark for evaluating AI coding agents on end-to-end project development tasks.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: ProjDevBench benchmark evaluates six coding agents on end-to-end project tasks. (2026-08-18) — The benchmark, detailed in an arXiv paper, assesses agents on tasks like from-scratch construction, finding performance varies significantly across models and tasks, with Codex+GPT-5 achieving the best overall score of 77.85%.

Sources: arXiv paper

METR

Organization conducting empirical research on AI safety and impact, including developer productivity studies.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: METR 2026 study summarizes AI productivity data, reiterating 19% slowdown finding. (2026-10-03) — A blog post summarizes METR's 2025 randomized controlled trial which found experienced open-source developers using AI tools were on average 19% slower.

Sources: METR blog

Visual Studio Code

Code editor with integrated AI coding agent support via extensions and GitHub Copilot.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: GitHub enables multi-agent AI coding inside repository workflows. (2026-09-14) — A news article reports that GitHub has enabled Claude and Codex agents within repository workflows, allowing contributors to submit requests and select multiple agents to execute tasks in parallel, with context attached to the work.

Sources: VS Code blog

Claude Code

AI coding agent by Anthropic, available as a CLI and integrated into various IDEs.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: Claude Code v2.1.288 adds UI selection API, built-in GitHub API, and mods system. (2026-10-03) — Anthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh API for cloud sessions, recovery for prompts cleared with Ctrl+C, and --max-findings flag for code review.

Sources: OpenHands comparison blog

Codex

AI coding agent by OpenAI, often integrated into tools like GitHub Copilot and VS Code.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: OpenAI releases Codex 0.162.0-alpha.9. (2026-10-03) — OpenAI published release 0.162.0-alpha.9 of the Codex CLI/desktop app.

Sources: OpenAI harness engineering post

SWE-bench

Benchmark for evaluating AI models on real-world software engineering issues from GitHub.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: Scale AI releases SWE-Bench Pro V2 with a cleaner, harder-to-game leaderboard. (2026-10-03) — Scale AI launched SWE-Bench Pro v2, a refreshed version with the same repositories and tasks but improved data quality and evaluation process to close channels that were open.

Sources: CodeSOTA guide

The Pair

Open-source desktop app for AI pair programming using a multi-agent cross-checking architecture.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: Open-source tool 'The Pair' provides multi-agent coding with cross-checking. (2026-08-18) — A GitHub repository describes The Pair, an open-source desktop app that uses a Mentor+Executor agent architecture to cross-check code and catch hallucinations, compatible with various AI models.

Sources: GitHub repository

Moderne

Company offering a multi-repo AI agent ('Moddy') for large-scale codebase transformation and maintenance.

First seen 2026-08-24 · confidence unverified · uncategorised

Why it's here: Moderne publishes documentation for its multi-repo AI agent 'Moddy'. (2026-09-14) — Moderne has published user documentation for 'Moddy', describing it as an AI agent that modernizes multi-repository codebases by combining LLMs with structured LST code data.

Sources: Moderne blog

Cursor

AI-powered IDE with integrated coding agents, competing with VS Code and Windsurf.

First seen 2026-08-24 · confidence unverified · uncategorised

Why it's here: Hacker News discussion questions if anyone is still using Cursor in 2026. (2026-09-21) — A Hacker News comment from September 2026 asks if anyone is still actually using Cursor, noting that people they've spoken with are now using Claude Code, Codex, or Copilot.

Sources: Cursor vs. Claude Code comparison

Aider

Open-source, terminal-first AI pair programming agent with deep git integration and multi-model support.

First seen 2026-08-24 · confidence unverified · uncategorised

Why it's here: Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. (2026-09-21) — A 2026 update notes Aider remains at version 0.86.2 with a stable maintenance cadence, while Claude Code has advanced past v2.1.200 with auto-mode support and Opus 5 model integration, narrowing the model gap.

Sources: Aider skill page on LobeHub

Kotlin Benchmark

JetBrains' official benchmark for evaluating AI coding agents on real-world Kotlin software engineering tasks.

First seen 2026-09-07 · confidence unverified · uncategorised

Why it's here: Kotlin Benchmark leaderboard updated with new agent results. (2026-09-21) — The official Kotlin Benchmark leaderboard shows updated results for various AI coding agents, with Claude Code + Opus 4.6 medium achieving 71.43% resolution and Codex + GPT 5.3 Codex medium achieving 68.57%.

Sources: JetBrains Blog Announcement

OpenHands Enterprise

Enterprise platform for running governed AI coding agent benchmarks in a private cloud with audit logs and cost attribution.

First seen 2026-09-07 · confidence unverified · uncategorised

Why it's here: OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. (2026-09-21) — OpenHands Enterprise introduces an Agent Control Plane, providing a fully self-hosted, on-premise system for running, controlling, observing, and scaling AI agents across an organization with enforced policies and audit logs.

Sources: OpenHands Blog

ProdCodeBench

Benchmark for evaluating AI coding agents on tasks derived from production code changes.

First seen 2026-09-14 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: ProdCodeBench paper on arXiv

Coding Agent Index

Independent benchmark claiming to evaluate full agent stacks (model + harness).

First seen 2026-09-14 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: Coding Agent Index 2026 article on Medium

AugmentCode

Provider of guides and resources on AI coding practices, including multi-agent coordination.

First seen 2026-09-14 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: AugmentCode guide on multi-agent coding workspace

Scale AI

Company providing AI data and evaluation platforms, including the SWE-bench Pro public dataset.

First seen 2026-09-14 · confidence unverified · uncategorised

Why it's here: Scale AI releases SWE-Bench Pro V2 with a cleaner, harder-to-game leaderboard. (2026-10-03) — Scale AI launched SWE-Bench Pro v2, a refreshed version with the same repositories and tasks but improved data quality and evaluation process to close channels that were open.

Sources: Scale AI SWE-bench Pro leaderboard

BenchLM

Platform providing AI model benchmark leaderboards and evaluations, including for SWE-bench Pro.

First seen 2026-09-14 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: BenchLM SWE-bench Pro leaderboard

Bito

Company building deep context graphs for coding agents, with an AI Architect context engine.

First seen 2026-09-21 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: Company mention (PR)

Autoheal

Platform claiming to manage the work AI coding agents leave behind, aiming to reduce costs.

First seen 2026-10-03 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: VentureBeat article

Argo-Bench

AI agent benchmark evaluating models on agentic tasks.

First seen 2026-10-03 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: Shattered.io article

miii

Local, terminal-based AI coding agent that runs offline as an alternative to Claude Code.

First seen 2026-10-03 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: YouTube video description

Xsolla AI Toolkit

Set of agent skills and plugins embedding game commerce API integration paths into AI coding assistants.

First seen 2026-10-03 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: The Fintech Times

Promising

Early-stage, but the evidence or the approach is genuinely interesting. The only editorial bucket on this page — an item only lands here with a reason recorded in the changelog.

Nothing in this bucket right now.

Open questions

Publishing what we don't know is the point. These are things the job is actively watching for; when one gets answered it becomes a changelog entry and moves down here to the answered list.

No open questions recorded. Either everything we wanted to know got answered, or nobody has written down what we don't know — the second is more likely on a young tracker.

Full changelog

Everything, newest first, grouped by the month we found it. Long by design — it's the receipts.

October 2026

Date Subject Finding
Oct 3 Codex
GitHub release 0.162.0-alpha.9
OpenAI releases Codex 0.162.0-alpha.9.
Oct 3 SWE-bench, Scale AI
Scale Labs blog post
Scale AI releases SWE-Bench Pro V2 with a cleaner, harder-to-game leaderboard.
Oct 3 discovery
VentureBeat article
Autoheal launches, claiming to manage the work AI coding agents leave behind. (secondary)
Oct 3 autonomous
VentureBeat article
MIT and Sakana AI propose SIFT framework using an LLM judge to cut evaluation costs for self-improving coding agents. (secondary)
Oct 3 autonomous
Shattered.io article
Argo-Bench AI benchmark released, top model clears 34.8%. (secondary)
Oct 3 METR
Ingenire blog post
METR 2026 study summarizes AI productivity data, reiterating 19% slowdown finding. (secondary)
Oct 3 orchestration
Cloudflare Blog
Cloudflare blog post calls for building the next Git platform to handle multi-agent coordination.
Oct 3 negative
SecNews article
AI coding agents expose thousands of corporate images to public GitHub repositories. (secondary)
Oct 3 discovery
The Fintech Times
Xsolla launches AI Toolkit embedding game commerce integration into AI coding tools. (secondary)
Oct 3 Claude Code
The Next Web article
Claude services experience outage, Anthropic investigates high error rates. (secondary)
Oct 3 local
YouTube video description
miii, a free offline AI coding agent for the terminal, launched as a Claude Code alternative. (secondary)
Oct 3 autonomous
Thoughtworks blog post
Thoughtworks publishes a practical pattern for reliable coding agents with analyze-blueprint-red-green workflow.
Oct 2 Claude Code
GitHub release v2.1.288
Claude Code v2.1.288 adds UI selection API, built-in GitHub API, and mods system.
Oct 1 Claude Code
GitHub release v2.1.287
Claude Code v2.1.287 adds Claude Mods and 'You should know' built-in mod.
Sep 30 Claude Code
GitHub release v2.1.286
Claude Code v2.1.286 fixes API errors, remote control, and cloud session issues.

September 2026

Date Subject Finding
Sep 21 Kotlin Benchmark
Kotlin Benchmark page
Kotlin Benchmark leaderboard updated with new agent results.
Sep 21 study
Anthropic research
Anthropic publishes RCT on AI assistance's impact on coding skill formation.
Sep 21 Cursor
Hacker News comment
Hacker News discussion questions if anyone is still using Cursor in 2026. (secondary)
Sep 21 Aider, Claude Code
Developers Digest blog
Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. (secondary)
Sep 14 discovery
ProdCodeBench paper on arXiv
ProdCodeBench benchmark introduced to evaluate AI coding agents on production-derived tasks.
Sep 14 discovery
medium.com
Coding Agent Index 2026 claims to be the first independent benchmark of full agent stacks. (secondary)
Sep 14 METR
cerbos.dev
Blog post cites METR randomized trial finding AI-assisted developers 19% slower. (secondary)
Sep 14 Visual Studio Code, Claude Code, Codex
helpnetsecurity.com
GitHub enables multi-agent AI coding inside repository workflows. (secondary)
Sep 14 discovery
raffertyuy.com
Repo-of-Repos pattern described for multi-repo workspace with AI coding agents. (secondary)
Sep 14 SWE-bench
labs.scale.com
Scale AI publishes public dataset and leaderboard for SWE-bench Pro.
Sep 14 Moderne
Moderne Docs for Moddy
Moderne publishes documentation for its multi-repo AI agent 'Moddy'. (vendor-claim)
Sep 7 launch
JetBrains Blog
JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks. (vendor-claim)
Sep 7 update
OpenHands Blog
OpenHands Enterprise platform enables governed AI agent benchmarking in private cloud. (vendor-claim)
Sep 7 study
arXiv
Google RCT estimates impact of three AI features on developer time for complex tasks.
Sep 7 study
alphaXiv
Randomized controlled experiment measures GitHub Copilot's effect on developer productivity.
Sep 7 orchestration
Medium
Submodule-based multi-repo workspace strategy proposed to avoid file conflicts with concurrent AI agents. (secondary)
Sep 7 orchestration
GitHub Discussion
GitHub acknowledges multi-repo agent capabilities are on the roadmap, recommends task decomposition for now.
Sep 7 Aider, Claude Code, Cursor
MorphLLM
Independent comparison benchmarks Claude Code, Cursor, and Aider on different task types. (secondary)
Sep 7 Claude Code, Cursor, Codex, Aider
Requesty AI Blog
Analysis highlights high operational costs of autonomous agents and need for LLM gateway routing. (secondary)
Sep 7 SWE-bench
Paddo.dev
Analysis illustrates large performance gap between SWE-bench Verified and Pro, highlighting contamination. (secondary)
Sep 7 local
Medium
Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension. (secondary)
Sep 7 pairing
Graphite Guides
Guide outlines best practices for pair programming with AI assistants. (vendor-claim)
Sep 7 pairing
ForgeCode Blog
ForgeCode publishes 12 practical lessons from six months of daily AI pair programming. (vendor-claim)
Sep 7 discovery
Martin Fowler
Multiple sources document concepts and architectures for 'agent harness engineering'. (secondary)
Sep 7 Claude Code
Claude Code Docs
Claude Code docs detail updates for week 20 of 2026, including default fast model change and agent view.
Sep 7 Codex
OpenAI Blog
OpenAI announces Codex integration in ChatGPT mobile app for working from anywhere. (vendor-claim)
May 6 OpenHands Enterprise
OpenHands blog
OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. (vendor-claim)
Feb 3 SWE-bench
Yahoo Finance (PRNewswire)
Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. (vendor-claim)
Feb 2 Codex
CNBC article
OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center. (secondary)

August 2026

Date Subject Finding
Aug 31 autonomous
SoftwareSeni
Article highlights gap between Pass@k and Pass^k metrics, inflating vendor-reported scores. (secondary)
Aug 31 SWE-bench
OpenAI blog
OpenAI announces it no longer evaluates on SWE-bench Verified, citing contamination issues.
Aug 24 Claude Code
news.ycombinator.com
Claude Code promotional weekly usage limits are ending, reverting to previous levels. (secondary)
Aug 24 study
deepsource.com
Independent analysis criticizes vendor-run AI code review benchmarks as biased and lacking trustworthy methodology. (secondary)
Aug 24 study
arxiv.org
Google randomized controlled trial estimates the impact of three AI features on developer time for complex tasks.
Aug 24 Moderne
moderne.ai
Moderne launches 'Moddy', a multi-repo AI agent for transforming enterprise codebases at scale. (vendor-claim)
Aug 24 SWE-bench
benchlm.ai
SWE-bench Pro leaderboard for August 2026 shows Claude Mythos 5 leading with 80.3%. (secondary)
Aug 19 Visual Studio Code
code.visualstudio.com
VS Code 1.134 adds features for organizing chats across windows and navigating long conversations.
Aug 18 ProjDevBench
arXiv ProjDevBench paper
ProjDevBench benchmark evaluates six coding agents on end-to-end project tasks.
Aug 18 study
METR blog post
METR randomized trial finds AI assistance made experienced open-source developers 19% slower.
Aug 18 discovery
AugmentCode guide
Guide documents failure modes of running multiple AI coding agents on one repository. (secondary)
Aug 18 Visual Studio Code, Claude Code, Codex
VS Code blog
VS Code adds support for Claude and Codex agents under GitHub Copilot subscription.
Aug 18 SWE-bench
arxiv.org
Criticism highlights contamination concerns in SWE-bench benchmark. (secondary)
Aug 18 The Pair
GitHub repository for The Pair
Open-source tool 'The Pair' provides multi-agent coding with cross-checking.
Aug 18 discovery
OpenAI blog post
OpenAI publishes article on 'harness engineering' for coding agents.
Mar 4 Codex
openai.com
OpenAI's Codex app, initially for macOS, is now available on Windows.

Dreamlab Live Research uses autonomous systems to track the state of the art in fields of interest. New trackers appear as the interests do.


Generated from a versioned dataset in a git repository: every state this page has ever been in is a commit, and a bad run is revertible. The dataset is the source; this page is build output and is not itself edited.