2026-08-18

AI Programming

A live tracker of how software actually gets written with machines — coding agents, interactive pairing, multi-agent orchestration, and whether any of it measurably helps.

Last updated 2026-09-21 · updated weekly by an automated research job · run #6.

The verdicts

This topic judges "best" along 4 axes, each with its own standing verdict and history.

Autonomous

Judged on: Best evidenced agent for multi-step work handed off with a goal rather than a diff — plans, edits many files, runs its own verification. Judged on independent task benchmarks and reported real-world completion, not vendor demos.

The current state of the art for Autonomous, as of 2026-08-18 (confidence primary-source):

Based on the ProjDevBench benchmark evaluating six coding agents, Codex+GPT-5 achieves the highest overall performance (77.85%) on end-to-end project tasks, though performance gaps widen on from-scratch construction tasks. However, independent empirical evidence from a METR randomized trial suggests AI assistance may actually slow down experienced developers in some contexts.

What triggered or contributed to this call:

Contenders:

Pairing

Judged on: Best evidenced tool for real-time human-and-agent work in one session — interruptibility, shared context, how the human steers mid-run.

The current state of the art for Pairing, as of 2026-08-18 (confidence primary-source):

The open-source tool 'The Pair' introduces a multi-agent architecture with Mentor+Executor cross-checking to catch hallucinations, providing a concrete implementation for AI pair programming. VS Code's integration of Claude and Codex agents offers a widely accessible platform for real-time collaborative work.

What triggered or contributed to this call:

Contenders:

Orchestration

Judged on: Best evidenced way to run several agents on one codebase at once — coordination, isolation, conflict and drift handling.

The current state of the art for Orchestration, as of 2026-08-24 (confidence vendor-claim):

The launch of Moderne's 'Moddy' agent provides a vendor-claimed, concrete implementation for multi-repo, large-scale codebase orchestration, moving beyond single-repo coordination patterns. The previously noted guide from AugmentCode remains the primary independent documentation of failure modes for multi-agent work on a single repository.

What triggered or contributed to this call:

Contenders:

Dethronement history — Orchestration

Every verdict this facet has ever retired, and the receipts for retiring it.

Dethroned 2026-08-24 — held since 2026-08-18

A guide from AugmentCode documents specific failure modes when running multiple AI coding agents on one repository without coordination, highlighting the need for deliberate orchestration. The open-source tool 'The Pair' demonstrates one architectural approach using multiple agents with defined roles.

How it was beaten, and how we know the new one is better: The previous verdict highlighted a guide documenting failure modes and The Pair's single-repo architecture. The new item, Moderne's Moddy, is claimed to operate at a multi-repo, enterprise-scale level, representing a different and more complex orchestration target, though its effectiveness is not yet independently verified.

Triggering entries:

Local

Judged on: Best coding setup that runs entirely on hardware you own, judged on usable quality at a stated VRAM budget, not on leaderboard position.

The current state of the art for Local, as of 2026-09-07 (confidence secondary):

Effective autonomous coding agents require 20B+ parameter models, with 32B+ offering significantly better performance, creating a fundamental tension with VRAM constraints on local hardware. Guides discuss hardware requirements and model options, but no single leading local setup or 'best' model is identified from current evidence that balances high coding capability with typical consumer VRAM budgets.

What triggered or contributed to this call:

Dethronement history — Local

Every verdict this facet has ever retired, and the receipts for retiring it.

Dethroned 2026-09-07 — held since 2026-08-18

Several guides discuss hardware requirements and model options for self-hosted AI coding, noting that effective autonomous agents typically require 20B+ parameter models, creating a tension with VRAM constraints. No single leading local setup is identified from the current results.

How it was beaten, and how we know the new one is better: The previous verdict stated that 'No single leading local setup is identified from the current results.' New evidence from a Medium article based on practical lessons provides a more concrete technical threshold, specifying that true autonomous agents require 20B parameter models minimum, which clarifies the VRAM-performance tension but does not yet identify a leading setup that resolves it.

What this is, and how it works

This page is generated, not written. A scheduled job runs weekly on a machine in a homelab. Each run it:

  1. Searches the open web for both the products already tracked here and for category-level terms designed to turn up ones we've never heard of.
  2. Feeds those results to a language model along with everything already on this page, and asks it what is genuinely new. Finding nothing is an acceptable answer, and most runs should find little.
  3. Writes the result into a versioned JSON dataset (data/live-research/ai-programming.json) and commits it. This page is re-rendered from that file.

So: a machine wrote the prose here. Every changelog entry carries at least one source link, and every item and entry carries a confidence label:

Those labels are always shown, including on the weak ones. A tracker that hides its own uncertainty is worse than no tracker.

What this is not. Not medical advice. Not a review site — nothing here has been tested in our hands. No affiliate relationships, no sponsored placements, nothing bought. Vendor claims are attributed to the vendor rather than restated as findings. Where a price, a date, or a regulatory status is unknown, it is left blank instead of guessed.

What's new

29 developments in the last 30 days, newest first.

Kotlin Benchmark leaderboard updated with new agent results.

Discovered 2026-09-21 · publication date unknown · update · confidence: primary-source

The official Kotlin Benchmark leaderboard shows updated results for various AI coding agents, with Claude Code + Opus 4.6 medium achieving 71.43% resolution and Codex + GPT 5.3 Codex medium achieving 68.57%.

Sources: Kotlin Benchmark page

Anthropic publishes RCT on AI assistance's impact on coding skill formation.

Discovered 2026-09-21 · publication date unknown · study · confidence: primary-source

Anthropic conducted a randomized controlled trial investigating whether AI assistance prevents software developers from growing their skills or understanding the systems they're building when learning a new Python library.

Sources: Anthropic research

Hacker News discussion questions if anyone is still using Cursor in 2026.

Discovered 2026-09-21 · publication date unknown · negative · confidence: secondary

A Hacker News comment from September 2026 asks if anyone is still actually using Cursor, noting that people they've spoken with are now using Claude Code, Codex, or Copilot.

Sources: Hacker News comment

Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration.

Discovered 2026-09-21 · publication date unknown · update · confidence: secondary

A 2026 update notes Aider remains at version 0.86.2 with a stable maintenance cadence, while Claude Code has advanced past v2.1.200 with auto-mode support and Opus 5 model integration, narrowing the model gap.

Sources: Developers Digest blog

OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management.

Discovered 2026-09-21 · published 2026-05-06 · launch · confidence: vendor-claim

OpenHands Enterprise introduces an Agent Control Plane, providing a fully self-hosted, on-premise system for running, controlling, observing, and scaling AI agents across an organization with enforced policies and audit logs.

Sources: OpenHands blog

Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5.

Discovered 2026-09-21 · published 2026-02-03 · study · confidence: vendor-claim

Bito announced evaluation results showing a Claude Sonnet 4.5 agent augmented with its AI Architect context engine achieved a 60.8% success rate on the SWE-Bench Pro benchmark.

Sources: Yahoo Finance (PRNewswire)

OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center.

Discovered 2026-09-21 · published 2026-02-02 · launch · confidence: secondary

OpenAI launched a standalone desktop app for its Codex AI coding assistant, temporarily available to ChatGPT users with Apple computers, designed as a 'command center' to manage multiple AI agents.

Sources: CNBC article

ProdCodeBench benchmark introduced to evaluate AI coding agents on production-derived tasks.

Discovered 2026-09-14 · publication date unknown · discovery · confidence: primary-source

A new benchmark, ProdCodeBench, is described as evaluating AI coding agents on tasks derived from production code changes, with a focus on identifying self-consistency checks where agents may reproduce previously generated solutions rather than solving novel tasks.

Sources: ProdCodeBench paper on arXiv

Coding Agent Index 2026 claims to be the first independent benchmark of full agent stacks.

Discovered 2026-09-14 · publication date unknown · discovery · confidence: secondary

A Medium article describes the 'Coding Agent Index 2026' as an independent benchmark evaluating full agent stacks (model + harness), aiming to help developers evaluate stacks before purchasing.

Sources: Coding Agent Index 2026 article on Medium

Blog post cites METR randomized trial finding AI-assisted developers 19% slower.

Discovered 2026-09-14 · publication date unknown · study · confidence: secondary

A blog post from Cerbos references the METR randomized trial from July 2025, which found experienced open-source developers using AI tools (like Cursor Pro with Claude) were on average 19% slower, yet believed they were faster.

Sources: Cerbos blog post on AI productivity paradox

GitHub enables multi-agent AI coding inside repository workflows.

Discovered 2026-09-14 · publication date unknown · update · confidence: secondary · facet: orchestration

A news article reports that GitHub has enabled Claude and Codex agents within repository workflows, allowing contributors to submit requests and select multiple agents to execute tasks in parallel, with context attached to the work.

Sources: Help Net Security article on GitHub multi-agent coding

Repo-of-Repos pattern described for multi-repo workspace with AI coding agents.

Discovered 2026-09-14 · publication date unknown · discovery · confidence: secondary

A blog post details the 'Repo-of-Repos' pattern, a template that uses an outer 'agent' repo to pull in related repos as workspace folders, facilitating multi-repo work for AI coding agents while maintaining separate commit histories.

Sources: Rafferty Uy blog post on Repo-of-Repos pattern

Scale AI publishes public dataset and leaderboard for SWE-bench Pro.

Discovered 2026-09-14 · publication date unknown · update · confidence: primary-source

Scale AI has launched a public dataset and leaderboard for SWE-bench Pro, describing it as a benchmark designed to address limitations in existing benchmarks with 1865 tasks across 41 professional repositories.

Sources: Scale AI SWE-bench Pro leaderboard

Moderne publishes documentation for its multi-repo AI agent 'Moddy'.

Discovered 2026-09-14 · publication date unknown · discovery · confidence: vendor-claim · facet: orchestration

Moderne has published user documentation for 'Moddy', describing it as an AI agent that modernizes multi-repository codebases by combining LLMs with structured LST code data.

Sources: Moderne Docs for Moddy

JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks.

Discovered 2026-09-07 · publication date unknown · launch · confidence: vendor-claim

JetBrains released an official benchmark designed to assess how different AI coding agents perform on Kotlin software engineering tasks, aiming to provide a credible, public evaluation closer to day-to-day development work.

Sources: JetBrains Blog

OpenHands Enterprise platform enables governed AI agent benchmarking in private cloud.

Discovered 2026-09-07 · publication date unknown · update · confidence: vendor-claim

OpenHands Enterprise offers a platform for running AI coding agent benchmarks in a customer's own virtual private cloud with audit logs and cost attribution, positioning it for internal, governed evaluation at scale.

Sources: OpenHands Blog

Google RCT estimates impact of three AI features on developer time for complex tasks.

Discovered 2026-09-07 · publication date unknown · study · confidence: primary-source

A randomized controlled trial with 96 full-time Google software engineers provides an estimate of how three AI features affect the time developers spend on a complex, enterprise-grade task.

Sources: arXiv

Randomized controlled experiment measures GitHub Copilot's effect on developer productivity.

Discovered 2026-09-07 · publication date unknown · study · confidence: primary-source

A study employing an RCT with 95 professional programmers recruited via Upwork measures the effect of GitHub Copilot on developer productivity, contributing to the empirical evidence base.

Sources: alphaXiv

Submodule-based multi-repo workspace strategy proposed to avoid file conflicts with concurrent AI agents.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary · facet: orchestration

A Medium article outlines a workspace strategy using Git submodules to allow multiple AI coding agents to work simultaneously on different service repos within a larger project, structurally preventing file conflicts.

Sources: Medium

GitHub acknowledges multi-repo agent capabilities are on the roadmap, recommends task decomposition for now.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: primary-source · facet: orchestration

A GitHub community discussion reveals that background agents currently work best within a single repo context, with multi-repo capabilities on the roadmap; the recommended current approach is to break work into repo-specific sub-tasks.

Sources: GitHub Discussion

Independent comparison benchmarks Claude Code, Cursor, and Aider on different task types.

Discovered 2026-09-07 · publication date unknown · study · confidence: secondary

A benchmark analysis finds Claude Code excels at complex, multi-file tasks like refactoring and migrations, Cursor wins on quick, focused tasks like bug fixes, and Aider lands in between on speed and thoroughness.

Sources: MorphLLM

Analysis highlights high operational costs of autonomous agents and need for LLM gateway routing.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary

A tool comparison notes that autonomous coding agents can be expensive, with sessions making 50-200+ LLM calls, and emphasizes the critical role of LLM gateway routing for cost and performance management.

Sources: Requesty AI Blog

Analysis illustrates large performance gap between SWE-bench Verified and Pro, highlighting contamination.

Discovered 2026-09-07 · publication date unknown · study · confidence: secondary

A blog post notes Claude Opus 4.5 scored 80.9% on SWE-bench Verified but 45.89% on SWE-bench Pro, a 35-point collapse attributed to the contamination-free design of the latter, which uses GPL-licensed and private repos.

Sources: Paddo.dev

Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary · facet: local

A Medium article based on lessons learned concludes that smaller models work for chat-based assistance, but true autonomous agents require 20B parameter models minimum, with 32B+ being significantly better, creating a tension with VRAM constraints.

Sources: Medium

Guide outlines best practices for pair programming with AI assistants.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: vendor-claim · facet: pairing

A Graphite guide details best practices for leveraging AI coding agents in a pair programming context, covering roles, context provision, code review, and using AI as both a productivity and learning tool.

Sources: Graphite Guides

ForgeCode publishes 12 practical lessons from six months of daily AI pair programming.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: vendor-claim · facet: pairing

A blog post distills lessons from extensive daily AI pair programming, covering planning, prompt engineering, context management, and noting what doesn't work, based on practical experience across multiple codebases.

Sources: ForgeCode Blog

Multiple sources document concepts and architectures for 'agent harness engineering'.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary

Articles from Martin Fowler, Microsoft, Addy Osmani, and others explore the concept of 'harness engineering'—the scaffolding around a coding agent—including runtime design, self-evolution loops, and observability-driven automation.

Sources: Martin Fowler · Microsoft Learn · Addy Osmani

Claude Code docs detail updates for week 20 of 2026, including default fast model change and agent view.

Discovered 2026-09-07 · publication date unknown · update · confidence: primary-source

Claude Code documentation for week 20 (May 11–15, 2026) notes that /fast now defaults to Opus 4.7, and the claude agents command provides a unified view of running sessions and their status.

Sources: Claude Code Docs

OpenAI announces Codex integration in ChatGPT mobile app for working from anywhere.

Discovered 2026-09-07 · publication date unknown · update · confidence: vendor-claim

OpenAI announced a preview of Codex within the ChatGPT mobile app, allowing users to connect to Codex running on their machines (laptop, Mac mini, remote env) and work fluidly across active threads and project context from iOS or Android.

Sources: OpenAI Blog

The last 30 days

The evidence layer thickened this month, with two new randomized controlled trials (RCTs) and a major benchmark shift. Anthropic published an RCT (unverified) investigating whether AI assistance harms skill formation when learning a new library. More critically, the METR trial finding—that experienced developers were ~19% slower with AI—was cited again, reinforcing a negative result that vendors have not contradicted with independent data. In benchmarks, the contamination problem became unavoidable: OpenAI stated it no longer evaluates on SWE-bench Verified, recommending SWE-bench Pro instead. An analysis showed a 35-point performance collapse between the two for Claude Opus 4.5, illustrating how inflated vendor scores on contaminated benchmarks have been.

On the tool front, the multi-agent and multi-repo coordination problem saw practical, if incremental, movement. GitHub enabled multi-agent workflows inside repositories (vendor-claim), and Moderne published docs for its multi-repo agent 'Moddy'. The recommended pattern for current tools remains breaking work into repo-specific sub-tasks, as true multi-repo capabilities are still on roadmaps. In a notable market signal, a Hacker News comment questioned if anyone is still using Cursor in 2026, noting a shift to Claude Code, Codex, or Copilot. Aider's development appears stable, while Claude Code advanced with Opus 5 integration.

For a reader deciding where to spend attention, the priority should be scrutinizing any performance claim. Trust only benchmarks that address contamination, like SWE-bench Pro, and understand that Pass@k scores (used by vendors) are often 15-25 points higher than the more reliable Pass^k metric. The operational cost of autonomous agents remains high, requiring careful LLM gateway management. If you're evaluating tools for internal use, platforms like OpenHands Enterprise now offer governed benchmarking in a private cloud, which is a more credible approach than public vendor leaderboards. The most concrete finding is that the evidence continues to suggest AI assistance does not universally speed up experienced developers, despite their belief that it does.

Written 2026-09-21 from the changelog below, not from a fresh search.

Drawn from:

The last year

The last twelve months have been defined by a growing rift between vendor claims and independent evidence. The most significant finding is the METR randomized trial, which reported that experienced open-source developers were roughly 19% slower when using AI assistance, directly contradicting widespread productivity promises. This was joined by a Google RCT estimating AI's impact on complex tasks and another study measuring GitHub Copilot's effect, all pointing to a nuanced, often underwhelming reality. Meanwhile, vendors continued to launch new benchmarks, but the credibility of these metrics collapsed under scrutiny. OpenAI stopped evaluating on SWE-bench Verified due to contamination, and independent analyses highlighted a 15-25 point gap between the misleading Pass@k scores vendors report and the more rigorous Pass^k metric. JetBrains launched a Kotlin-specific benchmark, and OpenHands Enterprise offered a platform for private benchmarking, but the overarching lesson is that vendor-run benchmarks are, as one critic put it, as trustworthy as a pharmaceutical company grading its own drug trial.

The tools themselves saw incremental updates rather than breakthroughs. Claude Code, Cursor, and Aider solidified as the main contenders for interactive pair programming, with independent comparisons noting Claude Code excels at complex refactoring while Cursor wins on quick bug fixes. However, a consistent theme was high operational cost and technical constraint: autonomous agents require 20B+ parameter models, creating VRAM tension, and a single session can make 50-200+ LLM calls. VS Code integrated Claude and Codex agents under the Copilot umbrella, and OpenAI launched a Codex app for mobile, extending the workspace but not fundamentally changing the agent's capabilities. The architectural conversation shifted to "harness engineering"—the scaffolding needed to make agents reliable—with multiple sources, including OpenAI, detailing concepts for runtime design and observability.

Where the year delivered concrete progress was in documenting the real, messy failure modes of using these tools at scale, which is more valuable than any benchmark score. Multiple guides emerged from practitioners running concurrent agents, detailing how they built duplicate pipelines or caused file conflicts. In response, the community proposed structural solutions like Git submodules for multi-repo workspaces, while GitHub acknowledged multi-repo agent capabilities are still on the roadmap. Moderne launched 'Moddy', a multi-repo agent for large-scale transformation, and the open-source tool 'The Pair' offered a multi-agent architecture with cross-checking. For anyone deciding where to spend attention, these practical reports on coordination, collision, and cost management are the most load-bearing information available—they map the gap between the demo and the daily grind. The evidence layer now clearly shows that the biggest gains won't come from a faster model, but from better orchestration of the ones we have.

Written 2026-09-07 from the changelog below, not from a fresh search.

Month Entries
September 2026 29
August 2026 16

The ones that mattered:

All time

The evidence layer is now the most important part of the field. The METR randomized trial—a genuine controlled experiment—found that experienced open-source developers were 19% slower when using AI assistance, despite believing they were faster. This directly contradicts the productivity claims of every vendor and should be the starting point for any evaluation. If you are considering these tools for a team, your default assumption must now be that they slow down experienced developers; the burden of proof is on any claim to the contrary.

Benchmarks are in crisis, which makes the evidence problem worse. SWE-bench is widely criticized for potential training data contamination, meaning high scores may reflect memorization, not problem-solving. A new benchmark, ProjDevBench, attempts to measure end-to-end project development. In its initial evaluation, Codex+GPT-5 scored highest at 77.85%, but performance varied wildly by task. Treat all benchmark results as unverified and vendor-adjacent until independent replications are done.

The tool landscape is consolidating into platform plays. Visual Studio Code now supports Claude and Codex agents under a GitHub Copilot subscription, bringing multiple “agentic” models into a single editor interface. The open-source counterpoint is tools like The Pair, which uses a multi-agent Mentor+Executor architecture to cross-check code for hallucinations. The actual execution model—how an agent holds context, plans, and verifies work—still varies drastically between tools, but the market is coalescing around IDEs as the orchestration layer.

The major operational finding this period is the concrete failure mode of multi-agent workspaces. Running multiple AI coding sessions on a single repository, as done in this estate, leads to coordination failures: agents independently build duplicate systems without awareness, both reporting success. An AugmentCode guide documents this, and OpenAI’s article on “harness engineering” implicitly acknowledges the need for serious infrastructure to manage agents. The current tools do not solve this; they create the problem. If you move beyond a single human/agent pair, you are entering uncharted territory with real collision and drift risks.

What does not work is assuming AI coding agents universally increase productivity—the best available evidence says the opposite for experienced devs. Also failing is relying on existing benchmarks to gauge real capability. What has changed is the hardening of these negative findings and the emergence of multi-agent chaos as a practical, not theoretical, problem. The open gaps are massive: we have no proven model for agent coordination, no uncontaminated benchmarks, and no replicated evidence that these tools improve outcomes for seasoned developers. The field is currently defined by its failures and uncertainties.

Written 2026-08-18 from the changelog below, not from a fresh search.

How the field breaks down, by what we're actually tracking:

Tracked products

Nothing on this beat is buyable right now, at least nothing we've verified. Everything currently tracked sits in the forward-looking sections below.

Upcoming

Announced, no date.

Nothing in this bucket right now.

Coming Soon

Announced with a date, or an open pre-order.

Nothing in this bucket right now.

What We're Watching

Exists, unproven, or newly discovered. This is where auto-discovered items land.

ProjDevBench

Benchmark for evaluating AI coding agents on end-to-end project development tasks.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: ProjDevBench benchmark evaluates six coding agents on end-to-end project tasks. (2026-08-18) — The benchmark, detailed in an arXiv paper, assesses agents on tasks like from-scratch construction, finding performance varies significantly across models and tasks, with Codex+GPT-5 achieving the best overall score of 77.85%.

Sources: arXiv paper

METR

Organization conducting empirical research on AI safety and impact, including developer productivity studies.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: Blog post cites METR randomized trial finding AI-assisted developers 19% slower. (2026-09-14) — A blog post from Cerbos references the METR randomized trial from July 2025, which found experienced open-source developers using AI tools (like Cursor Pro with Claude) were on average 19% slower, yet believed they were faster.

Sources: METR blog

Visual Studio Code

Code editor with integrated AI coding agent support via extensions and GitHub Copilot.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: GitHub enables multi-agent AI coding inside repository workflows. (2026-09-14) — A news article reports that GitHub has enabled Claude and Codex agents within repository workflows, allowing contributors to submit requests and select multiple agents to execute tasks in parallel, with context attached to the work.

Sources: VS Code blog

Claude Code

AI coding agent by Anthropic, available as a CLI and integrated into various IDEs.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. (2026-09-21) — A 2026 update notes Aider remains at version 0.86.2 with a stable maintenance cadence, while Claude Code has advanced past v2.1.200 with auto-mode support and Opus 5 model integration, narrowing the model gap.

Sources: OpenHands comparison blog

Codex

AI coding agent by OpenAI, often integrated into tools like GitHub Copilot and VS Code.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center. (2026-09-21) — OpenAI launched a standalone desktop app for its Codex AI coding assistant, temporarily available to ChatGPT users with Apple computers, designed as a 'command center' to manage multiple AI agents.

Sources: OpenAI harness engineering post

SWE-bench

Benchmark for evaluating AI models on real-world software engineering issues from GitHub.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. (2026-09-21) — Bito announced evaluation results showing a Claude Sonnet 4.5 agent augmented with its AI Architect context engine achieved a 60.8% success rate on the SWE-Bench Pro benchmark.

Sources: CodeSOTA guide

The Pair

Open-source desktop app for AI pair programming using a multi-agent cross-checking architecture.

First seen 2026-08-18 · confidence unverified · uncategorised

Why it's here: Open-source tool 'The Pair' provides multi-agent coding with cross-checking. (2026-08-18) — A GitHub repository describes The Pair, an open-source desktop app that uses a Mentor+Executor agent architecture to cross-check code and catch hallucinations, compatible with various AI models.

Sources: GitHub repository

Moderne

Company offering a multi-repo AI agent ('Moddy') for large-scale codebase transformation and maintenance.

First seen 2026-08-24 · confidence unverified · uncategorised

Why it's here: Moderne publishes documentation for its multi-repo AI agent 'Moddy'. (2026-09-14) — Moderne has published user documentation for 'Moddy', describing it as an AI agent that modernizes multi-repository codebases by combining LLMs with structured LST code data.

Sources: Moderne blog

Cursor

AI-powered IDE with integrated coding agents, competing with VS Code and Windsurf.

First seen 2026-08-24 · confidence unverified · uncategorised

Why it's here: Hacker News discussion questions if anyone is still using Cursor in 2026. (2026-09-21) — A Hacker News comment from September 2026 asks if anyone is still actually using Cursor, noting that people they've spoken with are now using Claude Code, Codex, or Copilot.

Sources: Cursor vs. Claude Code comparison

Aider

Open-source, terminal-first AI pair programming agent with deep git integration and multi-model support.

First seen 2026-08-24 · confidence unverified · uncategorised

Why it's here: Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. (2026-09-21) — A 2026 update notes Aider remains at version 0.86.2 with a stable maintenance cadence, while Claude Code has advanced past v2.1.200 with auto-mode support and Opus 5 model integration, narrowing the model gap.

Sources: Aider skill page on LobeHub

Kotlin Benchmark

JetBrains' official benchmark for evaluating AI coding agents on real-world Kotlin software engineering tasks.

First seen 2026-09-07 · confidence unverified · uncategorised

Why it's here: Kotlin Benchmark leaderboard updated with new agent results. (2026-09-21) — The official Kotlin Benchmark leaderboard shows updated results for various AI coding agents, with Claude Code + Opus 4.6 medium achieving 71.43% resolution and Codex + GPT 5.3 Codex medium achieving 68.57%.

Sources: JetBrains Blog Announcement

OpenHands Enterprise

Enterprise platform for running governed AI coding agent benchmarks in a private cloud with audit logs and cost attribution.

First seen 2026-09-07 · confidence unverified · uncategorised

Why it's here: OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. (2026-09-21) — OpenHands Enterprise introduces an Agent Control Plane, providing a fully self-hosted, on-premise system for running, controlling, observing, and scaling AI agents across an organization with enforced policies and audit logs.

Sources: OpenHands Blog

ProdCodeBench

Benchmark for evaluating AI coding agents on tasks derived from production code changes.

First seen 2026-09-14 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: ProdCodeBench paper on arXiv

Coding Agent Index

Independent benchmark claiming to evaluate full agent stacks (model + harness).

First seen 2026-09-14 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: Coding Agent Index 2026 article on Medium

AugmentCode

Provider of guides and resources on AI coding practices, including multi-agent coordination.

First seen 2026-09-14 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: AugmentCode guide on multi-agent coding workspace

Scale AI

Company providing AI data and evaluation platforms, including the SWE-bench Pro public dataset.

First seen 2026-09-14 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: Scale AI SWE-bench Pro leaderboard

BenchLM

Platform providing AI model benchmark leaderboards and evaluations, including for SWE-bench Pro.

First seen 2026-09-14 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: BenchLM SWE-bench Pro leaderboard

Bito

Company building deep context graphs for coding agents, with an AI Architect context engine.

First seen 2026-09-21 · confidence unverified · uncategorised

Why it's here: No changelog entry explains this status yet.

Sources: Company mention (PR)

Promising

Early-stage, but the evidence or the approach is genuinely interesting. The only editorial bucket on this page — an item only lands here with a reason recorded in the changelog.

Nothing in this bucket right now.

Open questions

Publishing what we don't know is the point. These are things the job is actively watching for; when one gets answered it becomes a changelog entry and moves down here to the answered list.

No open questions recorded. Either everything we wanted to know got answered, or nobody has written down what we don't know — the second is more likely on a young tracker.

Full changelog

Everything, newest first, grouped by the month we found it. Long by design — it's the receipts.

September 2026

Kotlin Benchmark leaderboard updated with new agent results.

Discovered 2026-09-21 · publication date unknown · update · confidence: primary-source

The official Kotlin Benchmark leaderboard shows updated results for various AI coding agents, with Claude Code + Opus 4.6 medium achieving 71.43% resolution and Codex + GPT 5.3 Codex medium achieving 68.57%.

Sources: Kotlin Benchmark page

Anthropic publishes RCT on AI assistance's impact on coding skill formation.

Discovered 2026-09-21 · publication date unknown · study · confidence: primary-source

Anthropic conducted a randomized controlled trial investigating whether AI assistance prevents software developers from growing their skills or understanding the systems they're building when learning a new Python library.

Sources: Anthropic research

Hacker News discussion questions if anyone is still using Cursor in 2026.

Discovered 2026-09-21 · publication date unknown · negative · confidence: secondary

A Hacker News comment from September 2026 asks if anyone is still actually using Cursor, noting that people they've spoken with are now using Claude Code, Codex, or Copilot.

Sources: Hacker News comment

Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration.

Discovered 2026-09-21 · publication date unknown · update · confidence: secondary

A 2026 update notes Aider remains at version 0.86.2 with a stable maintenance cadence, while Claude Code has advanced past v2.1.200 with auto-mode support and Opus 5 model integration, narrowing the model gap.

Sources: Developers Digest blog

OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management.

Discovered 2026-09-21 · published 2026-05-06 · launch · confidence: vendor-claim

OpenHands Enterprise introduces an Agent Control Plane, providing a fully self-hosted, on-premise system for running, controlling, observing, and scaling AI agents across an organization with enforced policies and audit logs.

Sources: OpenHands blog

Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5.

Discovered 2026-09-21 · published 2026-02-03 · study · confidence: vendor-claim

Bito announced evaluation results showing a Claude Sonnet 4.5 agent augmented with its AI Architect context engine achieved a 60.8% success rate on the SWE-Bench Pro benchmark.

Sources: Yahoo Finance (PRNewswire)

OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center.

Discovered 2026-09-21 · published 2026-02-02 · launch · confidence: secondary

OpenAI launched a standalone desktop app for its Codex AI coding assistant, temporarily available to ChatGPT users with Apple computers, designed as a 'command center' to manage multiple AI agents.

Sources: CNBC article

ProdCodeBench benchmark introduced to evaluate AI coding agents on production-derived tasks.

Discovered 2026-09-14 · publication date unknown · discovery · confidence: primary-source

A new benchmark, ProdCodeBench, is described as evaluating AI coding agents on tasks derived from production code changes, with a focus on identifying self-consistency checks where agents may reproduce previously generated solutions rather than solving novel tasks.

Sources: ProdCodeBench paper on arXiv

Coding Agent Index 2026 claims to be the first independent benchmark of full agent stacks.

Discovered 2026-09-14 · publication date unknown · discovery · confidence: secondary

A Medium article describes the 'Coding Agent Index 2026' as an independent benchmark evaluating full agent stacks (model + harness), aiming to help developers evaluate stacks before purchasing.

Sources: Coding Agent Index 2026 article on Medium

Blog post cites METR randomized trial finding AI-assisted developers 19% slower.

Discovered 2026-09-14 · publication date unknown · study · confidence: secondary

A blog post from Cerbos references the METR randomized trial from July 2025, which found experienced open-source developers using AI tools (like Cursor Pro with Claude) were on average 19% slower, yet believed they were faster.

Sources: Cerbos blog post on AI productivity paradox

GitHub enables multi-agent AI coding inside repository workflows.

Discovered 2026-09-14 · publication date unknown · update · confidence: secondary · facet: orchestration

A news article reports that GitHub has enabled Claude and Codex agents within repository workflows, allowing contributors to submit requests and select multiple agents to execute tasks in parallel, with context attached to the work.

Sources: Help Net Security article on GitHub multi-agent coding

Repo-of-Repos pattern described for multi-repo workspace with AI coding agents.

Discovered 2026-09-14 · publication date unknown · discovery · confidence: secondary

A blog post details the 'Repo-of-Repos' pattern, a template that uses an outer 'agent' repo to pull in related repos as workspace folders, facilitating multi-repo work for AI coding agents while maintaining separate commit histories.

Sources: Rafferty Uy blog post on Repo-of-Repos pattern

Scale AI publishes public dataset and leaderboard for SWE-bench Pro.

Discovered 2026-09-14 · publication date unknown · update · confidence: primary-source

Scale AI has launched a public dataset and leaderboard for SWE-bench Pro, describing it as a benchmark designed to address limitations in existing benchmarks with 1865 tasks across 41 professional repositories.

Sources: Scale AI SWE-bench Pro leaderboard

Moderne publishes documentation for its multi-repo AI agent 'Moddy'.

Discovered 2026-09-14 · publication date unknown · discovery · confidence: vendor-claim · facet: orchestration

Moderne has published user documentation for 'Moddy', describing it as an AI agent that modernizes multi-repository codebases by combining LLMs with structured LST code data.

Sources: Moderne Docs for Moddy

JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks.

Discovered 2026-09-07 · publication date unknown · launch · confidence: vendor-claim

JetBrains released an official benchmark designed to assess how different AI coding agents perform on Kotlin software engineering tasks, aiming to provide a credible, public evaluation closer to day-to-day development work.

Sources: JetBrains Blog

OpenHands Enterprise platform enables governed AI agent benchmarking in private cloud.

Discovered 2026-09-07 · publication date unknown · update · confidence: vendor-claim

OpenHands Enterprise offers a platform for running AI coding agent benchmarks in a customer's own virtual private cloud with audit logs and cost attribution, positioning it for internal, governed evaluation at scale.

Sources: OpenHands Blog

Google RCT estimates impact of three AI features on developer time for complex tasks.

Discovered 2026-09-07 · publication date unknown · study · confidence: primary-source

A randomized controlled trial with 96 full-time Google software engineers provides an estimate of how three AI features affect the time developers spend on a complex, enterprise-grade task.

Sources: arXiv

Randomized controlled experiment measures GitHub Copilot's effect on developer productivity.

Discovered 2026-09-07 · publication date unknown · study · confidence: primary-source

A study employing an RCT with 95 professional programmers recruited via Upwork measures the effect of GitHub Copilot on developer productivity, contributing to the empirical evidence base.

Sources: alphaXiv

Submodule-based multi-repo workspace strategy proposed to avoid file conflicts with concurrent AI agents.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary · facet: orchestration

A Medium article outlines a workspace strategy using Git submodules to allow multiple AI coding agents to work simultaneously on different service repos within a larger project, structurally preventing file conflicts.

Sources: Medium

GitHub acknowledges multi-repo agent capabilities are on the roadmap, recommends task decomposition for now.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: primary-source · facet: orchestration

A GitHub community discussion reveals that background agents currently work best within a single repo context, with multi-repo capabilities on the roadmap; the recommended current approach is to break work into repo-specific sub-tasks.

Sources: GitHub Discussion

Independent comparison benchmarks Claude Code, Cursor, and Aider on different task types.

Discovered 2026-09-07 · publication date unknown · study · confidence: secondary

A benchmark analysis finds Claude Code excels at complex, multi-file tasks like refactoring and migrations, Cursor wins on quick, focused tasks like bug fixes, and Aider lands in between on speed and thoroughness.

Sources: MorphLLM

Analysis highlights high operational costs of autonomous agents and need for LLM gateway routing.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary

A tool comparison notes that autonomous coding agents can be expensive, with sessions making 50-200+ LLM calls, and emphasizes the critical role of LLM gateway routing for cost and performance management.

Sources: Requesty AI Blog

Analysis illustrates large performance gap between SWE-bench Verified and Pro, highlighting contamination.

Discovered 2026-09-07 · publication date unknown · study · confidence: secondary

A blog post notes Claude Opus 4.5 scored 80.9% on SWE-bench Verified but 45.89% on SWE-bench Pro, a 35-point collapse attributed to the contamination-free design of the latter, which uses GPL-licensed and private repos.

Sources: Paddo.dev

Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary · facet: local

A Medium article based on lessons learned concludes that smaller models work for chat-based assistance, but true autonomous agents require 20B parameter models minimum, with 32B+ being significantly better, creating a tension with VRAM constraints.

Sources: Medium

Guide outlines best practices for pair programming with AI assistants.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: vendor-claim · facet: pairing

A Graphite guide details best practices for leveraging AI coding agents in a pair programming context, covering roles, context provision, code review, and using AI as both a productivity and learning tool.

Sources: Graphite Guides

ForgeCode publishes 12 practical lessons from six months of daily AI pair programming.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: vendor-claim · facet: pairing

A blog post distills lessons from extensive daily AI pair programming, covering planning, prompt engineering, context management, and noting what doesn't work, based on practical experience across multiple codebases.

Sources: ForgeCode Blog

Multiple sources document concepts and architectures for 'agent harness engineering'.

Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary

Articles from Martin Fowler, Microsoft, Addy Osmani, and others explore the concept of 'harness engineering'—the scaffolding around a coding agent—including runtime design, self-evolution loops, and observability-driven automation.

Sources: Martin Fowler · Microsoft Learn · Addy Osmani

Claude Code docs detail updates for week 20 of 2026, including default fast model change and agent view.

Discovered 2026-09-07 · publication date unknown · update · confidence: primary-source

Claude Code documentation for week 20 (May 11–15, 2026) notes that /fast now defaults to Opus 4.7, and the claude agents command provides a unified view of running sessions and their status.

Sources: Claude Code Docs

OpenAI announces Codex integration in ChatGPT mobile app for working from anywhere.

Discovered 2026-09-07 · publication date unknown · update · confidence: vendor-claim

OpenAI announced a preview of Codex within the ChatGPT mobile app, allowing users to connect to Codex running on their machines (laptop, Mac mini, remote env) and work fluidly across active threads and project context from iOS or Android.

Sources: OpenAI Blog

August 2026

Article highlights gap between Pass@k and Pass^k metrics, inflating vendor-reported scores.

Discovered 2026-08-31 · publication date unknown · study · confidence: secondary · facet: autonomous

A SoftwareSeni article explains that Pass@k measures success across k attempts (rewarding luck), while Pass^k requires success on every attempt (measuring reliability), with a typical gap of 15-25 percentage points that makes vendor scores misleading.

Sources: SoftwareSeni

OpenAI announces it no longer evaluates on SWE-bench Verified, citing contamination issues.

Discovered 2026-08-31 · publication date unknown · negative · confidence: primary-source · facet: autonomous

OpenAI states it has stopped reporting results on SWE-bench Verified due to contamination concerns and now recommends SWE-bench Pro, which empirically suffers less from contamination, though it is not perfect.

Sources: OpenAI blog

VS Code 1.134 adds features for organizing chats across windows and navigating long conversations.

Discovered 2026-08-24 · published 2026-08-19 · update · confidence: primary-source

The release, dated August 19, 2026, introduces capabilities to work across windows, organize related chats side by side, and navigate long conversations faster, as per the official VS Code update notes.

Sources: Visual Studio Code 1.134 release notes

Claude Code promotional weekly usage limits are ending, reverting to previous levels.

Discovered 2026-08-24 · publication date unknown · update · confidence: secondary

A Hacker News discussion notes that a promotion offering 50% higher weekly usage limits from May 13 to August 19, 2026, is ending, with limits reverting to pre-promotion levels.

Sources: Hacker News discussion on Claude Code limits

Independent analysis criticizes vendor-run AI code review benchmarks as biased and lacking trustworthy methodology.

Discovered 2026-08-24 · publication date unknown · study · confidence: secondary

A DeepSource blog post argues that self-evaluation by vendors is inherently biased, drawing parallels to pharmaceutical clinical trials, and calls for independent evaluation, published datasets, and reproducible methodology.

Sources: DeepSource blog on AI code review benchmarks

Google randomized controlled trial estimates the impact of three AI features on developer time for complex tasks.

Discovered 2026-08-24 · publication date unknown · study · confidence: primary-source

An arXiv paper details a randomized controlled trial with 96 full-time Google software engineers, contributing an estimate of how AI features affect the time developers spend on a complex, enterprise-grade task.

Sources: arXiv paper on AI impact on developer productivity

Moderne launches 'Moddy', a multi-repo AI agent for transforming enterprise codebases at scale.

Discovered 2026-08-24 · publication date unknown · launch · confidence: vendor-claim

Moderne introduces an AI agent designed to handle transformations across entire codebases simultaneously, contrasting with tools focused on single repositories.

Sources: Moderne blog introducing multi-repo AI agent

OpenAI's Codex app, initially for macOS, is now available on Windows.

Discovered 2026-08-24 · published 2026-03-04 · update · confidence: primary-source

An update to OpenAI's February 2026 announcement states the Codex app, an interface for managing multiple agents and parallel work, is now available on Windows.

Sources: OpenAI blog introducing the Codex app

SWE-bench Pro leaderboard for August 2026 shows Claude Mythos 5 leading with 80.3%.

Discovered 2026-08-24 · publication date unknown · update · confidence: secondary

A BenchLM.ai leaderboard update lists Claude Mythos 5 ahead of Claude Fable 5 and Claude Opus 5 on the SWE-bench Pro benchmark, which is designed for long-horizon, realistic software engineering work.

Sources: BenchLM.ai SWE-bench Pro leaderboard

ProjDevBench benchmark evaluates six coding agents on end-to-end project tasks.

Discovered 2026-08-18 · publication date unknown · study · confidence: primary-source

The benchmark, detailed in an arXiv paper, assesses agents on tasks like from-scratch construction, finding performance varies significantly across models and tasks, with Codex+GPT-5 achieving the best overall score of 77.85%.

Sources: arXiv ProjDevBench paper

METR randomized trial finds AI assistance made experienced open-source developers 19% slower.

Discovered 2026-08-18 · publication date unknown · study · confidence: primary-source

A METR blog post reports a randomized controlled trial where developers using AI tools took longer on tasks than those without, contradicting common productivity claims.

Sources: METR blog post

Guide documents failure modes of running multiple AI coding agents on one repository.

Discovered 2026-08-18 · publication date unknown · discovery · confidence: secondary

An AugmentCode guide outlines coordination challenges and failure modes when multiple agents work concurrently on the same codebase without proper orchestration.

Sources: AugmentCode guide

VS Code adds support for Claude and Codex agents under GitHub Copilot subscription.

Discovered 2026-08-18 · publication date unknown · update · confidence: primary-source

A VS Code blog post announces that version 1.109 allows running Claude and Codex agents locally or in the cloud as part of Copilot Pro+ and Enterprise subscriptions.

Sources: VS Code blog

Criticism highlights contamination concerns in SWE-bench benchmark.

Discovered 2026-08-18 · publication date unknown · study · confidence: secondary

Multiple sources, including an arXiv paper and blog posts, discuss how training data contamination may inflate model performance on SWE-bench, questioning what the benchmark actually measures.

Sources: arXiv paper on SWE-bench illusion · Latent Space article

Open-source tool 'The Pair' provides multi-agent coding with cross-checking.

Discovered 2026-08-18 · publication date unknown · discovery · confidence: primary-source

A GitHub repository describes The Pair, an open-source desktop app that uses a Mentor+Executor agent architecture to cross-check code and catch hallucinations, compatible with various AI models.

Sources: GitHub repository for The Pair

OpenAI publishes article on 'harness engineering' for coding agents.

Discovered 2026-08-18 · publication date unknown · discovery · confidence: primary-source

An OpenAI blog post discusses the architectural constraints and infrastructure needed to effectively leverage coding agents like Codex in production, emphasizing early investment in harness design.

Sources: OpenAI blog post


Dreamlab Live Research uses autonomous systems to track the state of the art in fields of interest. New trackers appear as the interests do.

Subscribe / join for details — placeholder: there is nothing to sign up to yet.


Generated from a versioned dataset in a git repository: every state this page has ever been in is a commit, and a bad run is revertible. The dataset is the source; this page is build output and is not itself edited.