2026-08-18
AI Programming
A live tracker of how software actually gets written with machines — coding agents, interactive pairing, multi-agent orchestration, and whether any of it measurably helps.
Last updated 2026-09-21 · updated weekly by an automated research job · run #6.
The verdicts
This topic judges "best" along 4 axes, each with its own standing verdict and history.
Autonomous
Judged on: Best evidenced agent for multi-step work handed off with a goal rather than a diff — plans, edits many files, runs its own verification. Judged on independent task benchmarks and reported real-world completion, not vendor demos.
The current state of the art for Autonomous, as of 2026-08-18 (confidence primary-source):
Based on the ProjDevBench benchmark evaluating six coding agents, Codex+GPT-5 achieves the highest overall performance (77.85%) on end-to-end project tasks, though performance gaps widen on from-scratch construction tasks. However, independent empirical evidence from a METR randomized trial suggests AI assistance may actually slow down experienced developers in some contexts.
What triggered or contributed to this call:
- 2026-08-18 — ProjDevBench benchmark evaluates six coding agents on end-to-end project tasks. (
study,primary-source) - 2026-08-18 — METR randomized trial finds AI assistance made experienced open-source developers 19% slower. (
study,primary-source)
Contenders:
- Claude Code — Shows strong performance in multi-file tasks and refactoring according to comparisons.
- Codex — Top performer in the ProjDevBench benchmark when combined with GPT-5.
Pairing
Judged on: Best evidenced tool for real-time human-and-agent work in one session — interruptibility, shared context, how the human steers mid-run.
The current state of the art for Pairing, as of 2026-08-18 (confidence primary-source):
The open-source tool 'The Pair' introduces a multi-agent architecture with Mentor+Executor cross-checking to catch hallucinations, providing a concrete implementation for AI pair programming. VS Code's integration of Claude and Codex agents offers a widely accessible platform for real-time collaborative work.
What triggered or contributed to this call:
- 2026-08-18 — Open-source tool 'The Pair' provides multi-agent coding with cross-checking. (
discovery,primary-source) - 2026-08-18 — VS Code adds support for Claude and Codex agents under GitHub Copilot subscription. (
update,primary-source)
Contenders:
- The Pair — Open-source, cross-model support with built-in review agent for safety.
- Visual Studio Code — Broad IDE integration with popular agents under a Copilot subscription.
Orchestration
Judged on: Best evidenced way to run several agents on one codebase at once — coordination, isolation, conflict and drift handling.
The current state of the art for Orchestration, as of 2026-08-24 (confidence vendor-claim):
The launch of Moderne's 'Moddy' agent provides a vendor-claimed, concrete implementation for multi-repo, large-scale codebase orchestration, moving beyond single-repo coordination patterns. The previously noted guide from AugmentCode remains the primary independent documentation of failure modes for multi-agent work on a single repository.
What triggered or contributed to this call:
- 2026-08-24 — Moderne launches 'Moddy', a multi-repo AI agent for transforming enterprise codebases at scale. (
launch,vendor-claim)
Contenders:
- The Pair — Open-source implementation with a multi-agent, cross-checking architecture for single-repo work.
- Moderne — Claims to orchestrate transformations across multiple repositories at scale for enterprise codebases.
Dethronement history — Orchestration
Every verdict this facet has ever retired, and the receipts for retiring it.
Dethroned 2026-08-24 — held since 2026-08-18
A guide from AugmentCode documents specific failure modes when running multiple AI coding agents on one repository without coordination, highlighting the need for deliberate orchestration. The open-source tool 'The Pair' demonstrates one architectural approach using multiple agents with defined roles.
How it was beaten, and how we know the new one is better: The previous verdict highlighted a guide documenting failure modes and The Pair's single-repo architecture. The new item, Moderne's Moddy, is claimed to operate at a multi-repo, enterprise-scale level, representing a different and more complex orchestration target, though its effectiveness is not yet independently verified.
Triggering entries:
- 2026-08-18 — Guide documents failure modes of running multiple AI coding agents on one repository. (
discovery,secondary) - 2026-08-18 — Open-source tool 'The Pair' provides multi-agent coding with cross-checking. (
discovery,primary-source)
Local
Judged on: Best coding setup that runs entirely on hardware you own, judged on usable quality at a stated VRAM budget, not on leaderboard position.
The current state of the art for Local, as of 2026-09-07 (confidence secondary):
Effective autonomous coding agents require 20B+ parameter models, with 32B+ offering significantly better performance, creating a fundamental tension with VRAM constraints on local hardware. Guides discuss hardware requirements and model options, but no single leading local setup or 'best' model is identified from current evidence that balances high coding capability with typical consumer VRAM budgets.
What triggered or contributed to this call:
- 2026-09-07 — Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension. (
discovery,secondary)
Dethronement history — Local
Every verdict this facet has ever retired, and the receipts for retiring it.
Dethroned 2026-09-07 — held since 2026-08-18
Several guides discuss hardware requirements and model options for self-hosted AI coding, noting that effective autonomous agents typically require 20B+ parameter models, creating a tension with VRAM constraints. No single leading local setup is identified from the current results.
How it was beaten, and how we know the new one is better: The previous verdict stated that 'No single leading local setup is identified from the current results.' New evidence from a Medium article based on practical lessons provides a more concrete technical threshold, specifying that true autonomous agents require 20B parameter models minimum, which clarifies the VRAM-performance tension but does not yet identify a leading setup that resolves it.
What this is, and how it works
This page is generated, not written. A scheduled job runs weekly on a machine in a homelab. Each run it:
- Searches the open web for both the products already tracked here and for category-level terms designed to turn up ones we've never heard of.
- Feeds those results to a language model along with everything already on this page, and asks it what is genuinely new. Finding nothing is an acceptable answer, and most runs should find little.
- Writes the result into a versioned JSON dataset (
data/live-research/ai-programming.json) and commits it. This page is re-rendered from that file.
So: a machine wrote the prose here. Every changelog entry carries at least one source link, and every item and entry carries a confidence label:
primary-source— vendor documentation, a published paper, or a regulator.secondary— reputable press.vendor-claim— marketing or an unreplicated vendor statement.unverified— a single low-quality source, usually auto-discovered.
Those labels are always shown, including on the weak ones. A tracker that hides its own uncertainty is worse than no tracker.
What this is not. Not medical advice. Not a review site — nothing here has been tested in our hands. No affiliate relationships, no sponsored placements, nothing bought. Vendor claims are attributed to the vendor rather than restated as findings. Where a price, a date, or a regulatory status is unknown, it is left blank instead of guessed.
What's new
29 developments in the last 30 days, newest first.
Kotlin Benchmark leaderboard updated with new agent results.
Discovered 2026-09-21 · publication date unknown · update · confidence: primary-source
The official Kotlin Benchmark leaderboard shows updated results for various AI coding agents, with Claude Code + Opus 4.6 medium achieving 71.43% resolution and Codex + GPT 5.3 Codex medium achieving 68.57%.
Sources: Kotlin Benchmark page
Anthropic publishes RCT on AI assistance's impact on coding skill formation.
Discovered 2026-09-21 · publication date unknown · study · confidence: primary-source
Anthropic conducted a randomized controlled trial investigating whether AI assistance prevents software developers from growing their skills or understanding the systems they're building when learning a new Python library.
Sources: Anthropic research
Hacker News discussion questions if anyone is still using Cursor in 2026.
Discovered 2026-09-21 · publication date unknown · negative · confidence: secondary
A Hacker News comment from September 2026 asks if anyone is still actually using Cursor, noting that people they've spoken with are now using Claude Code, Codex, or Copilot.
Sources: Hacker News comment
Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration.
Discovered 2026-09-21 · publication date unknown · update · confidence: secondary
A 2026 update notes Aider remains at version 0.86.2 with a stable maintenance cadence, while Claude Code has advanced past v2.1.200 with auto-mode support and Opus 5 model integration, narrowing the model gap.
Sources: Developers Digest blog
OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management.
Discovered 2026-09-21 · published 2026-05-06 · launch · confidence: vendor-claim
OpenHands Enterprise introduces an Agent Control Plane, providing a fully self-hosted, on-premise system for running, controlling, observing, and scaling AI agents across an organization with enforced policies and audit logs.
Sources: OpenHands blog
Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5.
Discovered 2026-09-21 · published 2026-02-03 · study · confidence: vendor-claim
Bito announced evaluation results showing a Claude Sonnet 4.5 agent augmented with its AI Architect context engine achieved a 60.8% success rate on the SWE-Bench Pro benchmark.
Sources: Yahoo Finance (PRNewswire)
OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center.
Discovered 2026-09-21 · published 2026-02-02 · launch · confidence: secondary
OpenAI launched a standalone desktop app for its Codex AI coding assistant, temporarily available to ChatGPT users with Apple computers, designed as a 'command center' to manage multiple AI agents.
Sources: CNBC article
ProdCodeBench benchmark introduced to evaluate AI coding agents on production-derived tasks.
Discovered 2026-09-14 · publication date unknown · discovery · confidence: primary-source
A new benchmark, ProdCodeBench, is described as evaluating AI coding agents on tasks derived from production code changes, with a focus on identifying self-consistency checks where agents may reproduce previously generated solutions rather than solving novel tasks.
Sources: ProdCodeBench paper on arXiv
Coding Agent Index 2026 claims to be the first independent benchmark of full agent stacks.
Discovered 2026-09-14 · publication date unknown · discovery · confidence: secondary
A Medium article describes the 'Coding Agent Index 2026' as an independent benchmark evaluating full agent stacks (model + harness), aiming to help developers evaluate stacks before purchasing.
Sources: Coding Agent Index 2026 article on Medium
Blog post cites METR randomized trial finding AI-assisted developers 19% slower.
Discovered 2026-09-14 · publication date unknown · study · confidence: secondary
A blog post from Cerbos references the METR randomized trial from July 2025, which found experienced open-source developers using AI tools (like Cursor Pro with Claude) were on average 19% slower, yet believed they were faster.
Sources: Cerbos blog post on AI productivity paradox
GitHub enables multi-agent AI coding inside repository workflows.
Discovered 2026-09-14 · publication date unknown · update · confidence: secondary · facet: orchestration
A news article reports that GitHub has enabled Claude and Codex agents within repository workflows, allowing contributors to submit requests and select multiple agents to execute tasks in parallel, with context attached to the work.
Sources: Help Net Security article on GitHub multi-agent coding
Repo-of-Repos pattern described for multi-repo workspace with AI coding agents.
Discovered 2026-09-14 · publication date unknown · discovery · confidence: secondary
A blog post details the 'Repo-of-Repos' pattern, a template that uses an outer 'agent' repo to pull in related repos as workspace folders, facilitating multi-repo work for AI coding agents while maintaining separate commit histories.
Sources: Rafferty Uy blog post on Repo-of-Repos pattern
Scale AI publishes public dataset and leaderboard for SWE-bench Pro.
Discovered 2026-09-14 · publication date unknown · update · confidence: primary-source
Scale AI has launched a public dataset and leaderboard for SWE-bench Pro, describing it as a benchmark designed to address limitations in existing benchmarks with 1865 tasks across 41 professional repositories.
Sources: Scale AI SWE-bench Pro leaderboard
Moderne publishes documentation for its multi-repo AI agent 'Moddy'.
Discovered 2026-09-14 · publication date unknown · discovery · confidence: vendor-claim · facet: orchestration
Moderne has published user documentation for 'Moddy', describing it as an AI agent that modernizes multi-repository codebases by combining LLMs with structured LST code data.
Sources: Moderne Docs for Moddy
JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks.
Discovered 2026-09-07 · publication date unknown · launch · confidence: vendor-claim
JetBrains released an official benchmark designed to assess how different AI coding agents perform on Kotlin software engineering tasks, aiming to provide a credible, public evaluation closer to day-to-day development work.
Sources: JetBrains Blog
OpenHands Enterprise platform enables governed AI agent benchmarking in private cloud.
Discovered 2026-09-07 · publication date unknown · update · confidence: vendor-claim
OpenHands Enterprise offers a platform for running AI coding agent benchmarks in a customer's own virtual private cloud with audit logs and cost attribution, positioning it for internal, governed evaluation at scale.
Sources: OpenHands Blog
Google RCT estimates impact of three AI features on developer time for complex tasks.
Discovered 2026-09-07 · publication date unknown · study · confidence: primary-source
A randomized controlled trial with 96 full-time Google software engineers provides an estimate of how three AI features affect the time developers spend on a complex, enterprise-grade task.
Sources: arXiv
Randomized controlled experiment measures GitHub Copilot's effect on developer productivity.
Discovered 2026-09-07 · publication date unknown · study · confidence: primary-source
A study employing an RCT with 95 professional programmers recruited via Upwork measures the effect of GitHub Copilot on developer productivity, contributing to the empirical evidence base.
Sources: alphaXiv
Submodule-based multi-repo workspace strategy proposed to avoid file conflicts with concurrent AI agents.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary · facet: orchestration
A Medium article outlines a workspace strategy using Git submodules to allow multiple AI coding agents to work simultaneously on different service repos within a larger project, structurally preventing file conflicts.
Sources: Medium
GitHub acknowledges multi-repo agent capabilities are on the roadmap, recommends task decomposition for now.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: primary-source · facet: orchestration
A GitHub community discussion reveals that background agents currently work best within a single repo context, with multi-repo capabilities on the roadmap; the recommended current approach is to break work into repo-specific sub-tasks.
Sources: GitHub Discussion
Independent comparison benchmarks Claude Code, Cursor, and Aider on different task types.
Discovered 2026-09-07 · publication date unknown · study · confidence: secondary
A benchmark analysis finds Claude Code excels at complex, multi-file tasks like refactoring and migrations, Cursor wins on quick, focused tasks like bug fixes, and Aider lands in between on speed and thoroughness.
Sources: MorphLLM
Analysis highlights high operational costs of autonomous agents and need for LLM gateway routing.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary
A tool comparison notes that autonomous coding agents can be expensive, with sessions making 50-200+ LLM calls, and emphasizes the critical role of LLM gateway routing for cost and performance management.
Sources: Requesty AI Blog
Analysis illustrates large performance gap between SWE-bench Verified and Pro, highlighting contamination.
Discovered 2026-09-07 · publication date unknown · study · confidence: secondary
A blog post notes Claude Opus 4.5 scored 80.9% on SWE-bench Verified but 45.89% on SWE-bench Pro, a 35-point collapse attributed to the contamination-free design of the latter, which uses GPL-licensed and private repos.
Sources: Paddo.dev
Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary · facet: local
A Medium article based on lessons learned concludes that smaller models work for chat-based assistance, but true autonomous agents require 20B parameter models minimum, with 32B+ being significantly better, creating a tension with VRAM constraints.
Sources: Medium
Guide outlines best practices for pair programming with AI assistants.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: vendor-claim · facet: pairing
A Graphite guide details best practices for leveraging AI coding agents in a pair programming context, covering roles, context provision, code review, and using AI as both a productivity and learning tool.
Sources: Graphite Guides
ForgeCode publishes 12 practical lessons from six months of daily AI pair programming.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: vendor-claim · facet: pairing
A blog post distills lessons from extensive daily AI pair programming, covering planning, prompt engineering, context management, and noting what doesn't work, based on practical experience across multiple codebases.
Sources: ForgeCode Blog
Multiple sources document concepts and architectures for 'agent harness engineering'.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary
Articles from Martin Fowler, Microsoft, Addy Osmani, and others explore the concept of 'harness engineering'—the scaffolding around a coding agent—including runtime design, self-evolution loops, and observability-driven automation.
Sources: Martin Fowler · Microsoft Learn · Addy Osmani
Claude Code docs detail updates for week 20 of 2026, including default fast model change and agent view.
Discovered 2026-09-07 · publication date unknown · update · confidence: primary-source
Claude Code documentation for week 20 (May 11–15, 2026) notes that /fast now defaults to Opus 4.7, and the claude agents command provides a unified view of running sessions and their status.
Sources: Claude Code Docs
OpenAI announces Codex integration in ChatGPT mobile app for working from anywhere.
Discovered 2026-09-07 · publication date unknown · update · confidence: vendor-claim
OpenAI announced a preview of Codex within the ChatGPT mobile app, allowing users to connect to Codex running on their machines (laptop, Mac mini, remote env) and work fluidly across active threads and project context from iOS or Android.
Sources: OpenAI Blog
The last 30 days
The evidence layer thickened this month, with two new randomized controlled trials (RCTs) and a major benchmark shift. Anthropic published an RCT (unverified) investigating whether AI assistance harms skill formation when learning a new library. More critically, the METR trial finding—that experienced developers were ~19% slower with AI—was cited again, reinforcing a negative result that vendors have not contradicted with independent data. In benchmarks, the contamination problem became unavoidable: OpenAI stated it no longer evaluates on SWE-bench Verified, recommending SWE-bench Pro instead. An analysis showed a 35-point performance collapse between the two for Claude Opus 4.5, illustrating how inflated vendor scores on contaminated benchmarks have been.
On the tool front, the multi-agent and multi-repo coordination problem saw practical, if incremental, movement. GitHub enabled multi-agent workflows inside repositories (vendor-claim), and Moderne published docs for its multi-repo agent 'Moddy'. The recommended pattern for current tools remains breaking work into repo-specific sub-tasks, as true multi-repo capabilities are still on roadmaps. In a notable market signal, a Hacker News comment questioned if anyone is still using Cursor in 2026, noting a shift to Claude Code, Codex, or Copilot. Aider's development appears stable, while Claude Code advanced with Opus 5 integration.
For a reader deciding where to spend attention, the priority should be scrutinizing any performance claim. Trust only benchmarks that address contamination, like SWE-bench Pro, and understand that Pass@k scores (used by vendors) are often 15-25 points higher than the more reliable Pass^k metric. The operational cost of autonomous agents remains high, requiring careful LLM gateway management. If you're evaluating tools for internal use, platforms like OpenHands Enterprise now offer governed benchmarking in a private cloud, which is a more credible approach than public vendor leaderboards. The most concrete finding is that the evidence continues to suggest AI assistance does not universally speed up experienced developers, despite their belief that it does.
Written 2026-09-21 from the changelog below, not from a fresh search.
Drawn from:
- Kotlin Benchmark leaderboard updated with new agent results. — 2026-09-21, confidence
primary-source - Anthropic publishes RCT on AI assistance's impact on coding skill formation. — 2026-09-21, confidence
primary-source - Hacker News discussion questions if anyone is still using Cursor in 2026. — 2026-09-21, confidence
secondary - Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. — 2026-09-21, confidence
secondary - OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. — 2026-09-21, confidence
vendor-claim - Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. — 2026-09-21, confidence
vendor-claim - OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center. — 2026-09-21, confidence
secondary - ProdCodeBench benchmark introduced to evaluate AI coding agents on production-derived tasks. — 2026-09-14, confidence
primary-source - Coding Agent Index 2026 claims to be the first independent benchmark of full agent stacks. — 2026-09-14, confidence
secondary - Blog post cites METR randomized trial finding AI-assisted developers 19% slower. — 2026-09-14, confidence
secondary - GitHub enables multi-agent AI coding inside repository workflows. — 2026-09-14, confidence
secondary - Repo-of-Repos pattern described for multi-repo workspace with AI coding agents. — 2026-09-14, confidence
secondary - Scale AI publishes public dataset and leaderboard for SWE-bench Pro. — 2026-09-14, confidence
primary-source - Moderne publishes documentation for its multi-repo AI agent 'Moddy'. — 2026-09-14, confidence
vendor-claim - JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks. — 2026-09-07, confidence
vendor-claim - OpenHands Enterprise platform enables governed AI agent benchmarking in private cloud. — 2026-09-07, confidence
vendor-claim - Google RCT estimates impact of three AI features on developer time for complex tasks. — 2026-09-07, confidence
primary-source - Randomized controlled experiment measures GitHub Copilot's effect on developer productivity. — 2026-09-07, confidence
primary-source - Submodule-based multi-repo workspace strategy proposed to avoid file conflicts with concurrent AI agents. — 2026-09-07, confidence
secondary - GitHub acknowledges multi-repo agent capabilities are on the roadmap, recommends task decomposition for now. — 2026-09-07, confidence
primary-source - Independent comparison benchmarks Claude Code, Cursor, and Aider on different task types. — 2026-09-07, confidence
secondary - Analysis highlights high operational costs of autonomous agents and need for LLM gateway routing. — 2026-09-07, confidence
secondary - Analysis illustrates large performance gap between SWE-bench Verified and Pro, highlighting contamination. — 2026-09-07, confidence
secondary - Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension. — 2026-09-07, confidence
secondary - Guide outlines best practices for pair programming with AI assistants. — 2026-09-07, confidence
vendor-claim - ForgeCode publishes 12 practical lessons from six months of daily AI pair programming. — 2026-09-07, confidence
vendor-claim - Multiple sources document concepts and architectures for 'agent harness engineering'. — 2026-09-07, confidence
secondary - Claude Code docs detail updates for week 20 of 2026, including default fast model change and agent view. — 2026-09-07, confidence
primary-source - OpenAI announces Codex integration in ChatGPT mobile app for working from anywhere. — 2026-09-07, confidence
vendor-claim
The last year
The last twelve months have been defined by a growing rift between vendor claims and independent evidence. The most significant finding is the METR randomized trial, which reported that experienced open-source developers were roughly 19% slower when using AI assistance, directly contradicting widespread productivity promises. This was joined by a Google RCT estimating AI's impact on complex tasks and another study measuring GitHub Copilot's effect, all pointing to a nuanced, often underwhelming reality. Meanwhile, vendors continued to launch new benchmarks, but the credibility of these metrics collapsed under scrutiny. OpenAI stopped evaluating on SWE-bench Verified due to contamination, and independent analyses highlighted a 15-25 point gap between the misleading Pass@k scores vendors report and the more rigorous Pass^k metric. JetBrains launched a Kotlin-specific benchmark, and OpenHands Enterprise offered a platform for private benchmarking, but the overarching lesson is that vendor-run benchmarks are, as one critic put it, as trustworthy as a pharmaceutical company grading its own drug trial.
The tools themselves saw incremental updates rather than breakthroughs. Claude Code, Cursor, and Aider solidified as the main contenders for interactive pair programming, with independent comparisons noting Claude Code excels at complex refactoring while Cursor wins on quick bug fixes. However, a consistent theme was high operational cost and technical constraint: autonomous agents require 20B+ parameter models, creating VRAM tension, and a single session can make 50-200+ LLM calls. VS Code integrated Claude and Codex agents under the Copilot umbrella, and OpenAI launched a Codex app for mobile, extending the workspace but not fundamentally changing the agent's capabilities. The architectural conversation shifted to "harness engineering"—the scaffolding needed to make agents reliable—with multiple sources, including OpenAI, detailing concepts for runtime design and observability.
Where the year delivered concrete progress was in documenting the real, messy failure modes of using these tools at scale, which is more valuable than any benchmark score. Multiple guides emerged from practitioners running concurrent agents, detailing how they built duplicate pipelines or caused file conflicts. In response, the community proposed structural solutions like Git submodules for multi-repo workspaces, while GitHub acknowledged multi-repo agent capabilities are still on the roadmap. Moderne launched 'Moddy', a multi-repo agent for large-scale transformation, and the open-source tool 'The Pair' offered a multi-agent architecture with cross-checking. For anyone deciding where to spend attention, these practical reports on coordination, collision, and cost management are the most load-bearing information available—they map the gap between the demo and the daily grind. The evidence layer now clearly shows that the biggest gains won't come from a faster model, but from better orchestration of the ones we have.
Written 2026-09-07 from the changelog below, not from a fresh search.
| Month | Entries |
|---|---|
| September 2026 | 29 |
| August 2026 | 16 |
The ones that mattered:
- 2026-09-21 — Anthropic publishes RCT on AI assistance's impact on coding skill formation. (
study,primary-source) - 2026-09-21 — Hacker News discussion questions if anyone is still using Cursor in 2026. (
negative,secondary) - 2026-09-21 — OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. (
launch,vendor-claim) - 2026-09-21 — Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. (
study,vendor-claim) - 2026-09-21 — OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center. (
launch,secondary) - 2026-09-14 — Blog post cites METR randomized trial finding AI-assisted developers 19% slower. (
study,secondary) - 2026-09-07 — JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks. (
launch,vendor-claim) - 2026-09-07 — Google RCT estimates impact of three AI features on developer time for complex tasks. (
study,primary-source)
All time
The evidence layer is now the most important part of the field. The METR randomized trial—a genuine controlled experiment—found that experienced open-source developers were 19% slower when using AI assistance, despite believing they were faster. This directly contradicts the productivity claims of every vendor and should be the starting point for any evaluation. If you are considering these tools for a team, your default assumption must now be that they slow down experienced developers; the burden of proof is on any claim to the contrary.
Benchmarks are in crisis, which makes the evidence problem worse. SWE-bench is widely criticized for potential training data contamination, meaning high scores may reflect memorization, not problem-solving. A new benchmark, ProjDevBench, attempts to measure end-to-end project development. In its initial evaluation, Codex+GPT-5 scored highest at 77.85%, but performance varied wildly by task. Treat all benchmark results as unverified and vendor-adjacent until independent replications are done.
The tool landscape is consolidating into platform plays. Visual Studio Code now supports Claude and Codex agents under a GitHub Copilot subscription, bringing multiple “agentic” models into a single editor interface. The open-source counterpoint is tools like The Pair, which uses a multi-agent Mentor+Executor architecture to cross-check code for hallucinations. The actual execution model—how an agent holds context, plans, and verifies work—still varies drastically between tools, but the market is coalescing around IDEs as the orchestration layer.
The major operational finding this period is the concrete failure mode of multi-agent workspaces. Running multiple AI coding sessions on a single repository, as done in this estate, leads to coordination failures: agents independently build duplicate systems without awareness, both reporting success. An AugmentCode guide documents this, and OpenAI’s article on “harness engineering” implicitly acknowledges the need for serious infrastructure to manage agents. The current tools do not solve this; they create the problem. If you move beyond a single human/agent pair, you are entering uncharted territory with real collision and drift risks.
What does not work is assuming AI coding agents universally increase productivity—the best available evidence says the opposite for experienced devs. Also failing is relying on existing benchmarks to gauge real capability. What has changed is the hardening of these negative findings and the emergence of multi-agent chaos as a practical, not theoretical, problem. The open gaps are massive: we have no proven model for agent coordination, no uncontaminated benchmarks, and no replicated evidence that these tools improve outcomes for seasoned developers. The field is currently defined by its failures and uncertainties.
Written 2026-08-18 from the changelog below, not from a fresh search.
How the field breaks down, by what we're actually tracking:
- uncategorised — 18 items: ProjDevBench, METR, Visual Studio Code, Claude Code, Codex, SWE-bench, The Pair, Moderne, Cursor, Aider, Kotlin Benchmark, OpenHands Enterprise, ProdCodeBench, Coding Agent Index, AugmentCode, Scale AI, BenchLM, Bito.
Tracked products
Nothing on this beat is buyable right now, at least nothing we've verified. Everything currently tracked sits in the forward-looking sections below.
Upcoming
Announced, no date.
Nothing in this bucket right now.
Coming Soon
Announced with a date, or an open pre-order.
Nothing in this bucket right now.
What We're Watching
Exists, unproven, or newly discovered. This is where auto-discovered items land.
ProjDevBench
Benchmark for evaluating AI coding agents on end-to-end project development tasks.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: ProjDevBench benchmark evaluates six coding agents on end-to-end project tasks. (2026-08-18) — The benchmark, detailed in an arXiv paper, assesses agents on tasks like from-scratch construction, finding performance varies significantly across models and tasks, with Codex+GPT-5 achieving the best overall score of 77.85%.
Sources: arXiv paper
METR
Organization conducting empirical research on AI safety and impact, including developer productivity studies.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: Blog post cites METR randomized trial finding AI-assisted developers 19% slower. (2026-09-14) — A blog post from Cerbos references the METR randomized trial from July 2025, which found experienced open-source developers using AI tools (like Cursor Pro with Claude) were on average 19% slower, yet believed they were faster.
Sources: METR blog
Visual Studio Code
Code editor with integrated AI coding agent support via extensions and GitHub Copilot.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: GitHub enables multi-agent AI coding inside repository workflows. (2026-09-14) — A news article reports that GitHub has enabled Claude and Codex agents within repository workflows, allowing contributors to submit requests and select multiple agents to execute tasks in parallel, with context attached to the work.
Sources: VS Code blog
Claude Code
AI coding agent by Anthropic, available as a CLI and integrated into various IDEs.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. (2026-09-21) — A 2026 update notes Aider remains at version 0.86.2 with a stable maintenance cadence, while Claude Code has advanced past v2.1.200 with auto-mode support and Opus 5 model integration, narrowing the model gap.
Sources: OpenHands comparison blog
Codex
AI coding agent by OpenAI, often integrated into tools like GitHub Copilot and VS Code.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center. (2026-09-21) — OpenAI launched a standalone desktop app for its Codex AI coding assistant, temporarily available to ChatGPT users with Apple computers, designed as a 'command center' to manage multiple AI agents.
Sources: OpenAI harness engineering post
SWE-bench
Benchmark for evaluating AI models on real-world software engineering issues from GitHub.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5. (2026-09-21) — Bito announced evaluation results showing a Claude Sonnet 4.5 agent augmented with its AI Architect context engine achieved a 60.8% success rate on the SWE-Bench Pro benchmark.
Sources: CodeSOTA guide
The Pair
Open-source desktop app for AI pair programming using a multi-agent cross-checking architecture.
First seen 2026-08-18 · confidence unverified · uncategorised
Why it's here: Open-source tool 'The Pair' provides multi-agent coding with cross-checking. (2026-08-18) — A GitHub repository describes The Pair, an open-source desktop app that uses a Mentor+Executor agent architecture to cross-check code and catch hallucinations, compatible with various AI models.
Sources: GitHub repository
Moderne
Company offering a multi-repo AI agent ('Moddy') for large-scale codebase transformation and maintenance.
First seen 2026-08-24 · confidence unverified · uncategorised
Why it's here: Moderne publishes documentation for its multi-repo AI agent 'Moddy'. (2026-09-14) — Moderne has published user documentation for 'Moddy', describing it as an AI agent that modernizes multi-repository codebases by combining LLMs with structured LST code data.
Sources: Moderne blog
Cursor
AI-powered IDE with integrated coding agents, competing with VS Code and Windsurf.
First seen 2026-08-24 · confidence unverified · uncategorised
Why it's here: Hacker News discussion questions if anyone is still using Cursor in 2026. (2026-09-21) — A Hacker News comment from September 2026 asks if anyone is still actually using Cursor, noting that people they've spoken with are now using Claude Code, Codex, or Copilot.
Sources: Cursor vs. Claude Code comparison
Aider
Open-source, terminal-first AI pair programming agent with deep git integration and multi-model support.
First seen 2026-08-24 · confidence unverified · uncategorised
Why it's here: Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration. (2026-09-21) — A 2026 update notes Aider remains at version 0.86.2 with a stable maintenance cadence, while Claude Code has advanced past v2.1.200 with auto-mode support and Opus 5 model integration, narrowing the model gap.
Sources: Aider skill page on LobeHub
Kotlin Benchmark
JetBrains' official benchmark for evaluating AI coding agents on real-world Kotlin software engineering tasks.
First seen 2026-09-07 · confidence unverified · uncategorised
Why it's here: Kotlin Benchmark leaderboard updated with new agent results. (2026-09-21) — The official Kotlin Benchmark leaderboard shows updated results for various AI coding agents, with Claude Code + Opus 4.6 medium achieving 71.43% resolution and Codex + GPT 5.3 Codex medium achieving 68.57%.
Sources: JetBrains Blog Announcement
OpenHands Enterprise
Enterprise platform for running governed AI coding agent benchmarks in a private cloud with audit logs and cost attribution.
First seen 2026-09-07 · confidence unverified · uncategorised
Why it's here: OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management. (2026-09-21) — OpenHands Enterprise introduces an Agent Control Plane, providing a fully self-hosted, on-premise system for running, controlling, observing, and scaling AI agents across an organization with enforced policies and audit logs.
Sources: OpenHands Blog
ProdCodeBench
Benchmark for evaluating AI coding agents on tasks derived from production code changes.
First seen 2026-09-14 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: ProdCodeBench paper on arXiv
Coding Agent Index
Independent benchmark claiming to evaluate full agent stacks (model + harness).
First seen 2026-09-14 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: Coding Agent Index 2026 article on Medium
AugmentCode
Provider of guides and resources on AI coding practices, including multi-agent coordination.
First seen 2026-09-14 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: AugmentCode guide on multi-agent coding workspace
Scale AI
Company providing AI data and evaluation platforms, including the SWE-bench Pro public dataset.
First seen 2026-09-14 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: Scale AI SWE-bench Pro leaderboard
BenchLM
Platform providing AI model benchmark leaderboards and evaluations, including for SWE-bench Pro.
First seen 2026-09-14 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: BenchLM SWE-bench Pro leaderboard
Bito
Company building deep context graphs for coding agents, with an AI Architect context engine.
First seen 2026-09-21 · confidence unverified · uncategorised
Why it's here: No changelog entry explains this status yet.
Sources: Company mention (PR)
Promising
Early-stage, but the evidence or the approach is genuinely interesting. The only editorial bucket on this page — an item only lands here with a reason recorded in the changelog.
Nothing in this bucket right now.
Open questions
Publishing what we don't know is the point. These are things the job is actively watching for; when one gets answered it becomes a changelog entry and moves down here to the answered list.
No open questions recorded. Either everything we wanted to know got answered, or nobody has written down what we don't know — the second is more likely on a young tracker.
Full changelog
Everything, newest first, grouped by the month we found it. Long by design — it's the receipts.
September 2026
Kotlin Benchmark leaderboard updated with new agent results.
Discovered 2026-09-21 · publication date unknown · update · confidence: primary-source
The official Kotlin Benchmark leaderboard shows updated results for various AI coding agents, with Claude Code + Opus 4.6 medium achieving 71.43% resolution and Codex + GPT 5.3 Codex medium achieving 68.57%.
Sources: Kotlin Benchmark page
Anthropic publishes RCT on AI assistance's impact on coding skill formation.
Discovered 2026-09-21 · publication date unknown · study · confidence: primary-source
Anthropic conducted a randomized controlled trial investigating whether AI assistance prevents software developers from growing their skills or understanding the systems they're building when learning a new Python library.
Sources: Anthropic research
Hacker News discussion questions if anyone is still using Cursor in 2026.
Discovered 2026-09-21 · publication date unknown · negative · confidence: secondary
A Hacker News comment from September 2026 asks if anyone is still actually using Cursor, noting that people they've spoken with are now using Claude Code, Codex, or Copilot.
Sources: Hacker News comment
Comparison notes Aider remains at v0.86.2 while Claude Code advances with Opus 5 integration.
Discovered 2026-09-21 · publication date unknown · update · confidence: secondary
A 2026 update notes Aider remains at version 0.86.2 with a stable maintenance cadence, while Claude Code has advanced past v2.1.200 with auto-mode support and Opus 5 model integration, narrowing the model gap.
Sources: Developers Digest blog
OpenHands Enterprise launches Agent Control Plane for self-hosted, on-premise AI agent management.
Discovered 2026-09-21 · published 2026-05-06 · launch · confidence: vendor-claim
OpenHands Enterprise introduces an Agent Control Plane, providing a fully self-hosted, on-premise system for running, controlling, observing, and scaling AI agents across an organization with enforced policies and audit logs.
Sources: OpenHands blog
Bito's AI Architect achieves 60.8% success rate on SWE-Bench Pro with Claude Sonnet 4.5.
Discovered 2026-09-21 · published 2026-02-03 · study · confidence: vendor-claim
Bito announced evaluation results showing a Claude Sonnet 4.5 agent augmented with its AI Architect context engine achieved a 60.8% success rate on the SWE-Bench Pro benchmark.
Sources: Yahoo Finance (PRNewswire)
OpenAI launches standalone Codex app for Apple computers, serving as a multi-agent command center.
Discovered 2026-09-21 · published 2026-02-02 · launch · confidence: secondary
OpenAI launched a standalone desktop app for its Codex AI coding assistant, temporarily available to ChatGPT users with Apple computers, designed as a 'command center' to manage multiple AI agents.
Sources: CNBC article
ProdCodeBench benchmark introduced to evaluate AI coding agents on production-derived tasks.
Discovered 2026-09-14 · publication date unknown · discovery · confidence: primary-source
A new benchmark, ProdCodeBench, is described as evaluating AI coding agents on tasks derived from production code changes, with a focus on identifying self-consistency checks where agents may reproduce previously generated solutions rather than solving novel tasks.
Sources: ProdCodeBench paper on arXiv
Coding Agent Index 2026 claims to be the first independent benchmark of full agent stacks.
Discovered 2026-09-14 · publication date unknown · discovery · confidence: secondary
A Medium article describes the 'Coding Agent Index 2026' as an independent benchmark evaluating full agent stacks (model + harness), aiming to help developers evaluate stacks before purchasing.
Sources: Coding Agent Index 2026 article on Medium
Blog post cites METR randomized trial finding AI-assisted developers 19% slower.
Discovered 2026-09-14 · publication date unknown · study · confidence: secondary
A blog post from Cerbos references the METR randomized trial from July 2025, which found experienced open-source developers using AI tools (like Cursor Pro with Claude) were on average 19% slower, yet believed they were faster.
Sources: Cerbos blog post on AI productivity paradox
GitHub enables multi-agent AI coding inside repository workflows.
Discovered 2026-09-14 · publication date unknown · update · confidence: secondary · facet: orchestration
A news article reports that GitHub has enabled Claude and Codex agents within repository workflows, allowing contributors to submit requests and select multiple agents to execute tasks in parallel, with context attached to the work.
Sources: Help Net Security article on GitHub multi-agent coding
Repo-of-Repos pattern described for multi-repo workspace with AI coding agents.
Discovered 2026-09-14 · publication date unknown · discovery · confidence: secondary
A blog post details the 'Repo-of-Repos' pattern, a template that uses an outer 'agent' repo to pull in related repos as workspace folders, facilitating multi-repo work for AI coding agents while maintaining separate commit histories.
Sources: Rafferty Uy blog post on Repo-of-Repos pattern
Scale AI publishes public dataset and leaderboard for SWE-bench Pro.
Discovered 2026-09-14 · publication date unknown · update · confidence: primary-source
Scale AI has launched a public dataset and leaderboard for SWE-bench Pro, describing it as a benchmark designed to address limitations in existing benchmarks with 1865 tasks across 41 professional repositories.
Sources: Scale AI SWE-bench Pro leaderboard
Moderne publishes documentation for its multi-repo AI agent 'Moddy'.
Discovered 2026-09-14 · publication date unknown · discovery · confidence: vendor-claim · facet: orchestration
Moderne has published user documentation for 'Moddy', describing it as an AI agent that modernizes multi-repository codebases by combining LLMs with structured LST code data.
Sources: Moderne Docs for Moddy
JetBrains launches Kotlin Benchmark for evaluating AI coding agents on real-world Kotlin tasks.
Discovered 2026-09-07 · publication date unknown · launch · confidence: vendor-claim
JetBrains released an official benchmark designed to assess how different AI coding agents perform on Kotlin software engineering tasks, aiming to provide a credible, public evaluation closer to day-to-day development work.
Sources: JetBrains Blog
OpenHands Enterprise platform enables governed AI agent benchmarking in private cloud.
Discovered 2026-09-07 · publication date unknown · update · confidence: vendor-claim
OpenHands Enterprise offers a platform for running AI coding agent benchmarks in a customer's own virtual private cloud with audit logs and cost attribution, positioning it for internal, governed evaluation at scale.
Sources: OpenHands Blog
Google RCT estimates impact of three AI features on developer time for complex tasks.
Discovered 2026-09-07 · publication date unknown · study · confidence: primary-source
A randomized controlled trial with 96 full-time Google software engineers provides an estimate of how three AI features affect the time developers spend on a complex, enterprise-grade task.
Sources: arXiv
Randomized controlled experiment measures GitHub Copilot's effect on developer productivity.
Discovered 2026-09-07 · publication date unknown · study · confidence: primary-source
A study employing an RCT with 95 professional programmers recruited via Upwork measures the effect of GitHub Copilot on developer productivity, contributing to the empirical evidence base.
Sources: alphaXiv
Submodule-based multi-repo workspace strategy proposed to avoid file conflicts with concurrent AI agents.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary · facet: orchestration
A Medium article outlines a workspace strategy using Git submodules to allow multiple AI coding agents to work simultaneously on different service repos within a larger project, structurally preventing file conflicts.
Sources: Medium
GitHub acknowledges multi-repo agent capabilities are on the roadmap, recommends task decomposition for now.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: primary-source · facet: orchestration
A GitHub community discussion reveals that background agents currently work best within a single repo context, with multi-repo capabilities on the roadmap; the recommended current approach is to break work into repo-specific sub-tasks.
Sources: GitHub Discussion
Independent comparison benchmarks Claude Code, Cursor, and Aider on different task types.
Discovered 2026-09-07 · publication date unknown · study · confidence: secondary
A benchmark analysis finds Claude Code excels at complex, multi-file tasks like refactoring and migrations, Cursor wins on quick, focused tasks like bug fixes, and Aider lands in between on speed and thoroughness.
Sources: MorphLLM
Analysis highlights high operational costs of autonomous agents and need for LLM gateway routing.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary
A tool comparison notes that autonomous coding agents can be expensive, with sessions making 50-200+ LLM calls, and emphasizes the critical role of LLM gateway routing for cost and performance management.
Sources: Requesty AI Blog
Analysis illustrates large performance gap between SWE-bench Verified and Pro, highlighting contamination.
Discovered 2026-09-07 · publication date unknown · study · confidence: secondary
A blog post notes Claude Opus 4.5 scored 80.9% on SWE-bench Verified but 45.89% on SWE-bench Pro, a 35-point collapse attributed to the contamination-free design of the latter, which uses GPL-licensed and private repos.
Sources: Paddo.dev
Analysis states effective autonomous coding agents require 20B+ parameter models, creating VRAM tension.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary · facet: local
A Medium article based on lessons learned concludes that smaller models work for chat-based assistance, but true autonomous agents require 20B parameter models minimum, with 32B+ being significantly better, creating a tension with VRAM constraints.
Sources: Medium
Guide outlines best practices for pair programming with AI assistants.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: vendor-claim · facet: pairing
A Graphite guide details best practices for leveraging AI coding agents in a pair programming context, covering roles, context provision, code review, and using AI as both a productivity and learning tool.
Sources: Graphite Guides
ForgeCode publishes 12 practical lessons from six months of daily AI pair programming.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: vendor-claim · facet: pairing
A blog post distills lessons from extensive daily AI pair programming, covering planning, prompt engineering, context management, and noting what doesn't work, based on practical experience across multiple codebases.
Sources: ForgeCode Blog
Multiple sources document concepts and architectures for 'agent harness engineering'.
Discovered 2026-09-07 · publication date unknown · discovery · confidence: secondary
Articles from Martin Fowler, Microsoft, Addy Osmani, and others explore the concept of 'harness engineering'—the scaffolding around a coding agent—including runtime design, self-evolution loops, and observability-driven automation.
Sources: Martin Fowler · Microsoft Learn · Addy Osmani
Claude Code docs detail updates for week 20 of 2026, including default fast model change and agent view.
Discovered 2026-09-07 · publication date unknown · update · confidence: primary-source
Claude Code documentation for week 20 (May 11–15, 2026) notes that /fast now defaults to Opus 4.7, and the claude agents command provides a unified view of running sessions and their status.
Sources: Claude Code Docs
OpenAI announces Codex integration in ChatGPT mobile app for working from anywhere.
Discovered 2026-09-07 · publication date unknown · update · confidence: vendor-claim
OpenAI announced a preview of Codex within the ChatGPT mobile app, allowing users to connect to Codex running on their machines (laptop, Mac mini, remote env) and work fluidly across active threads and project context from iOS or Android.
Sources: OpenAI Blog
August 2026
Article highlights gap between Pass@k and Pass^k metrics, inflating vendor-reported scores.
Discovered 2026-08-31 · publication date unknown · study · confidence: secondary · facet: autonomous
A SoftwareSeni article explains that Pass@k measures success across k attempts (rewarding luck), while Pass^k requires success on every attempt (measuring reliability), with a typical gap of 15-25 percentage points that makes vendor scores misleading.
Sources: SoftwareSeni
OpenAI announces it no longer evaluates on SWE-bench Verified, citing contamination issues.
Discovered 2026-08-31 · publication date unknown · negative · confidence: primary-source · facet: autonomous
OpenAI states it has stopped reporting results on SWE-bench Verified due to contamination concerns and now recommends SWE-bench Pro, which empirically suffers less from contamination, though it is not perfect.
Sources: OpenAI blog
VS Code 1.134 adds features for organizing chats across windows and navigating long conversations.
Discovered 2026-08-24 · published 2026-08-19 · update · confidence: primary-source
The release, dated August 19, 2026, introduces capabilities to work across windows, organize related chats side by side, and navigate long conversations faster, as per the official VS Code update notes.
Sources: Visual Studio Code 1.134 release notes
Claude Code promotional weekly usage limits are ending, reverting to previous levels.
Discovered 2026-08-24 · publication date unknown · update · confidence: secondary
A Hacker News discussion notes that a promotion offering 50% higher weekly usage limits from May 13 to August 19, 2026, is ending, with limits reverting to pre-promotion levels.
Sources: Hacker News discussion on Claude Code limits
Independent analysis criticizes vendor-run AI code review benchmarks as biased and lacking trustworthy methodology.
Discovered 2026-08-24 · publication date unknown · study · confidence: secondary
A DeepSource blog post argues that self-evaluation by vendors is inherently biased, drawing parallels to pharmaceutical clinical trials, and calls for independent evaluation, published datasets, and reproducible methodology.
Sources: DeepSource blog on AI code review benchmarks
Google randomized controlled trial estimates the impact of three AI features on developer time for complex tasks.
Discovered 2026-08-24 · publication date unknown · study · confidence: primary-source
An arXiv paper details a randomized controlled trial with 96 full-time Google software engineers, contributing an estimate of how AI features affect the time developers spend on a complex, enterprise-grade task.
Sources: arXiv paper on AI impact on developer productivity
Moderne launches 'Moddy', a multi-repo AI agent for transforming enterprise codebases at scale.
Discovered 2026-08-24 · publication date unknown · launch · confidence: vendor-claim
Moderne introduces an AI agent designed to handle transformations across entire codebases simultaneously, contrasting with tools focused on single repositories.
Sources: Moderne blog introducing multi-repo AI agent
OpenAI's Codex app, initially for macOS, is now available on Windows.
Discovered 2026-08-24 · published 2026-03-04 · update · confidence: primary-source
An update to OpenAI's February 2026 announcement states the Codex app, an interface for managing multiple agents and parallel work, is now available on Windows.
Sources: OpenAI blog introducing the Codex app
SWE-bench Pro leaderboard for August 2026 shows Claude Mythos 5 leading with 80.3%.
Discovered 2026-08-24 · publication date unknown · update · confidence: secondary
A BenchLM.ai leaderboard update lists Claude Mythos 5 ahead of Claude Fable 5 and Claude Opus 5 on the SWE-bench Pro benchmark, which is designed for long-horizon, realistic software engineering work.
Sources: BenchLM.ai SWE-bench Pro leaderboard
ProjDevBench benchmark evaluates six coding agents on end-to-end project tasks.
Discovered 2026-08-18 · publication date unknown · study · confidence: primary-source
The benchmark, detailed in an arXiv paper, assesses agents on tasks like from-scratch construction, finding performance varies significantly across models and tasks, with Codex+GPT-5 achieving the best overall score of 77.85%.
Sources: arXiv ProjDevBench paper
METR randomized trial finds AI assistance made experienced open-source developers 19% slower.
Discovered 2026-08-18 · publication date unknown · study · confidence: primary-source
A METR blog post reports a randomized controlled trial where developers using AI tools took longer on tasks than those without, contradicting common productivity claims.
Sources: METR blog post
Guide documents failure modes of running multiple AI coding agents on one repository.
Discovered 2026-08-18 · publication date unknown · discovery · confidence: secondary
An AugmentCode guide outlines coordination challenges and failure modes when multiple agents work concurrently on the same codebase without proper orchestration.
Sources: AugmentCode guide
VS Code adds support for Claude and Codex agents under GitHub Copilot subscription.
Discovered 2026-08-18 · publication date unknown · update · confidence: primary-source
A VS Code blog post announces that version 1.109 allows running Claude and Codex agents locally or in the cloud as part of Copilot Pro+ and Enterprise subscriptions.
Sources: VS Code blog
Criticism highlights contamination concerns in SWE-bench benchmark.
Discovered 2026-08-18 · publication date unknown · study · confidence: secondary
Multiple sources, including an arXiv paper and blog posts, discuss how training data contamination may inflate model performance on SWE-bench, questioning what the benchmark actually measures.
Sources: arXiv paper on SWE-bench illusion · Latent Space article
Open-source tool 'The Pair' provides multi-agent coding with cross-checking.
Discovered 2026-08-18 · publication date unknown · discovery · confidence: primary-source
A GitHub repository describes The Pair, an open-source desktop app that uses a Mentor+Executor agent architecture to cross-check code and catch hallucinations, compatible with various AI models.
Sources: GitHub repository for The Pair
OpenAI publishes article on 'harness engineering' for coding agents.
Discovered 2026-08-18 · publication date unknown · discovery · confidence: primary-source
An OpenAI blog post discusses the architectural constraints and infrastructure needed to effectively leverage coding agents like Codex in production, emphasizing early investment in harness design.
Sources: OpenAI blog post
Dreamlab Live Research uses autonomous systems to track the state of the art in fields of interest. New trackers appear as the interests do.
Subscribe / join for details — placeholder: there is nothing to sign up to yet.
Generated from a versioned dataset in a git repository: every state this page has ever been in is a commit, and a bad run is revertible. The dataset is the source; this page is build output and is not itself edited.
- Dataset:
data/live-research/ai-programming.json(schema version 2) - Runs completed: 6 · cadence: weekly
- Items tracked: 18 · changelog entries: 45
- Distinct sources seen: 204
- Last run recorded: 2026-09-21T05:35:49Z (status:
ok)