AINEWS Search
Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time AI score54

Meta paper says mixed coding-agent review catches more silent bugsMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

A Meta paper says having different coding agents review each other’s patches catches more silent bugs than giving one agent a bigger budget. At matched spend, mixing Claude Code and Codex raised fully correct patches from 45.8% to 62.5%. The material says such errors in training code can leave programs running while wasting GPU time and producing results that appear valid.

Article · Original

The article text is unavailable in this language; an existing version is shown.

New Meta paper finds that having 2 different coding agents review each other's patches catches far more silent bugs than giving 1 agent a bigger budget.

Mixing Claude Code and Codex on the same task raised fully correct patches from 45.8% to 62.5% at matched spend.

Agents editing real training code can leak test data, break a gradient, or miswire a flag.

The code still runs, so you burn GPU hours and get numbers that look valid.

The reason is uncorrelated mistakes: agents from the same product fail the same way, so there is little for review to catch. The effect repeated on an unrelated training codebase.

– arxiv. org/abs/2609.39551

Title: "RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models"

来源:Rohan Paul · x.com

Research
Found an error? Send a correction