AINEWS 搜索
返回 Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · 原发布时间 AI 评分54

Meta 论文称:混用两种编程智能体审查代码可减少隐蔽错误

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 导读

一篇 Meta 论文称,让不同编程智能体互审代码补丁,比给单个智能体更多预算更能发现隐蔽错误。在相同支出下,混用 Claude Code 和 Codex 使完全正确的补丁比例从 45.8% 升至 62.5%。材料称,训练代码中的这类错误可能不会让程序停止运行,却会耗费 GPU 时间并产生看似有效的结果。

正文 · 原文

该语言的正文暂不可用,当前显示已有版本。

New Meta paper finds that having 2 different coding agents review each other's patches catches far more silent bugs than giving 1 agent a bigger budget.

Mixing Claude Code and Codex on the same task raised fully correct patches from 45.8% to 62.5% at matched spend.

Agents editing real training code can leak test data, break a gradient, or miswire a flag.

The code still runs, so you burn GPU hours and get numbers that look valid.

The reason is uncorrelated mistakes: agents from the same product fail the same way, so there is little for review to catch. The effect repeated on an unrelated training codebase.

– arxiv. org/abs/2609.39551

Title: "RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models"

来源:Rohan Paul · x.com

论文
发现内容有误?提交纠错