正文 · 原文
该语言的正文暂不可用,当前显示已有版本。
@IntuitMachine Seems like good strategy even now in late 2026
自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。
一则介绍 Meta 论文的帖子称,让不同编码代理互审补丁,比增加单个代理的预算更能发现隐蔽错误。在相同支出下,Claude Code 与 Codex 配合使完全正确的补丁比例从 45.8% 升至 62.5%;帖子称,这一效果也在另一套训练代码中出现。
该语言的正文暂不可用,当前显示已有版本。
@IntuitMachine Seems like good strategy even now in late 2026
Tip: Use different AI vendors to critique your work.
在 X 查看回复的帖子
New Meta paper finds that having 2 different coding agents review each other's patches catches far more silent bugs than giving 1 agent a bigger budget.
Mixing Claude Code and Codex on the same task raised fully correct patches from 45.8% to 62.5% at matched spend.
Agents editing real training code can leak test data, break a gradient, or miswire a flag.
The code still runs, so you burn GPU hours and get numbers that look valid.
The reason is uncorrelated mistakes: agents from the same product fail the same way, so there is little for review to catch. The effect repeated on an unrelated training codebase.
– arxiv. org/abs/2609.39551
Title: "RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models"
在 X 查看上下文