AINEWS 搜索
返回 Rohan Paul (@rohanpaul_ai)
Rohan Paul (@rohanpaul_ai)· · 原发布时间

研究者推出 AgentBug-Smith,将智能体故障报告转成可运行测试

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 辅助摘要

AgentBug-Smith 将 GitHub 上的真实故障报告转为可运行测试,构建了一个包含 200 个故障、可持续扩充的基准。材料称,三款编程智能体中表现最好的一款只修复了其中 9% 的故障;一份总结过往修复经验的简短指南,则让一款智能体在 79 个未见过的故障中正确修复的数量从 1 个增至 6 个。

正文 · 原文

– https://t.co/quiAZ2O9Ac

Title: "AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems"

引用或回复的背景(作者 ID 2588345408,https://x.com/i/status/2106240755857801348): Self-improving AI agents will need to fix their own code.

And this paper from top US+China labs, shows coding agents miss most such bugs but improve with lessons from past fixes.

that real bugs in agent harnesses, can be automatically turned into a growing set of runnable tests.

An agent's own code is everything around the model: tool calls, memory, and prompts. Its bugs depend on live model calls, which makes them hard to recreate and test.

So the researchers built AgentBug-Smith, which turns real GitHub bug reports into runnable tests. The result is a 200-bug benchmark that keeps growing.

The best of 3 coding agents fixed just 9% of those bugs, versus about 40% reported on regular software bugs. A short guide of lessons from past fixes lifted an agent from 1 to 6 correct fixes on 79 unseen bugs.

Before trusting a coding agent with your agent's code, try it on bugs you've already fixed.

发现内容有误?提交纠错