AINEWS 搜索
返回 Rohan Paul (@rohanpaul_ai)
Rohan Paul (@rohanpaul_ai)· · 原发布时间

ActiveSaddler 用未解决的失败案例动态选择智能体训练任务

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 辅助摘要

一则介绍微软论文的帖子称,ActiveSaddler 会追踪智能体的失败模式,优先选择值得修复的问题,也会尝试未见过的任务。该帖报告,在相同优化器下,它使 GAIA2 和 Terminal-Bench 2.0 的测试通过率分别提高 4.4 和 7.5 个百分点;达到 58.5% 的 GAIA2 开发集准确率花费 298 美元,固定任务顺序则花费 1,360 美元。

正文 · 原文

– https://t.co/wciUS3C7fW

Title: "ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization"

引用或回复的背景(作者 ID 2588345408,https://x.com/i/status/2107132234373468545): New Microsoft paper on Automated harness optimization for agents.

Most harness auto-tuners focus on how to patch prompts and tools, but which tasks produce the feedback also changes how good the final harness gets.

But you will get stronger AI agents when you pick training tasks based on which failures are still unfixed, so stop feeding them a fixed task list.

ActiveSaddler tracks failure patterns and works on the one most worth fixing, or tries unseen tasks to find new ones.

On the same optimizer, ActiveSaddler raised test pass rates by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0. Reaching 58.5% GAIA2 dev accuracy cost $298, versus $1,360 with a fixed order.

If you auto-tune an agent, aim your run budget at the failures that are still open.

发现内容有误?提交纠错