AINEWS Search
Back Rohan Paul (@rohanpaul_ai)
Rohan Paul (@rohanpaul_ai)· · Original publication time

ActiveSaddler Dynamically Selects Agent Training Tasks Based on Unresolved FailuresMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI-assisted summary

A post describing a Microsoft paper says ActiveSaddler tracks agent failure patterns, prioritizes problems worth fixing, and tries unseen tasks. The post reports that, using the same optimizer, it raised test pass rates by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0. Reaching 58.5% accuracy on the GAIA2 development set cost $298, compared with $1,360 using a fixed task order.

Article · Original

– https://t.co/wciUS3C7fW

Title: "ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization"

引用或回复的背景(作者 ID 2588345408,https://x.com/i/status/2107132234373468545): New Microsoft paper on Automated harness optimization for agents.

Most harness auto-tuners focus on how to patch prompts and tools, but which tasks produce the feedback also changes how good the final harness gets.

But you will get stronger AI agents when you pick training tasks based on which failures are still unfixed, so stop feeding them a fixed task list.

ActiveSaddler tracks failure patterns and works on the one most worth fixing, or tries unseen tasks to find new ones.

On the same optimizer, ActiveSaddler raised test pass rates by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0. Reaching 58.5% GAIA2 dev accuracy cost $298, versus $1,360 with a fixed order.

If you auto-tune an agent, aim your run budget at the failures that are still open.

Found an error? Send a correction