AINEWS Search
Back Rohan Paul (@rohanpaul_ai)
Rohan Paul (@rohanpaul_ai)· · Original publication time

Study trains a small model to revise AI agent harness code from run resultsMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI-assisted summary

The study has a small model read an agent’s harness code and failure reports, then revise the code that controls what the model sees and which tools it calls. Across 21 unseen reasoning task types, a 4B editor’s average edit score rose from 0.32 to 0.62; the source also says repeated edits beat single fixes.

Article · Original

– https://t.co/G6AG9yusKj

Title: "Harness Learning Enables Generalizable Test-Time Adaptation"

引用或回复的背景(作者 ID 2588345408,https://x.com/i/status/2107252281507012973): New John Hopkins and Carnegie Mellon University Paper Shows that a small model can learn to improve an agent's harness code from run results, and the skill transfers to new tasks.

A small model trained to rewrite an agent's harness code from failure reports can adapt agents to new tasks, so let it tune your harness instead of doing it by hand.

The harness is the code that decides what the model sees and which tools it calls. The editor reads the harness and what failed, then writes a code change, rewarded by how well the new harness scores.

On 21 unseen reasoning task types, a 4B editor's average edit score rose from 0.32 to 0.62, above its 35B teacher. A separate editor trained on HotpotQA kept improving harnesses on 2 other QA benchmarks.

Run it for several rounds per task, since repeated edits beat single fixes.

Found an error? Send a correction