AINEWS 搜索
返回 Rohan Paul (@rohanpaul_ai)
Rohan Paul (@rohanpaul_ai)· · 原发布时间

研究提出让小模型依据运行结果改写 AI 智能体配置代码

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 辅助摘要

这项研究让小模型读取智能体的配置代码和失败报告,再通过改写代码调整模型可见的信息与工具调用。在21类未见过的推理任务上,4B 编辑模型的平均修改得分从0.32升至0.62;材料还称,多轮修改优于单次修改。

正文 · 原文

– https://t.co/G6AG9yusKj

Title: "Harness Learning Enables Generalizable Test-Time Adaptation"

引用或回复的背景(作者 ID 2588345408,https://x.com/i/status/2107252281507012973): New John Hopkins and Carnegie Mellon University Paper Shows that a small model can learn to improve an agent's harness code from run results, and the skill transfers to new tasks.

A small model trained to rewrite an agent's harness code from failure reports can adapt agents to new tasks, so let it tune your harness instead of doing it by hand.

The harness is the code that decides what the model sees and which tools it calls. The editor reads the harness and what failed, then writes a code change, rewarded by how well the new harness scores.

On 21 unseen reasoning task types, a 4B editor's average edit score rose from 0.32 to 0.62, above its 35B teacher. A separate editor trained on HotpotQA kept improving harnesses on 2 other QA benchmarks.

Run it for several rounds per task, since repeated edits beat single fixes.

发现内容有误?提交纠错