AINEWS 搜索
主题与信息类型

智能体 研究 AI 安全与评测

返回 Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · 原发布时间 AI 评分54

VERA 研究提出交替训练模型与修改技能文件,改进多步骤任务智能体

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 导读

一则对 NVIDIA 论文的介绍称,VERA 将基准测试转为逾 9,000 个可重启的沙盒,依据真实文件和日志逐步评分。在所述医疗研究基准测试中,9B 智能体同时采用模型训练和技能文件修改得分为 69.1,分别只采用其中一种方法时得分为 43.3 和 56.1。

正文 · 原文

该语言的正文暂不可用,当前显示已有版本。

– https://t.co/dJrFo7NStw

Title: "VERA: Scaling Verifiable Environments for Agentic co-Evolution"

回复Rohan Paul@rohanpaul_ai
New Nvidia paper shows agents for long, multi-step work improve most when you alternate between training the model and editing its skill files, using step-by-step scores from real evidence to pick each fix. Training only the model or only the harness leaves about half the gain on the table, compared with updating both in alternating rounds. Most environments score only the final result, which hides which step broke. VERA builds over 9,000 restartable sandboxes that check each step against real files and logs, then improves the agent in rounds. VERA turns benchmark runs into over 9,000 restartable sandboxes where each step of a long workflow gets its own checklist score from real evidence. On a medical research benchmark, a 9B agent scored 69.1 with both kinds of updates, versus 56.1 with skill edits alone and 43.3 with training alone. Score each step against real artifacts, and let those scores decide whether the next fix goes into the model or its skills.
在 X 查看回复的帖子

来源:Rohan Paul · x.com

论文
发现内容有误?提交纠错