AINEWS Search
Topics and information types

智能体 研究 AI 安全与评测

Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time AI score54

VERA study proposes alternating model training and skill-file edits for multistep agentsMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

A post describing an NVIDIA paper says VERA turns benchmark runs into more than 9,000 restartable sandboxes that score each step against real files and logs. On the reported medical research benchmark, a 9B agent scored 69.1 with both model training and skill-file edits, versus 43.3 and 56.1 with either approach alone, respectively.

Article · Original

The article text is unavailable in this language; an existing version is shown.

– https://t.co/dJrFo7NStw

Title: "VERA: Scaling Verifiable Environments for Agentic co-Evolution"

ReplyRohan Paul@rohanpaul_ai
New Nvidia paper shows agents for long, multi-step work improve most when you alternate between training the model and editing its skill files, using step-by-step scores from real evidence to pick each fix. Training only the model or only the harness leaves about half the gain on the table, compared with updating both in alternating rounds. Most environments score only the final result, which hides which step broke. VERA builds over 9,000 restartable sandboxes that check each step against real files and logs, then improves the agent in rounds. VERA turns benchmark runs into over 9,000 restartable sandboxes where each step of a long workflow gets its own checklist score from real evidence. On a medical research benchmark, a 9B agent scored 69.1 with both kinds of updates, versus 56.1 with skill edits alone and 43.3 with training alone. Score each step against real artifacts, and let those scores decide whether the next fix goes into the model or its skills.
View replied-to post on X

来源:Rohan Paul · x.com

Research
Found an error? Send a correction