正文 · 原文
该语言的正文暂不可用,当前显示已有版本。
– https://t.co/zmwn4jdUp2
Title: "AIM: Agentic Idea Management for Automated Research"
自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。
一则介绍 Google 论文的帖子称,AIM 将研究思路按主题排序,并在每轮尝试中兼顾已有较好结果的方向和未经测试的方向;审查环节会剔除被取巧的结果,并核对代码与思路是否一致。帖子称,在两组任务中,AIM 比此前表现最好的智能体 ScientistOne 分别高出 1.6 和 4.9 分,达到后者最佳分数的速度最高快 3.1 倍。
该语言的正文暂不可用,当前显示已有版本。
– https://t.co/zmwn4jdUp2
Title: "AIM: Agentic Idea Management for Automated Research"
New Google paper shows research agents get better results sooner when they keep a ranked map of ideas and check that code matches each idea, so build both in.
Most agents just keep editing code.
When code drifts from the idea it's scored as, the agent learns the wrong lesson.
AIM sorts ideas into ranked themes and splits each round between strong themes and untested ones. An auditor tosses gamed results and relabels ideas to match the code.
It beat the best prior agent, ScientistOne, by 1.6 and 4.9 points on 2 task groups, and matched ScientistOne's best score up to 3.1x sooner.
Expect the biggest gains on tasks with many possible approaches and few good ones.
在 X 查看回复的帖子