AINEWS Search
Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time AI score54

Google paper proposes AIM to rank research ideas and check their codeMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

The research says AIM ranks ideas by theme, explores both strong and untested themes, and checks whether code matches the idea being evaluated. It reports scores 1.6 and 4.9 points higher than the previous best agent, ScientistOne, on two task groups, and says AIM matched ScientistOne’s best score up to 3.1 times sooner.

Article · Original

The article text is unavailable in this language; an existing version is shown.

New Google paper shows research agents get better results sooner when they keep a ranked map of ideas and check that code matches each idea, so build both in.

Most agents just keep editing code.

When code drifts from the idea it's scored as, the agent learns the wrong lesson.

AIM sorts ideas into ranked themes and splits each round between strong themes and untested ones. An auditor tosses gamed results and relabels ideas to match the code.

It beat the best prior agent, ScientistOne, by 1.6 and 4.9 points on 2 task groups, and matched ScientistOne's best score up to 3.1x sooner.

Expect the biggest gains on tasks with many possible approaches and few good ones.

来源:Rohan Paul · x.com

Research
Found an error? Send a correction