AINEWS Search
Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time AI score47

AutomationBench Verified released with fixes for 206 task-grading bugsMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

The project says it checked the graders for all 600 public tasks in Zapier's AutomationBench and fixed 206 bugs confirmed by human review. Regrading 1,235 Kimi K3 runs with the old and fixed graders changed 344 grades (27.9%).

Article · Original

The article text is unavailable in this language; an existing version is shown.

@shaped @zapier This is solid.

A benchmark score is only as honest as the code that grades it.

Replyshaped@shaped
AutomationBench Verified is out. We went through all 600 public tasks in @Zapier's AutomationBench and checked every verifier. • Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed. • We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade • Where verifiers were too strict, pass rate went from 18.8% to 43.8% • Where they were too lenient, it dropped from 60.2% to 49.7% Audit and dataset links below.
View replied-to post on X

来源:Rohan Paul · x.com

Research
Found an error? Send a correction