AINEWS Search
Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time SelectedAI score60

Parsewave releases AutomationBench Verified, fixing 206 grading errorsMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

Parsewave audited all 600 public tasks in Zapier's AutomationBench and says it confirmed and fixed 206 grading errors. Regrading 1,235 Kimi K3 runs changed 344 verdicts (27.9%), showing how graders can affect benchmark results.

Why it matters

对关注模型评测的读者,1,235 次运行中有 27.9% 因判分器修复而改变结果,是解读该基准成绩时需要留意的具体影响。

Article · Original

The article text is unavailable in this language; an existing version is shown.

Agent benchmarks need audits of the graders, not only the tasks.

Parsewave reviewed all 600 public tasks in Zapier's AutomationBench, which frontier labs cite on their model cards.

It used agents to write convincing wrong answers, then had humans confirm which graders actually failed: 206 did.

AutomationBench Verified fixes every one.

On 1,235 Kimi K3 runs, the fixed graders gave a different verdict 27.9% of the time.

The clearest case is task 813. The grader checked the Salesforce notes but never looked at the DocuSign template. A submission that sent four contracts on the Standard template, instead of the required GDPR, HIPAA, SOC2 and Enterprise ones, passed 13 of 13 checks and scored 1.0. After the fix it scores 0.09.

Quoteshaped@shaped
AutomationBench Verified is out. We went through all 600 public tasks in @Zapier's AutomationBench and checked every verifier. • Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed. • We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade • Where verifiers were too strict, pass rate went from 18.8% to 43.8% • Where they were too lenient, it dropped from 60.2% to 49.7% Audit and dataset links below.
View quoted post on X

来源:Rohan Paul · x.com

Research
Found an error? Send a correction