Article · Original
The article text is unavailable in this language; an existing version is shown.
Agent benchmarks need audits of the graders, not only the tasks.
Parsewave reviewed all 600 public tasks in Zapier's AutomationBench, which frontier labs cite on their model cards.
It used agents to write convincing wrong answers, then had humans confirm which graders actually failed: 206 did.
AutomationBench Verified fixes every one.
On 1,235 Kimi K3 runs, the fixed graders gave a different verdict 27.9% of the time.
The clearest case is task 813. The grader checked the Salesforce notes but never looked at the DocuSign template. A submission that sent four contracts on the Standard template, instead of the required GDPR, HIPAA, SOC2 and Enterprise ones, passed 13 of 13 checks and scored 1.0. After the fix it scores 0.09.