Article · Original
The article text is unavailable in this language; an existing version is shown.
@shaped @zapier This is solid.
A benchmark score is only as honest as the code that grades it.
Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.
The project says it checked the graders for all 600 public tasks in Zapier's AutomationBench and fixed 206 bugs confirmed by human review. Regrading 1,235 Kimi K3 runs with the old and fixed graders changed 344 grades (27.9%).
The article text is unavailable in this language; an existing version is shown.
@shaped @zapier This is solid.
A benchmark score is only as honest as the code that grades it.
AutomationBench Verified is out.
We went through all 600 public tasks in @Zapier's AutomationBench and checked every verifier.
• Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed.
• We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade
• Where verifiers were too strict, pass rate went from 18.8% to 43.8%
• Where they were too lenient, it dropped from 60.2% to 49.7%
Audit and dataset links below.
View replied-to post on X