Parsewave reviewed all 600 public tasks in Zapier's AutomationBench and fixed 206 grader bugs. When 1,235 Kimi K3 runs were graded again with the fixed verifiers, 344 (27.9%) received a different grade, showing how grading rules can affect assessments of agent performance.
同一事件,精选展示《Parsewave 发布 AutomationBench Verified,修复 206 个任务判分错误》Parsewave 发布 AutomationBench Verified,修复 206 个任务判分错误
Parsewave 审查了 Zapier 的 AutomationBench 全部 600 个公开任务,经人工确认并修复 206 处判分错误。用修复前后的判分器重评 1,235 次 Kimi K3 运行时,344 次(27.9%)得到不同结果,显示基准测试的评分方式会影响模型成绩。
Event developments
Coverage of the same occurrence shares a node; subsequent developments have their own nodes.
Parsewave releases AutomationBench Verified, fixing 206 grading errors
Parsewave audited all 600 public tasks in Zapier's AutomationBench and says it confirmed and fixed 206 grading errors. Regrading 1,235 Kimi K3 runs changed 344 verdicts (27.9%), showing how graders can affect benchmark results.
View source reports · 1 sources · 3 reports
- Rohan Paul (@rohanpaul_ai) · AutomationBench Verified released with fixes for 206 task-grading bugs
- Rohan Paul (@rohanpaul_ai) · Parsewave releases AutomationBench Verified, fixing 206 grading errors
- Rohan Paul (@rohanpaul_ai) · Parsewave releases AutomationBench Verified, fixing 206 grading bugs
All reports · 3
By original publication time, with each report's bookmarks and feedback preserved.
Parsewave audited all 600 public tasks in Zapier's AutomationBench and says it confirmed and fixed 206 grading errors. Regrading 1,235 Kimi K3 runs changed 344 verdicts (27.9%), showing how graders can affect benchmark results.
The project says it checked the graders for all 600 public tasks in Zapier's AutomationBench and fixed 206 bugs confirmed by human review. Regrading 1,235 Kimi K3 runs with the old and fixed graders changed 344 grades (27.9%).
同一事件,精选展示《Parsewave 发布 AutomationBench Verified,修复 206 个任务判分错误》