AINEWS Search
← Back to eventsLatest development 10月07日 05:55
STORY IN FOCUS

Parsewave 发布 AutomationBench Verified,修复 206 个任务判分错误

Parsewave 审查了 Zapier 的 AutomationBench 全部 600 个公开任务,经人工确认并修复 206 处判分错误。用修复前后的判分器重评 1,235 次 Kimi K3 运行时,344 次(27.9%)得到不同结果,显示基准测试的评分方式会影响模型成绩。

1 developments3 reports1 sources

Event developments

Coverage of the same occurrence shares a node; subsequent developments have their own nodes.

Parsewave releases AutomationBench Verified, fixing 206 grading errors

Parsewave audited all 600 public tasks in Zapier's AutomationBench and says it confirmed and fixed 206 grading errors. Regrading 1,235 Kimi K3 runs changed 344 verdicts (27.9%), showing how graders can affect benchmark results.

View source reports · 1 sources · 3 reports
All reports · 3

By original publication time, with each report's bookmarks and feedback preserved.

Rohan Paul (@rohanpaul_ai) AI 辅助摘要 同事件 机器翻译
编辑优先级 54/100
Parsewave releases AutomationBench Verified, fixing 206 grading bugs

Parsewave reviewed all 600 public tasks in Zapier's AutomationBench and fixed 206 grader bugs. When 1,235 Kimi K3 runs were graded again with the fixed verifiers, 344 (27.9%) received a different grade, showing how grading rules can affect assessments of agent performance.

同一事件,精选展示《Parsewave 发布 AutomationBench Verified,修复 206 个任务判分错误》
Rohan Paul (@rohanpaul_ai) AI 辅助摘要 精选 机器翻译
编辑优先级 60/100
Parsewave releases AutomationBench Verified, fixing 206 grading errors

Parsewave audited all 600 public tasks in Zapier's AutomationBench and says it confirmed and fixed 206 grading errors. Regrading 1,235 Kimi K3 runs changed 344 verdicts (27.9%), showing how graders can affect benchmark results.


推荐理由:对关注模型评测的读者,1,235 次运行中有 27.9% 因判分器修复而改变结果,是解读该基准成绩时需要留意的具体影响。
Rohan Paul (@rohanpaul_ai) AI 辅助摘要 同事件 机器翻译
编辑优先级 47/100
AutomationBench Verified released with fixes for 206 task-grading bugs

The project says it checked the graders for all 600 public tasks in Zapier's AutomationBench and fixed 206 bugs confirmed by human review. Regrading 1,235 Kimi K3 runs with the old and fixed graders changed 344 grades (27.9%).

同一事件,精选展示《Parsewave 发布 AutomationBench Verified,修复 206 个任务判分错误》