正文 · 原文
该语言的正文暂不可用,当前显示已有版本。
Read their detailed technical blog on this
https://t.co/5YcB25Kamj
自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。
Parsewave 审查了 Zapier AutomationBench 的全部 600 道公开任务,发现并修复 206 个评分器错误。用修复前后的评分器重新评判 1,235 次 Kimi K3 运行时,344 次(27.9%)得到不同成绩,显示评分规则会影响对智能体表现的判断。
该语言的正文暂不可用,当前显示已有版本。
Read their detailed technical blog on this
https://t.co/5YcB25Kamj
Agent benchmarks need audits of the graders, not only the tasks.
Parsewave reviewed all 600 public tasks in Zapier's AutomationBench, which frontier labs cite on their model cards.
It used agents to write convincing wrong answers, then had humans confirm which graders actually failed: 206 did.
AutomationBench Verified fixes every one.
On 1,235 Kimi K3 runs, the fixed graders gave a different verdict 27.9% of the time.
The clearest case is task 813. The grader checked the Salesforce notes but never looked at the DocuSign template. A submission that sent four contracts on the Standard template, instead of the required GDPR, HIPAA, SOC2 and Enterprise ones, passed 13 of 13 checks and scored 1.0. After the fix it scores 0.09.
在 X 查看回复的帖子
AutomationBench Verified is out.
We went through all 600 public tasks in @Zapier's AutomationBench and checked every verifier.
• Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed.
• We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade
• Where verifiers were too strict, pass rate went from 18.8% to 43.8%
• Where they were too lenient, it dropped from 60.2% to 49.7%
Audit and dataset links below.
在 X 查看上下文