正文 · 原文
该语言的正文暂不可用,当前显示已有版本。
@shaped @zapier This is solid.
A benchmark score is only as honest as the code that grades it.
自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。
项目方称,他们检查了 Zapier 的 AutomationBench 全部 600 个公开任务的评分程序,并修复人工复核确认的 206 个错误。用修复前后的程序重新评判 1,235 次 Kimi K3 运行后,344 次(27.9%)的评分发生变化。
该语言的正文暂不可用,当前显示已有版本。
@shaped @zapier This is solid.
A benchmark score is only as honest as the code that grades it.
AutomationBench Verified is out.
We went through all 600 public tasks in @Zapier's AutomationBench and checked every verifier.
• Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed.
• We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade
• Where verifiers were too strict, pass rate went from 18.8% to 43.8%
• Where they were too lenient, it dropped from 60.2% to 49.7%
Audit and dataset links below.
在 X 查看回复的帖子