AINEWS 搜索
返回 Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · 原发布时间 AI 评分54

斯坦福研究发现,多用户代理分别行动时更难协调共享资源

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 导读

一项斯坦福研究在五种前沿模型上比较了共享预算、日历等任务:每位用户各用一个代理时,整体表现不如由一个代理统筹所有人。在有争议的 token 预算任务中,Opus 5 多代理团队取得可实现价值的 30%,单个协调代理取得 64%。

正文 · 原文

该语言的正文暂不可用,当前显示已有版本。

– https://t.co/awL07ExeOk

Title: "Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams"

回复Rohan Paul@rohanpaul_ai
New Stanford paper finds that when each person's agent acts alone on a shared resource, the group does worse than 1 agent serving everyone. A shared budget or calendar is handled better by 1 agent serving everyone than by 1 agent per user, across 5 frontier models. Each agent does a sensible job for its own user. Together they overwrite each other, stall as the team grows, and with no channel they collapsed outright in 2 environments. On a contested token budget, Opus 5 teams captured 30% of the achievable value against 64% for 1 coordinating agent. Agents invented facts about other users in more than half of Claude team episodes in the group-ordering environment. Prefer 1 agent holding everyone's constraints, and if you run 1 per user, make reading peers a condition of committing.
在 X 查看回复的帖子

来源:Rohan Paul · x.com

论文
发现内容有误?提交纠错