AINEWS 搜索
返回 Rohan Paul (@rohanpaul_ai)
Rohan Paul (@rohanpaul_ai)· · 原发布时间

个人代理研究发现:记忆笔记过长会增加违规则例

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 辅助摘要

研究材料称,Claude Haiku 4.5 在没有记忆时的偏好违反率为 77%,加入 10 行记忆后降至 20%,记忆继续加长后回升至约 25%。对于持续累计支出,文字规则仍有 44% 的失败率,而用代码维护总额的失败率为 0%。

正文 · 原文

– https://t.co/j5j7F4M79V

Title: "Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents"

引用或回复的背景(作者 ID 2588345408,https://x.com/i/status/2106781194616803778): More memory does not keep helping personal agents, because relevant notes help up to a point and then extra lines start making the agent miss rules.

A personal agent can't learn every user preference through memory notes, and more notes eventually make it worse, so use code for anything it must count or track and keep memory short.

Written rules work for style, like signing texts with the user's first name. For a running spending total, a stated rule still failed 44% of the time, while code that kept the total failed 0%.

Memory size has a sweet spot. With Claude Haiku 4.5, violations dropped from 77% with no memory to 20% at 10 lines, then rose to about 25% with longer memories.

Agents that rewrite their own memory from user complaints improve early, then stall. The best methods ended near 48% violations, against 7.1% when the agent was simply told every preference.

发现内容有误?提交纠错