AINEWS 搜索
返回 Rohan Paul (@rohanpaul_ai)
Rohan Paul (@rohanpaul_ai)· · 原发布时间

SkillGym 研究用通过检验的技能执行记录训练 AI 智能体

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 辅助摘要

这项上海人工智能实验室研究将人工编写的技能文件转成带有代码检验器的沙盒任务,再用通过检验的执行记录训练模型。在 Claude Code 中,Qwen3.5-35B-A3B 的 Terminal-Bench 2.1 成功率从 39.33% 升至 58.43%;未加载技能文件时,它在 SkillsBench 上得分 26.81%,高于加载技能文件的基础模型的 23.34%。

正文 · 原文

– https://t.co/dr1jJMbMEy

Title: "SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving"

引用或回复的背景(作者 ID 2588345408,https://x.com/i/status/2107236641006129333): New Shanghai AI Laboratory paper finds that training on verified runs of human-written agent skills makes a model a better agent, even without the skill files.

Prompt-time skills depend on retrieval and instruction following, and fine-tuning on verified skill runs reduces that dependence.

Turning each skill file into a sandboxed task with a pass-or-fail checker produced training data that lifted Terminal-Bench 2.1 success by 19.10 points in Claude Code.

Skill files usually sit in the prompt, so they only help if the agent finds and follows them.

SkillGym turns each skill into a sandboxed task with a code checker, then trains on the runs that pass.

In Claude Code, Qwen3.5-35B-A3B jumped from 39.33% to 58.43% on Terminal-Bench 2.1. With no skill files, it scored 26.81% on SkillsBench, beating the base model with skills at 23.34%.

Loading the skills on top still helps, lifting it to 51.47%.

发现内容有误?提交纠错