A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain 文章

ArXiv CS.AI2026-07-29PAPERen作者: Debjyoti Paul

详细信息

来源站点
ArXiv CS.AI
作者
Debjyoti Paul
文章类型
PAPER
语言
en
发布日期
2026-07-29

摘要

arXiv:2607.25415v1 Announce Type: new Abstract: Production LLM agents are increasingly assembled from a frozen model wrapped in a harness: a prompt template, a tool set, a memory/retrieval layer, a planning strategy, and a verification policy. Two 2026 systems, Meta-Harness (Lee et al., 2026) and HyperAgents (Meta AI, 2026), show that this harness can itself be optimized or even self-rewritten by an agentic proposer -- at the cost of either an expensive code-search loop or unconstrained self-modifying code, neither of which is auditable or usable with a fully black-box model API. We take a narrower, more constrained position: treat the harness as a small, fixed, human-legible action space and learn a policy over it online with classic sample-efficient reinforcement learning (an $\epsilon$-greedy contextual bandit and REINFORCE), scored against a multi-objective reward (task success, verifier score, policy compliance, cost, latency, and an unsupported-claim penalty).

相关事件

暂无数据

相关公司查看全部 (1)

M
Meta AICOMPANY

相关人物

暂无数据