RunningTab:通过环境侧标签实现直接工作空间交互
RunningTab: Direct Workspace Interaction with Environment-Side Tabs
Jinheon Baek · Soyeong Jeong · Yumin Choi · Dongsu Han · Sung Ju Hwang
中文摘要
许多知识工作基于工作空间中已有文件产出新成果,大语言模型(LLM)智能体正逐步接管这类任务。通过直接语料交互,智能体可在终端中搜索并读取这些文件,基于多份文件产出成果的这种方式被称作直接工作空间交互(DWI)。然而,触及文件仅是任务的一半:任务要求什么、已读取什么、列出却未打开什么,这些信息都未留下任何痕迹便滑过上下文窗口,智能体可能提取了图表却在交付报告时遗漏。为此,提出 RunningTab 框架,为直接工作空间交互配备一个环境侧标签:由环境与智能体共同维护的、针对每项任务的待办记录。具体而言,智能体登记任务需求,环境则将每次读取的文件记录为带来源信息的摘录,将列出却未打开的文件记录为候选;智能体可查看每条需求及其最佳匹配的摘录与未读候选,依据匹配内容解决需求或附带原因搁置,若尝试在仍有未关闭需求时结束任务,环境侧的完成检查会予以提示。在三个基准、三个 LLM 上的验证表明,其一致优于纯 DWI 以及在模型内部维护记录的基线方法,标签通常在查看一次后即持有交付物所需的值。
关键要点
- 01针对 LLM 智能体在 DWI 任务中需求与文件状态在上下文窗口中无痕丢失的问题,提出 RunningTab 框架
- 02(2) 通过环境侧标签记录每项任务需求、已读文件摘录(含来源)及列出未读的候选文件,由智能体与环境共同维护
- 03(3) 提供完成检查(标签机制),智能体尝试结束时若仍有未关闭需求会被提示,支持附带原因解决或搁置
- 04(4) 在三个 benchmark、三个 LLM 上验证,一致优于纯 DWI 及模型内部维护记录的基线
- 05(5) 标签通常在查看一次后即持有交付物所需的值,减少上下文窗口中的信息丢失
解读
尚无解读。
原始英文摘要
arXiv:2610.10444v1 Announce Type: cross Abstract: Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what we call direct workspace interaction (DWI). Reaching the files, however, is only half the task: nothing keeps track of what the task asks for, what has been read, and what was listed but never opened, all of which slip through the context window without leaving a trace, so an agent may extract a figure and still deliver a report without it. To address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent. Specifically, the agent adds its requirements, while the environment records every file read as an excerpt with its provenance and every listed but unopened file as a candidate; the agent can then see each requirement beside its best-matching excerpts and top unopened candidates, resolve it against matching content or set it aside with a reason, and, should it try to finish with requirements still open, receive them in a finish check. We validate RunningTab on three benchmarks with three LLMs, where it consistently outperforms plain DWI and baselines that keep the record in the model, while its tab usually holds the values a deliverable needs once seen.