Memory, context management and selective observation in agents

What information is retained, when is it selected, and how is selection trained?

JSON · CSV · BibTeX

WorkInformation retainedOperation and timingTraining or selection ruleTask setting
MemAgentCompact memory of previously read document chunks2.1 The MemAgent Workflow: RL-shaped Memory for Unbounded ContextsOverwrite memory after reading each chunk2.1 The MemAgent Workflow: RL-shaped Memory for Unbounded ContextsMulti-conversation DAPO; final-answer advantage shared across conversations2.2 Training MemAgent with Multi-conv RL · 2.2 Training MemAgent with Multi-conv RLDocument processing followed by answer generation2.1 The MemAgent Workflow: RL-shaped Memory for Unbounded Contexts
AgentFoldMultiscale history summaries and the complete latest interaction3.2 AgentFold’s Context: Multi-Scale State Summaries, Latest InteractionGenerate a folding directive alongside the next action3.3 AgentFold’s Response: Thinking, Folding, Explanation, ActionSupervised fine-tuning on generated and filtered trajectories3.4 AgentFold’s Training: Data Trajectory Collection · 3.4 AgentFold’s Training: Data Trajectory CollectionMulti-step web information seeking3.1 Overview
sigma-RevealTask-relevant files and their directory context3.2 $\sigma$ -Reveal Context HarnessSelect an initial workspace view before the first action3.2 $\sigma$ -Reveal Context HarnessRelevance scoring and subtree-closed selection under a token budget3.2 $\sigma$ -Reveal Context HarnessPartially observed filesystem workspaces for CLI agents3.2 $\sigma$ -Reveal Context Harness
HiMPOCompressed memory with the latest interaction3.1 Problem Setup and OverviewWrite memory before reasoning and tool use at each step3.1 Problem Setup and OverviewOutcome-based RL plus memory-specific utility and hindsight credit3.2 Local Counterfactual Utility · 3.3 Retrospective Filtering of Memory Utility · 3.4 Stabilized Token-Level Policy OptimizationOpen-domain agent tasks and compressive-memory question answeringBenchmarks and metrics.

Qualitative comparisons of methods, information requirements and task design in the cited paper versions.

References

  1. MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent. Yu, Hongli; Chen, Tinghong; Feng, Jiangtao; Chen, Jiangjie; Dai, Weinan; Yu, Qiying; Zhang, Ya-Qin; Ma, Wei-Ying; Liu, Jingjing; Wang, Mingxuan; Zhou, Hao. 2507.02259v2.
  2. AgentFold: Long-Horizon Web Agents with Proactive Context Management. Ye, Rui; Zhang, Zhongwang; Li, Kuan; Yin, Huifeng; Tao, Zhengwei; Zhao, Yida; Su, Liangcai; Zhang, Liwen; Qiao, Zile; Wang, Xinyu; Xie, Pengjun; Huang, Fei; Chen, Siheng; Zhou, Jingren; Jiang, Yong. 2510.24699v1.
  3. Learning CLI Agents with Structured Action Credit under Selective Observation. Su, Haoyang; Wen, Ying. 2605.08013v1.
  4. HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents. Yan, Jiangze; Shen, Yi; Zhang, Wenjing; Huang, Jieyun; Liu, Zhaoxiang; Wang, Ning; Wang, Kai; Lian, Shiguo. 2606.16285v2.