Wood Chen

How Agents Build Experience and Learn: From Reflection and Long-Term Memory to Background Curation

0 comments0 views2.5k words

This post was translated from Chinese by AI. If anything reads oddly, the Chinese original is authoritative. 中文原文

For an Agent to avoid repeating the same mistakes in its next task, it needs to turn execution results into reusable lessons and retrieve them when appropriate. For applications using a fixed model API, this is usually done through external memory, retrieval, and workflow updates, without changing model parameters after every interaction.[1][2]

For example, a coding Agent used the wrong working directory on its first attempt to modify a configuration. It then fixed the problem and recorded: in this project, configuration files are located in server/config/; before making changes, confirm the repository root and read the original file, then run configuration validation afterward. When handling a similar task next time, this lesson enters the context, giving the Agent a chance to avoid its previous mistake.

What accumulates here is experience the system can use. Whether the base model becomes more capable is a separate question that requires dedicated training and evaluation.

1. Distinguish Three Types of “Learning”

Approach What changes How long it persists Common uses
Adaptation within the current session The current context, plan, and execution strategy Depends on how sessions are retained and compressed Adjusting actions based on errors, understanding temporary requirements
Experience accumulated across sessions External memory, rules, skills, and retrieval indexes Stored persistently until updated or deleted Reusing project conventions, user preferences, and effective workflows
Parameter learning Model weights or adapter parameters Stored with the model version Training relatively stable domain capabilities or output behaviors

The Reflexion paper presents an implementation that does not update weights: the Agent generates textual reflections based on task feedback and stores them in episodic memory for later attempts. ACE treats context as a continuously maintained playbook, accumulating experience through generation, reflection, and curation. These approaches show that adjusting subsequent inputs can change an Agent's task performance, though the results still depend on the specific tasks and evaluation conditions.[2][3]

In engineering practice, parameter training can be handled through separate, evaluable version iterations, while day-to-day corrections and project experience live in an editable, reversible memory system. The two can work together.

2. Why Not Turn Every Correction into a Fine-Tuning Run?

Parameter updates require training data, training jobs, and validation procedures. OpenAI's model optimization guide places evaluation, prompt improvements, and fine-tuning within an iterative workflow, rather than treating every correction in a conversation directly as training data.[4]

A failure may result from the wrong working directory, outdated documentation, a temporary network issue, or missing permissions. Execution can fail even when the model's reasoning is sound. Training these events directly into the weights risks learning incorrect associations.

Continual fine-tuning also requires checking whether existing capabilities have degraded. An empirical study of continual instruction tuning for large language models observed catastrophic forgetting affecting domain knowledge, reasoning, and reading comprehension; it also noted that training schedules can mitigate this effect. Forgetting is therefore a risk to manage, not an inevitable outcome of every fine-tuning run.[5]

There is no single figure for training time, either. Model size, the number of trainable parameters, dataset size, and hardware all affect cost. A phrase like “from a few hundred milliseconds to tens of minutes” cannot cover every situation. For day-to-day experience, external memory makes targeted changes easier: replacing a broken command or retracting an incorrect inference does not require retraining the model.

3. Accumulating Experience Requires a Complete Feedback Loop

Below is one practical memory system design. Validation and filtering are added to the workflow so that not every retrospective note becomes a long-term rule.

flowchart TD
    A[执行任务] --> B[收集结果与反馈]
    B --> C[生成候选经验]
    C --> D{验证与筛选}
    D -->|通过| E[持久化记忆]
    D -->|不通过| F[保留日志或丢弃]
    E --> G[后台合并与更新]
    G --> H[按新任务检索]
    H --> I[注入有限上下文]
    I --> A

1. Reflection Must Start from Execution Results

Useful feedback includes compiler errors, test results, API responses, explicit user corrections, and read-back checks after execution. Model-generated explanations of failures are only candidate causes; they cannot replace these results.

For example, a tool returning “file not found” may indicate an incorrect path, but it could also mean the file has not been generated, a mount is not visible, or the execution environment is different. Simply recording “always use absolute paths from now on” turns a single failure into a general rule.

A better record should define its scope:

In project A's container environment, the command's default working directory is not the repository root. Before running configuration validation, switch to the confirmed repository directory. This workflow has passed one successful validation and needs to be rechecked if the container entrypoint changes.

Successful tasks are also worth distilling: effective troubleshooting sequences, repeatable build workflows, and reliable ways to verify results can all become candidate lessons. Reflexion supports feedback of different types and from different sources, not just error messages.[2]

2. Filtering Determines Which Lessons Are Retained

I recommend classifying new lessons into three states: pending validation, validated, and invalidated. Explicitly stated long-term user preferences can be recorded directly; technical causes inferred by the model should remain pending validation until supported by tool results or subsequent tasks.

Before recording a lesson, check:

  • Can it be reused in other tasks?
  • Can it be obtained directly from existing code or documentation, and is it worth storing again?
  • Does it specify the applicable project, environment, or version?
  • Is there traceable evidence?
  • Does it conflict with existing records or contain sensitive information that should not be stored?

“A download timed out today” can usually stay in the logs. “This internal network requires a specific proxy to download dependencies” may become an environment-specific lesson, but it still needs validation.

3. Archived Lessons Need Scope and Provenance

Chat histories are useful for tracing past activity, but should not be treated as an experience repository themselves. I recommend storing reusable lessons as separate records, each expressing one fact or action rule.

An illustrative record might look like this:

id: memory-0042
kind: procedure
scope: project-a/container
trigger: 修改配置文件或执行配置校验
lesson: 先确认仓库根目录并读取原配置,再修改目标文件
evidence: task-018 的路径错误与 task-019 的成功校验
status: validated
last_verified_at: 2026-10-11
recheck_when: 容器入口或仓库目录结构变化
supersedes: memory-0017

These fields are design examples, not a fixed format from any product. Actual storage can use Markdown, a relational database, or both. Text stores the content, the database manages status and versions, and the search index handles retrieval. There is no need to force every responsibility into one storage system.

4. Define the Scope Before Retrieval

Retrieval can follow this sequence: “permission and scope filtering → candidate retrieval → ranking and deduplication → context trimming.”

First identify the user, project, device, and environment, then search for records using keywords and semantic similarity. File paths, error codes, and command names are well suited to keyword matching; lessons expressed differently but with similar meanings are suited to semantic matching. OpenClaw's memory search is one example of hybrid retrieval.[8][9]

Retrieving project A's deployment experience for project B may be worse than having no memory at all. Multi-user systems especially should not rely solely on prompts for isolation; visibility must be restricted at the retrieval-query and storage-access layers.

5. Inject Only a Few Relevant Lessons

Anthropic's article on context engineering emphasizes that context is a finite resource and that high-signal information must be selected. It describes strategies for long-running tasks, including compaction, structured notes, and multiple Agents.[1]

One practical approach is to inject only the conditions and actions needed for the current task:

相关经验:
- 适用范围:项目 A,容器环境。
- 修改配置前,先确认仓库根目录并读取原文件。
- 修改后运行项目配置校验;容器入口变化时重新确认此流程。
- 来源:task-018、task-019;状态:已验证。

Keep the full logs outside the context and read them when needed. Injected content should also distinguish explicit user agreements, validated technical lessons, and unvalidated model inferences so that they are not all treated as equally trustworthy.

4. How Claude Code, Codex, and OpenClaw Handle This

These are specific products or tool implementations; they do not imply that all Agents use the same mechanisms.

System Persistence and curation How memory is read
Claude Code Manually maintained CLAUDE.md, plus MEMORY.md and topic files in the automatic memory directory Loads rules and a length-limited automatic memory index at session startup; reads topic files on demand
Codex Extracts session experience in the background and writes it to a state database; a curation Agent then updates local Markdown memory and optional skills Uses a summary to guide retrieval, looking up handbook entries, session summaries, and related files as needed
OpenClaw Stores user information, long-term memory, and daily notes in Markdown; the memory backend builds retrieval indexes Loads some memory at startup and retrieves other relevant content through search tools

Claude Code: A Short Index and Topic Files

Claude Code's official documentation distinguishes manually maintained rules from automatic memory. CLAUDE.md holds instructions such as coding conventions and workflows; automatic memory stores preferences, corrections, and project context that cannot be inferred directly from the code.[6]

Automatic memory's MEMORY.md serves as an index. At session startup, the first 200 lines or first 25KB are loaded, whichever limit is reached first; topic files are not all loaded at startup but are read on demand. This limit applies to the automatic memory index and should not be assumed to apply to all rule files.[6]

Codex: A Database Handles Candidates, Files Hold the Curated Results

OpenAI's Codex repository exposes a two-stage memory pipeline: the first stage extracts structured memories from eligible historical sessions and writes the results to a state database; the second selects candidates, synchronizes files, and runs a dedicated curation sub-Agent.[7]

The curated output includes MEMORY.md, memory_summary.md, and optional skills/. The curation prompt requires support for progressive reading and reuse of validated workflows and checklists. The database is therefore only part of the pipeline; the final usable memory also includes local files.[7][10]

These descriptions refer to the memory implementation in the public repository. The pipeline depends on conditions such as memory being enabled, so this does not mean every Codex session performs background learning.

OpenClaw's official documentation uses Markdown files in the workspace to store memory, distinguishing between user information, long-term memory, and daily notes. The search layer supports semantic retrieval and keyword matching, allowing an Agent to search first and then read specific snippets.[8][9]

This division of responsibilities keeps memory content open to manual inspection while the index improves lookup efficiency. The documentation also notes that memory can record operational boundaries, but cannot replace hard controls such as approvals and sandboxing.[8]

5. What the “Dreaming” Mechanism Actually Organizes

As memory accumulates, duplicate, conflicting, and outdated records emerge. Background consolidation reorganizes this content rather than endlessly expanding summaries.

Anthropic has introduced a Dreams mechanism in Claude Managed Agents: it reads an existing memory store and past sessions to generate a new, consolidated memory store, merging duplicates, replacing outdated records or those contradicted by new information, and distilling new findings. The input memory store is not modified directly, and the output can be reviewed or discarded. The official documentation labels it a research preview feature.[11]

“AutoDream”-style consolidation can describe this type of background task, but each product's triggers and capabilities need to be verified separately. Dreaming is also a functional metaphor; it does not imply that the system reproduces human memory processes during sleep.

For custom-built Agents, I recommend having background consolidation produce proposed additions, deletions, and updates rather than directly overwriting the entire store:

Operation Example Information to Preserve
Merge duplicates The same build command was recorded multiple times Each source and its scope of applicability
Replace outdated records A new entry point replaces an old one The relationship between old and new versions, and why the old one is no longer valid
Handle conflicts Two environments use different deployment workflows Environment conditions; do not force a single choice
Distill workflows Multiple successful tasks follow the same sequence of checks Acceptance criteria and known exceptions

Consolidation should also use a consistent input snapshot to prevent new memories and cleanup operations from overwriting each other. High-impact rule changes should be reviewed before being published to the active memory version. This makes rollback possible if consolidation goes wrong.

6. An Implementation That Can Start Small

The following are engineering recommendations, not default configurations shared by the products above.

An initial version can consist of just a task log, a candidate experience table, an active memory table, and a retrieval interface. There is no need to introduce a complex vector database from the start. First make keyword search and scope filtering reliable, then add semantic retrieval based on missed results.

Live tasks should read only published memory versions. After a task ends, a background process extracts candidates; only validated candidates enter active memory. Consolidation builds a new version and switches to it atomically after checks pass, preventing running Agents from reading a mix of old and new rules.

任务开始:
  确定用户、项目、环境和记忆版本
  检索允许访问的活动记忆
  排序、去重,按上下文预算注入

任务结束:
  保存执行结果和明确反馈
  提取候选经验,检查来源与敏感信息
  验证可复用性,合并或替换相关记录

后台整理:
  读取一致快照,生成变更提案
  执行冲突检查与回归评估
  发布新记忆版本,保留回滚记录

When a user asks the system to “forget a piece of information,” active records, indexes, and caches must be updated together. If historical logs are retained, the next round of consolidation must also be prevented from extracting that information again.

7. How to Tell Whether an Agent Has Really Learned

An increase in the number of recorded experiences only shows that writes succeeded. To assess whether the system has improved, compare three groups: no long-term memory, manually written rules only, and curated experience-based memory. Keep the model, tools, and task budgets identical, and compare performance on similar but not identical tasks.[4]

I recommend tracking task success rate, recurrence of similar errors, tool call count, execution cost, and the number of human corrections. Also check whether injected memories are relevant or outdated, and whether they cause failures on tasks the system could otherwise complete.

Evaluation must avoid mistaking “remembering the answer” for improved capability. Test tasks should not enter the memory store in advance; otherwise, the system may simply retrieve the test results.

Reflection on errors can also hurt subsequent performance. The 2026 VRL-Bench compared several text-based memory methods under limited trial budgets, observing improvements in some settings and declines in others. This suggests that memory systems need regression testing; not every reflection should be assumed helpful.[12]

For high-impact actions such as financial transactions, production deployments, and data deletion, experience should only assist planning. Authorization, approval, and execution limits should still be enforced by the tools and application layer.[8]

Conclusion

Agents can accumulate experience across sessions by combining a fixed model with external memory. What needs maintenance is not just the memory content, but also its provenance, scope of applicability, validity status, retrieval method, and version.

A useful feedback loop is to extract candidates from real outcomes, validate and archive them, retrieve them as needed for new tasks, and use execution performance to check whether they remain effective. Background consolidation compresses duplicates and updates invalid experience; evaluation identifies rules that look reasonable but are harmful in practice.[1][2][11][12]

References

  1. Anthropic: Effective context engineering for AI agents
  2. Reflexion: Language Agents with Verbal Reinforcement Learning
  3. Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
  4. OpenAI: Model optimization
  5. An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
  6. Claude Code: How Claude remembers your project
  7. OpenAI Codex: Memory pipeline
  8. OpenClaw: Memory overview
  9. OpenClaw: Memory search
  10. OpenAI Codex: Memory consolidation prompt
  11. Claude Managed Agents: Dreams
  12. VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets

Related posts

Comments 0