Wood Chen

How AI Agents Plan: From Task Decomposition and Dynamic Replanning to Search and State Machines

0 comments1 views2.7k words

This post was translated from Chinese by AI. If anything reads oddly, the Chinese original is authoritative. 中文原文

Give an AI Agent a goal, such as “analyze a batch of customer feedback, identify the main issues, and generate a report,” and it needs to decide which data to read first, how to group it, when to gather more information, and what counts as completion.

These decisions form the Agent's planning process. Turning the goal into a few to-do items covers only part of it. Once execution begins, the Agent also needs to handle dependencies, failures, changes in the environment, and permission boundaries.

To understand Agent Plan, we can explore four questions: how plans are generated, how execution progresses, how deviations are corrected, and how far the system is allowed to go. Planning before execution, feedback and replanning, multi-path search, and SOPs and state machines each offer different approaches. They can be combined; they are not mutually exclusive categories.[1][2]

1. What Questions Should a Plan Answer?

For “analyzing customer feedback,” a rough plan might be: read the feedback, summarize the issues, and generate a report.

This plan helps show progress, but it is not enough to drive execution directly. Which time period should the data cover? Should the grouped results retain the original feedback IDs? Can an issue without enough evidence be included in the report? If permissions are insufficient, should the system retry or pause?

When designing an executable plan, each task can record:

  • Its specific goal and input sources.
  • The prerequisite tasks it depends on.
  • The tools and permissions it is allowed to use.
  • Expected outputs and acceptance criteria.
  • How to handle failure.

For example, the following is an illustrative task structure, not a fixed interface from any particular framework:

{
  "id": "cluster_feedback",
  "depends_on": ["load_feedback"],
  "input_ref": "artifacts.feedback",
  "allowed_tools": ["analyze_text"],
  "expected_output": "带原始反馈编号的问题分组",
  "acceptance": [
    "每条输入都有明确归属或被标记为未分类",
    "每组结论能追溯到原始反馈"
  ],
  "on_failure": "保留已完成结果,提交失败原因"
}

The acceptance criteria still need to be enforced by an executor or validator. Writing acceptance criteria into JSON does not mean the system can already check them.

LLMCompiler, introduced by LangChain, puts tools, arguments, and dependencies into a task graph. A scheduler then executes tasks once their dependencies are satisfied. This shows that a plan's output can be a data structure for programmatic scheduling, rather than just a natural-language checklist.[2]

2. Plan-and-Execute: Separating Planning from Execution

The basic structure of Plan-and-Execute is for a Planner to generate a multi-step plan and hand it to an Executor. A single execution task may involve multiple tool calls or be handled by a local Agent.[2]

For example, customer feedback analysis could start with this plan:

  1. Retrieve feedback for the specified time range.
  2. Remove duplicate records while retaining the original IDs.
  3. Group the issues.
  4. Summarize frequencies and select representative examples.
  5. Generate a report and check whether the conclusions are supported by data.

This works well when the goal is clear and the task can be broken down in advance. An explicit plan also makes it easier to check for missing steps.

Planner and Executor are two responsibilities; they do not have to use different models. A more capable model can handle planning while a smaller model handles simple subtasks. The same model can also do both, or some steps can be delegated to conventional programs. Cost savings depend on whether the subtasks can actually be completed using cheaper execution methods.[2]

The trade-off is that upfront planning takes time. If the plan is based on incorrect assumptions, later steps will be affected. The more a task depends on information not yet available, the less suitable it is to lock down every detail in advance.

One approach is to define stage-level goals first and refine the next stage's actions after obtaining results. For example, read the codebase and locate the issue first, then plan changes based on its actual structure, rather than guessing every file path before reading the code.

How It Relates to ReAct

ReAct interleaves reasoning and action, using feedback from the environment to decide the next step. It also includes planning and adjustment, but usually does not produce a complete plan as a separate artifact.[3]

The two can be nested: an outer Planner organizes “locate the issue, make changes, test, and deliver,” while an inner Executor uses ReAct to repeatedly read files, search for symbols, and inspect call chains during the investigation stage.

The choice between an explicit plan and step-by-step decisions therefore depends mainly on how much advance coordination the task requires. There is no need to treat them as capabilities where you can choose only one.

3. Feedback and Dynamic Replanning: Execution Results Change the Path Ahead

When a plan is created, the Agent has limited information. After tools run, new facts may invalidate the remaining plan.

For example, the original plan may call for reading a full month of feedback, while the data source retains only the most recent seven days. The system should first record the data gap, then decide whether to use a backup source, narrow the scope of analysis, or wait for more data. Continuing to generate a “monthly report” would treat an assumption in the plan as a fact.

LangChain's Plan-and-Execute example already includes a replanning step: based on execution results, it decides whether to finish or generate a follow-up plan. Dynamic replanning can therefore be added directly to Plan-and-Execute without creating a completely different architecture.[2]

flowchart TD
    A[目标与约束] --> B[生成或更新计划]
    B --> C[执行任务]
    C --> D[验证结果]
    D -->|继续| C
    D -->|需要调整| B
    D -->|条件满足| E[交付]
    D -->|需要授权或补充信息| F[暂停等待]

Retries, Replanning, and Reflection Address Different Problems

If an API temporarily times out and the original approach remains sound, a limited number of retries may be appropriate. If the data source does not exist, further requests are pointless and the plan needs to change. To help later attempts avoid similar mistakes, the system can also save a short summary of lessons learned from the results.

Reflexion studies this last mechanism: an Agent generates written reflections based on task feedback, stores them in memory, and uses them to influence subsequent attempts. This process does not update the model's weights.[4]

In engineering practice, errors can be handled separately: retry transient failures, replan when an approach no longer works, and pause when permissions are insufficient or the goal is unclear. Do not let the model respond to every error with the same “try again.”

Reflection Needs Verifiable Evidence

Asking a model to evaluate its own answer can provide suggestions for improvement, but whether the problem has actually been solved still depends on external results.

Code tasks can run tests; data tasks can check row counts, fields, and statistics; publishing tasks can read back saved content. Open-ended content can be assessed with explicit scoring rubrics and human spot checks. Anthropic's Agent evaluation guide treats programmatic, model-based, and human judgments as different tools, and notes that model-based judgments need calibration.[5]

The frequency of reflection should also be designed around cost. One option is to use lightweight checks for routine steps and call a model for evaluation at the end of a stage, when a key assumption changes, or when a task fails. Adding a lengthy round of reflection after every file read will slow down simple tasks.

4. Multi-Path Search: Keeping Several Possible Solutions in Play

Some tasks are difficult because the first approach is likely to be wrong. For mathematical derivations, combinatorial problems, or code fixes with clear test criteria, it may be worth generating multiple candidates and progressively filtering them.

Tree of Thoughts has the model generate and evaluate intermediate candidates, then uses search methods to explore and backtrack. LATS combines tree search, actions, environmental feedback, and reflection in a single framework.[6][7]

For example, when fixing a parser bug, the system could first propose three hypotheses: incorrect input preprocessing, a missing state transition, or an incorrect boundary check. Each path gathers evidence separately. Paths that do not match the failing test case are eliminated, while the remaining paths are refined further.

What needs designing here is the search process, not just a prompt asking the model to “give three options”:

  • What does a node represent: a hypothesis, a partial solution, or an environment state?
  • How are candidates generated, and how are duplicates identified?
  • What is the basis for scoring, and can tests be run?
  • How many branches are retained at each step, and how long can the search run?
  • When does it stop, and how does it fall back if it fails?

These questions determine whether the search is worth its extra overhead.

Search Graphs, Task Graphs, and Flowcharts Are Not the Same Thing

A search graph records candidate solutions, and only one path may be selected for execution. A task dependency graph records work that must be completed, such as retrieving data first and then computing statistics for different fields in parallel. A state transition graph records the process the system permits, such as requiring a request to pass review before execution.

All of them can be drawn as graphs, but they mean different things. LangGraph provides graph-based orchestration, but that does not mean an Agent built with it automatically supports multi-path search. Candidate generation, scoring, and exploration logic still need to be designed.[2][8]

A Branch Score Is Not Its Actual Probability of Success

When a model says a path is “more promising,” that is only an evaluation signal. Candidate generation may miss good solutions, and the evaluator may make incorrect judgments. With a limited search budget, the goal should be to improve the chance of finding an acceptable solution, not to promise a global optimum.[6][7]

Search also requires an environment suited to exploration. Drafts, sandboxed code, and rollback-capable states are relatively easy to experiment with. Real payments, emails, and data deletion should not be repeated just to compare approaches. The system can search for an action plan first, then use permission and approval mechanisms to control actual actions.

5. SOPs and State Machines: Encoding Business Boundaries in Software

Some processes are already well defined. For example, a customer support ticket must first have its information verified, then have its issue type identified, and finally enter an allowed handling branch. Here, business rules should be expressed first, before deciding which nodes need a model.

Anthropic distinguishes workflows driven by predefined code paths from agents whose processes are dynamically directed by a model. SOPs and state machines are usually closer to the former, though autonomous Agents can also be embedded in individual nodes.[1]

Below is an illustrative ticket workflow:

flowchart LR
    A[接收工单] --> B[核实信息]
    B --> C{信息完整}
    C -->|否| D[等待补充]
    D --> B
    C -->|是| E[分类与拟定处理方案]
    E --> F[规则检查或人工审核]
    F -->|通过| G[执行允许的操作]
    F -->|未通过| H[人工处理]

The model can help classify issues, extract fields, and draft a handling plan; the program is responsible for checking required fields, verifying identity, restricting the scope of actions, and controlling state transitions.

LangGraph's official overview covers persistence, human intervention, and the ability to combine deterministic steps with Agent steps. It is well suited to expressing these workflows, but compliance rules still need to be enforced through business logic, permission systems, and audit mechanisms.[8]

The trade-off of fixed workflows is maintenance cost. Rule changes require workflow updates, and uncovered cases need human handling or explicit exception branches. The system needs a way to handle exceptions, rather than letting the model bypass rules on its own.

6. How to Choose and Combine the Four Approaches

Approach Main problem it solves Scenarios to try first Trade-offs
Plan-and-Execute Coordinate multiple subtasks in advance Long tasks with clear goals and decomposable steps The initial plan may rely on incorrect assumptions
Feedback and replanning Adjust subsequent actions based on new facts Research, troubleshooting, and tasks in changing environments Validation and replanning add overhead
Multi-path search Avoid committing to one solution too early Difficult problems with clear evaluation criteria and room for exploration Branch growth, scoring errors, and search costs
SOPs and state machines Restrict permitted business workflows Stable business processes with clear approval and permission boundaries Workflow maintenance and exception handling

This table offers guidance for choosing an approach, not a performance ranking. Evaluate a system on real tasks before deciding which layer of mechanisms to add.

For example, a system that handles customer feedback could use a fixed workflow to restrict data access; have a Planner break down tasks within the “analysis” node; process independent groups in parallel; replan when it finds insufficient samples; and finally use code to check reference numbers before a human reviews the report.

Neither multiple models nor multiple Agents are prerequisites. Start by completing tasks with one model, a set of tools, and an execution loop, then split roles based on actual bottlenecks. This makes problems easier to pinpoint. It also aligns with Anthropic's recommendation to start with simple, composable approaches.[1]

7. Production Systems Need Safeguards Beyond Planning

Save Execution State, Not Just a To-Do List

I recommend recording plan versions, task states, references to tool results, and failure reasons separately. Task states should at least distinguish between not started, running, succeeded, failed, awaiting input, and canceled.

When replanning, retain verified outputs and modify only the affected downstream tasks. Otherwise, a single failure could cause an entire stretch of work to run again.

LangGraph saves execution state through checkpoints, but resuming execution may re-enter a node. Do not assume that every external operation executes exactly once. The official documentation on resuming after an interrupt explains that logic runs again from the beginning of the node.[9]

Protect Actions with Side Effects Separately

One failure scenario worth testing in advance is this: an external system has accepted a request, but the local system has not received a response. Retrying immediately could send something twice or create duplicate resources.

The solution should match the external API's capabilities: use idempotency keys where possible; when idempotency is not supported, query the operation's status before deciding whether to retry. You also need an explicit strategy for maintaining consistency between checkpoint records, business operation records, and external state.

Enforce Permissions and Budgets at Runtime

Writing “requires approval” in a plan is not enough; unapproved actions must also be blocked at the tool-call entry point. Replanning must not become a way to expand authorization.

I recommend setting limits on tool calls, timeouts, failure retries, and search budgets. When a limit is reached, save the current outputs and explain what remains unfinished to avoid infinite loops.

Tool output can inform factual judgments, but it should not be treated as new system authorization. Retrieved web pages, emails, and logs may all contain text that asks for behavioral changes. The executor should still stay within its original permissions.

8. How to Tell Whether Planning Actually Helps

Evaluation should compare both final results and the execution process. A plan that looks complete does not mean the task was completed better.

Prepare a set of tasks with clear acceptance criteria, then introduce several types of failures: missing data, tool timeouts, permission denials, resuming midway through execution, and tools succeeding without their results being recorded locally.

I recommend tracking task success rate, total elapsed time, cost per task, ineffective tool calls, human intervention, and incorrect actions. Use programmatic checks for result correctness, and scoring rubrics plus human spot checks for open-ended quality. Anthropic's evaluation guide also recommends separating capability evaluations from regression evaluations to avoid breaking existing capabilities while improving one class of tasks.[5]

If explicit planning does not improve success rates but significantly increases execution time, reduce planning granularity or return to a simpler loop. If most errors come from invalid tool arguments, improving tool interfaces and validation is more direct than adding more planning roles.

Ultimately, Agent Plan is about organizing and completing work. Explicit decomposition helps coordinate tasks, feedback mechanisms correct faulty assumptions, search compares candidate solutions, and state machines enforce business boundaries. Putting this into practice also requires verifiable results, recoverable execution state, and permission controls that the model cannot bypass.

References

[1] Anthropic: Building effective agents

[2] LangChain: Plan-and-Execute Agents

[3] ReAct: Synergizing Reasoning and Acting in Language Models

[4] Reflexion: Language Agents with Verbal Reinforcement Learning

[5] Anthropic: Demystifying evals for AI agents

[6] Tree of Thoughts: Deliberate Problem Solving with Large Language Models

[7] Language Agent Tree Search

[8] LangChain: Open source agent stack

[9] LangGraph: interrupt reference documentation

Related posts

Comments 0