Wood Chen

Why AI Agent Stop Buttons Are Hard to Build: Cancellation Signals, Tool Processes, Session History, and Execution State

0 comments1 views4.2k words

This post was translated from Chinese by AI. If anything reads oddly, the Chinese original is authoritative. 中文原文

When a user clicks an AI Agent’s “Stop” button and the interface stops displaying text, has the task really stopped?

If the Agent is only generating a response, the question is relatively simple. But once it starts running commands, modifying files, or calling remote tools, stopping raises another issue: who is responsible for cleaning up work that has already started, and how should the next turn interpret its results?

This article explains the basics of cancellation and offers a set of engineering recommendations. Specific API rules are cited; state names, UI messages, and test checklists are design examples, not universal standards for any framework.

1. First, Clarify What the User Wants to Stop

Before designing a stop button, distinguish between these intentions:

User action What the system should do
Stop displaying output Stop updating the interface, but do not claim that background tasks have stopped
Cancel the current task Stop scheduling new work, handle work already started, and save execution state
Reject an operation Do not execute a tool call that has not yet been approved
Add instructions Inject a new message at an agreed safe boundary to adjust subsequent work
Undo a completed operation Perform a separate rollback or compensating operation, after assessing whether it is feasible

These are different product behaviors and should not all be represented by a single “stopped” state. Cancellation and undo are especially important to distinguish: stopping further execution should not lead users to believe that files already written or messages already sent have also been reverted.

Consider a hypothetical scenario: an Agent is modifying code, and the first two files have been saved while the third has not yet been processed. After the user cancels, a reasonable outcome is to stop further work and record which files were changed and which were not. Restoring the first two files requires backups, version history, or explicit undo steps.

2. A Cancellation Signal Is a Notification, Not Forced Termination

Cancellation typically uses a cooperative mechanism. An upper layer sends a signal; lower layers receive it, stop waiting or working, and then release resources.

JavaScript’s AbortController / AbortSignal and Go’s context.Context are examples of this mechanism. The official Go documentation explicitly states that CancelFunc does not wait for work to stop. So “cancel was called” and “all work has exited” are two different moments.[1][2]

In an Agent system, give each task its own cancellation scope and pass the signal to model requests, tool execution, and retry waits. Do not just store a global boolean on the chat page and expect every background operation to stop automatically.

Think of it in Go terms: passing ctx along the call chain does not mean that the functions being executed support cancellation. If a lower layer never checks ctx or passes it to cancellable I/O, the signal has no real effect on the work.

Pay particular attention to these checkpoints:

  • Before sending the next model request.
  • While receiving streamed output.
  • Before placing a tool call in the execution queue.
  • Before actually starting a tool.
  • During retry backoff waits.
  • Before performing a step with side effects.

The “check again before starting the tool” step is easy to miss. A task may still be active when queued, but the user may have canceled it by the time execution begins.

Checks are not a cure-all, though. Cancellation can still occur between a successful check and the start of an external operation. The closer you get to a step with side effects, the more you need execution records and result verification, rather than relying on a single check.

How Claude Code and Codex Expose Cancellation

Claude Code’s interactive documentation explicitly states that pressing Esc interrupts the current response or tool call, while completed work is retained. If messages are already queued, they may still be sent afterward. Pressing Esc at a permission prompt instead rejects that operation. The same key has different meanings in different contexts; interruption should not be equated with undoing changes or clearing the queue.[11]

For programmatic integration, the Claude Agent SDK provides explicit cancellation interfaces. The officially published TypeScript SDK 0.3.185 type definitions include Options.abortController, which cancels a query and cleans up resources. In streaming input/output mode, Query.interrupt() is a control request that interrupts the current execution and returns control.[12] The official Agent SDK overview says it provides the tools, Agent Loop, and context management that power Claude Code, so these interfaces align more closely with the Agent execution lifecycle than simply disconnecting a standard model client.[13]

These should not be treated as identical operations: use the corresponding cancellation mechanism when you need to end a query, and use the supported control interface when you need to interrupt the current execution within an ongoing session. Then handle subsequent input according to the specific version.

Codex’s public Rust source shows how cancellation reaches the tool layer: the dispatch function in tools/router.rs accepts a CancellationToken, places it in ToolInvocation, and passes it to the tool registry for execution.[14] This directly supports the approach of propagating cancellation signals down the call chain, but it does not mean that every external tool necessarily responds to cancellation.

The Codex source references in this article are pinned to commit ed59a6c1cdf5e6fc96351fd46dfc1ef8a16db385 from October 7, 2026, so readers can inspect the exact code rather than treating the evolving main branch as a permanent behavioral specification.

3. Stop Scheduling and Clean Up Existing Work Separately

I recommend handling cancellation in two phases.

The first phase closes the entry points: the current task stops dispatching new model requests and tool calls, and stops ordinary application retries. Work that is queued but has not started is marked as not executed.

The second phase handles cleanup: notify running tools to exit, allow a bounded cleanup period, save known output, and verify execution state.

The following is a suggested flow, not a fixed implementation in any SDK:

flowchart TD
    A[收到取消请求] --> B[禁止本轮新增工作]
    B --> C[向运行中的工作传递取消]
    C --> D[限时等待与必要的强制终止]
    D --> E[记录结果和未确认事项]
    E --> F[补齐会话并结束本轮]

Cleanup itself also needs a time budget. Otherwise, after a user cancels, the system may wait indefinitely for a tool that never exits, leaving the stop button ineffective.

Use a separate execution scope with a short deadline for cleanup, rather than relying on the already-canceled task scope to persist data and verify results. This scope should only allow saving results, releasing resources, and verifying state—not continuing the original task. This engineering arrangement follows from how cancellation propagates.[1]

Codex Separates “Interruption Requested” from “Turn Completed”

The official Codex App Server documentation provides turn/interrupt: the client specifies the turn to interrupt using threadId and turnId, a successful response is {}, and the turn eventually ends with an interrupted status. The turn lifecycle is reported through the turn/completed notification.[15]

This interface design suggests a practical integration guideline: after receiving a successful response to an interruption request, continue processing completion notifications and existing tool records rather than immediately destroying all session state. Interruption targets a specific turn and does not require killing the entire App Server.

Note that a turn’s interrupted status describes its execution lifecycle; it is not a declaration that “all application changes have been undone.” Remote side effects still need to be verified against actual tool results.

The Claude Agent SDK changelog also shows that queue cleanup needs separate handling: 0.3.219 added the optional cancel_queued parameter to the interrupt control request, requiring support for the corresponding capability, to also cancel queued and pending messages.[16] When implementing a stop button, you therefore need to specify whether it only stops the current turn or also cancels new instructions that have not yet been applied. Do not assume both always happen together.

4. Killing the Shell Does Not Mean All the Command’s Work Has Ended

Command tools often launch other programs through a shell. For example, a build command may actually involve a shell, a package manager, a Node process, and multiple worker processes.

The official Node.js documentation explicitly warns that terminating a parent process does not necessarily terminate its children. Successfully sending a kill signal also does not mean the process has exited.[3]

On Linux/macOS, a common approach is to create a separate process group for the tool and then terminate that group. Typically, you first request a graceful exit and wait for a while, then force termination if it has not finished. The Agent itself must not belong to the process group being terminated.

Process groups also have limits. On non-Windows platforms, Node’s detached option can make a child process the leader of a new session and process group. In other words, descendants in the process tree and members of the current process group are not the same set. Do not describe “killing the process group” as “guaranteed to kill all descendants.”[3]

On Windows, use a platform-appropriate management mechanism, such as a Job Object. It can manage and terminate processes as a group, but you still need to account for how processes join it, breakaway configuration, nested jobs, and other boundaries.[4]

From an engineering perspective, I recommend abstracting a “tool execution unit” to centrally manage startup, cancellation, waiting, output collection, and checks for leftover processes, with OS-specific implementations underneath. Do not let every tool implement its own ad hoc kill logic.

Process Group Cleanup in the Codex Source

Codex implements this as a shared helper module in utils/pty/src/process_group.rs. In the Unix code path, set_process_group creates a separate process group; terminate_process_group sends SIGTERM to the specified group, while kill_process_group uses SIGKILL. kill_process_group_by_pid first looks up the PGID, then signals the entire group.[17]

The module also handles platform differences. For example, on macOS, it attempts to signal individual group members if a group signal is rejected. These implementations show that “managing related processes as an execution unit” is real work in Agent engineering, not just an abstract recommendation.

However, these functions use best-effort semantics, and the non-Unix implementations in this file include no-op branches. This source alone does not justify claiming that Codex always uses the same shutdown sequence on every platform and execution path, much less treating a sent signal as proof that all descendants have exited. The executor still needs to implement the approach recommended here—“wait with a timeout, escalate termination if necessary, and verify the outcome”—based on the actual execution path.

5. The process may have exited while its output pipes remain open

Process management has another subtle pitfall: a descendant may inherit the stdout / stderr pipes. Even after the directly launched process exits, pipe readers may still never receive EOF.

Go's os/exec documentation describes this kind of waiting problem. WaitDelay can limit both the wait for exit after cancellation and the wait for I/O pipes that remain open after the process exits; its default value of zero imposes no such limit.[5]

A tool executor should therefore manage three things separately: whether the process has exited, whether all output has been read, and whether file descriptors have been released. Do not equate “stdout has not been fully read” with “the program is still running,” or “the main process has exited” with “everything has been cleaned up.”

Process exit codes must also be distinguished from business outcomes. Terminating a command only means execution has ended; it does not mean the command left all files unchanged.

6. Cancelling remote tools is a separate layer of the problem

A local Agent stopping its wait for a remote response does not, by itself, prove that the remote work has stopped. The remote side needs to receive the cancellation and propagate it to its own requests, processes, or background tasks.

For example, the 2025-06-18 MCP specification defines notifications/cancelled, which uses a request ID to identify the call to cancel. The specification also allows recipients to ignore the notification if the request has already completed, cannot be cancelled, or similar conditions apply, and requires both sides to handle races between cancellation and completion.[6]

The MCP Go SDK documentation likewise distinguishes between “the notification has been sent” and “the server has observed the notification”; the former does not guarantee the latter.[7]

This means a remote tool executor should ideally provide task IDs, status queries, and cancellation capabilities. If an interface only returns “connection lost,” the Agent should preserve the possibility that the outcome is unknown rather than automatically treating it as “not executed.”

MCP cancellation rules and transport behavior differ across versions, so check the negotiated protocol version when integrating. Do not treat the details of one specification version as a common implementation shared by all MCP services.

7. Complete the conversation history, but do not fabricate results

Once a tool call has entered the history, abruptly stopping execution may leave it without a result. Whether subsequent model requests accept this history depends on the specific API.

For Claude's client tools, for example, tool_use and tool_result are matched by ID. The official documentation requires tool results to immediately follow the corresponding tool-call message, with no other messages inserted between them. Parallel tool calls must also have individually matched results.[8]

But completing the history does not mean writing “the user rejected the operation” for every case. I recommend describing the actual state:

Known situation Suggested record
User denied authorization Not approved; the tool was not started
Cancelled while queued Not executed; do not claim execution failed
Terminated after starting Cancelled during execution; record known output and termination details
Tool already completed Preserve the actual completion result, even if the turn was subsequently cancelled
Remote outcome cannot be confirmed Outcome unknown; verify before continuing

These descriptions should be mapped to tool-result formats accepted by the target API. For platform-hosted tools executed on the server, do not fabricate client-side results either; Claude's documentation explicitly distinguishes how client and server tools are handled.[8]

There is another case: a streaming response stops halfway through, before the tool arguments form a valid call. I recommend keeping it in debug logs rather than forcibly parsing and executing it to “complete the history.” A complete execution record and a history that can be sent back to the model can be two different views of the data.

Codex repairs missing tool outputs, but this does not verify business outcomes

Codex's context_manager/normalize.rs contains an explicit repair function: ensure_call_outputs_present. In the commit cited here, it checks the correspondence between calls and outputs, constructs outputs containing aborted for types such as FunctionCall and CustomToolCall that lack results, and inserts them after the corresponding calls.[18]

This is a direct source-code example of “tool results must not be missing without explanation.” However, it is a fallback used when preparing model context; comments in the file also note that synthetic outputs may be used only for prompt normalization and not persisted. It should not be described as “every cancellation fully preserves the actual execution result.”

aborted can fill the gap in the protocol structure, but it cannot tell us how many lines in a file were changed or whether an email was submitted. Actual tool records should still preserve known output, the reason the outcome is unknown, and recovery steps.

For Claude, the Messages API's call-and-result rules are covered by the official documentation cited earlier in this section. The Agent SDK 0.3.216 release notes also added tool_result_meta, allowing integrators to distinguish denial, interruption, cancellation, and other cases without relying solely on matching result text.[8][16] This supports a state design that does not label all non-success outcomes as user denials, but it is not source-code proof of every internal history-repair path in Claude Code.

8. A task can be cancelled while an individual tool still succeeds

I recommend storing task status and tool status separately.

Suppose a turn performs three steps in sequence: read the project, write the configuration, and deploy the service. If the user cancels before deployment, the overall task can be marked as cancelled, but reading and writing have already completed and should not be relabeled as cancelled.

A task can have lifecycle states such as running, cancelling, and cancelled, while each tool separately records not started, running, succeeded, failed, cancelled during execution, or outcome unknown. The names can change; the information must not be lost.

For tools with side effects, I also recommend separately recording “whether side effects have been verified.” For example, the process may have been terminated, but whether the configuration file was fully written still needs confirmation. A single status field is often insufficient to express this state.

When cancellation and success arrive at the same time, preserve the result based on verifiable execution facts. Even if a transport protocol requires the client to ignore late responses, that should not be taken to mean no business side effects occurred.[6]

9. Cancellation is not rollback, and retries are not inherently safe

Consider a hypothetical email-sending tool: the request reaches the server and the email is submitted, but the client cancels before receiving the response. Retrying immediately risks sending the email twice.

The recommended recovery sequence is to first query the status using an operation ID or business record, then decide whether to retry. Interfaces that support idempotency can reduce the risk of duplicate execution.

Stripe's official documentation provides a concrete example: a request carries an idempotency key, and subsequent requests with the same key can return the previously stored result, avoiding duplicate creates or updates. However, keys have a retention period, and parameters must match; this is not a universal deduplication guarantee that lasts forever.[9]

For Agent tools, I recommend preserving the idempotency key associated with “the same business operation.” If a new key is generated every time a task resumes, the server may treat each request as a new operation.

File changes need their own safeguards. Options include recording versions before execution, preserving diffs, making changes in an isolated workspace, and verifying afterward. Before rolling back, also check that no one else has made further changes to the files, to avoid overwriting newer work with old content.

This is recovery and compensation design, not a capability automatically provided by cancellation.

Claude Code's rewind also has explicit boundaries

Claude Code treats interruption and rewind as separate operations. Its interaction documentation explains that pressing Esc twice when the input box is empty opens the rewind menu to restore or summarize earlier code and conversation state; this differs from the interruption triggered by a single Esc.[11]

The official checkpointing documentation also states that files modified through Bash commands are outside the scope of checkpoint tracking and cannot be restored with rewind; checkpoints track changes made directly by Claude's file-editing tools.[19]

This provides a concrete product example of “cancellation is not rollback”: even a dedicated rewind feature has limits, so an ordinary stop button should not promise that arbitrary operations can be undone. Effects caused by calls to external services require those services to provide their own undo, query, or compensation capabilities.

10. Steering changes direction; it does not replace cancellation

When a user says “add another explanatory section later,” they are usually adding a requirement; when they say “don't send it,” the relevant operation needs to stop. The product should clearly distinguish these cases.

One application-side implementation is to queue new messages and inject them after collecting tool results, before the next model call. This boundary is easy to manage, but not every system must wait for the entire batch of tools to finish.

OpenAI's Mid-turn steering documentation distinguishes between a message being accepted, queued, and actually applied. It also explains that Steering does not rewrite content already emitted, undo earlier actions, or cancel tools that have already started.[10]

I therefore recommend that the interface show separate states for “additional instructions queued” and “applied to subsequent execution,” rather than immediately displaying “instructions applied” as soon as a message arrives. If the user's request affects deletion, sending, or deployment operations that have not yet run, first prevent those operations from starting, then decide how to adjust the plan.

Codex explicitly distinguishes Queue, Steer, and Interrupt

OpenAI's official Codex usage article distinguishes Queue from Steer: Queue waits for the current response to finish and sends the new input as the next turn; Steer injects guidance into work in progress.[20] Neither should be confused with Interrupt.

Codex App Server exposes this distinction through its API: turn/steer appends user input to an active turn without starting a new one; expectedTurnId must match the current turn, and the request fails if no turn is active. To request cancellation, use turn/interrupt, introduced earlier.[15]

This shows that Steering is more than “sending another message” in a chat input box. The system needs to know which turn the message belongs to, whether it has been accepted, and whether it should be queued as a new task or influence the current work.

The Claude Agent SDK's Streaming Input documentation also describes long-lived sessions, queued messages, interrupts, and context retention across turns.[21] But this does not mean the two providers handle these operations at exactly the same points, and Codex's turn/steer interface, the OpenAI Responses API's Mid-turn steering, and Claude's input queue must not be described as a single protocol.

11. A practical cancellation design

Given the boundaries above, the implementation requirements can be organized into the following checklist. These are engineering recommendations:

  1. Give each task turn its own ID and cancellation scope so that stopping one session does not affect another.
  2. Once cancellation is requested, prevent new work first, then wind down work already in progress.
  3. Support cancellation for queues, model requests, retry waits, and tool execution.
  4. Allow repeated cancellation calls, and prevent duplicate cleanup and result writes.
  5. Manage local commands through a unified execution unit, handling process groups or Job Objects as appropriate for the platform.
  6. Put a time limit on cleanup, and record any resources whose release could not be confirmed.
  7. Write an execution record before starting a tool, and save the actual result when it finishes; if writing the record fails, do not continue claiming that the state has been saved.
  8. Have the session adapter fill in results according to the API's rules, without conflating cancellation, denial, and failure into a single state.
  9. Verify write operations with unknown outcomes before resuming; for tools that support idempotency, reuse the original operation key.
  10. Give Steering its own message queue and application status, rather than reusing the cancellation button's semantics.

The UI should also reflect actual progress. For example, “Stopping,” “Local process exited,” and “Remote execution status awaiting confirmation” are more accurate than a generic “Cancelled” notification. Unverified work can remain in the recovery record; there is no need to fabricate a definitive result just to end the wait for the current turn.

Ultimately, a stop button's reliability depends on what remains after cancellation: whether work is still running, whether any changes remain unconfirmed, whether the session can continue, and whether the next recovery will execute anything twice.

References

[1] Go: context package

[2] AbortController and AbortSignal

[3] Node.js: child_process

[4] Microsoft: Job Objects

[5] Go: os/exec package

[6] MCP 2025-06-18: Cancellation

[7] MCP Go SDK: LifeCycle

[8] Claude: Handle tool calls

[9] Stripe: Idempotent requests

[10] OpenAI: Mid-turn steering

Primary sources for Claude Code, Claude Agent SDK, and Codex

The sources below cover the product behavior, SDK interfaces, and specific source code discussed in the article. The documentation and source code illustrate existing implementations; they do not imply that all versions and third-party tools provide the same guarantees.

[11] Claude Code: Interactive mode, Esc interrupts, permission denial, and rewind

[12] Official Anthropic release package: Claude Agent SDK 0.3.185 type definitions, Options.abortController and Query.interrupt

[13] Claude Agent SDK: Overview, its relationship to Claude Code's runtime capabilities

[14] OpenAI Codex source code: tools/router.rs, CancellationToken propagation

[15] OpenAI: Codex App Server, turn/interrupt, turn/steer, and the turn lifecycle

[16] Anthropic: Claude Agent SDK TypeScript changelog, cancel_queued in 0.3.219 and tool_result_meta in 0.3.216

[17] OpenAI Codex source code: process_group.rs, process group signals and platform differences

[18] OpenAI Codex source code: context_manager/normalize.rs, patching missing tool results

[19] Claude Code: Checkpointing, rewind capabilities and limitations for Bash changes

[20] OpenAI: Mastering remote engineering work from your phone, the difference between Queue and Steer

[21] Claude Agent SDK: Streaming Input, persistent sessions, message queues, and interrupts


Originally published on the SunAI forum.

For additional terminology and implementation details, see the comments on the original post.

Last updated 2026-10-09

Related posts

Comments 0