What to Do When Agent Tool Calls Fail: An Engineering Guide to Error Classification, Timeout Handling, and Idempotent Retries
This post was translated from Chinese by AI. If anything reads oddly, the Chinese original is authoritative. 中文原文
When an Agent tool call fails, you cannot just add “retry three times on failure” to its instructions. Retrying a lookup might only cost a few extra seconds; retrying an email, order creation, or deployment might actually perform the operation twice.
A more reliable approach is to first determine where the failure occurred and whether the operation already happened, then decide whether to wait and retry, change the parameters, query the status, or stop and hand it over to a human.
The official guides from OpenAI and Anthropic both call for exit conditions, failure boundaries, and human handoff mechanisms for Agents. Backoff, process cleanup, and duplicate prevention also require established distributed systems engineering practices.[1][2]
1. First distinguish: did the model request fail, or did tool execution fail?
A tool call usually involves several steps: the application requests the model, the model returns a tool name and arguments, the application executes the tool, and the result is passed back to the model.
An “Agent error” can therefore occur in at least these places:
| Failure location | Common symptoms | What to check first |
|---|---|---|
| Model API | Rate limits, insufficient balance, oversized requests, service errors | API error codes, request ID, quota, and request size |
| Tool call format | Unknown tool, missing fields, type mismatches | Tool definitions, argument validation, protocol format |
| Tool runtime environment | Network disconnection, hung commands, unavailable dependencies | Connection status, process status, execution logs |
| Business operation | Insufficient inventory, order status restrictions, missing permissions | Business rules and actual state |
| Result delivery | Execution completed, but the response was lost or the session interrupted | Execution records, task ID, whether the business object was created |
OpenAI’s Function calling documentation explicitly places tool execution on the application side. The model proposing a call does not mean the model API has performed the business operation for the application.[3]
This distinction directly affects recovery: requesting the model again is no substitute for checking whether the previous order was created; incorrect tool arguments should not be addressed by repeatedly calling the same external service.
2. Assess error types and operation risks separately
“Network error,” “invalid arguments,” and “timeout” describe failures; “does this charge money or modify data?” describes the operation. These are not four mutually exclusive error categories.
The same network timeout requires different handling when reading the weather versus creating a payment.
Use the following table to decide the first step. Which errors can actually be recovered automatically should depend on the tool implementation and service API contract, not just the HTTP status code.[4][5]
| Situation | Default handling | What not to do |
|---|---|---|
| Temporary rate limiting or brief service unavailability | Retry with bounded backoff after confirming that repetition is safe | Immediately send identical requests in succession |
| Missing fields or invalid arguments | Return a specific error so the model can correct and resubmit | Replay the original arguments unchanged |
| Insufficient balance or permissions | Stop the affected operation and explain that external action is needed | Have the model repeatedly try to bypass restrictions |
| Timeout or lost response | Check execution status first | Assume “it did not execute” |
| Write operation with an unknown outcome | Query status, or recover using existing idempotency guarantees | Retry with a new request identifier |
| Recovery budget exceeded | Stop, fall back, or hand over to a human | Keep looping until “success” |
429 is an easy example to misinterpret. OpenAI’s error documentation distinguishes request rate limits from quota and billing limits; the latter require changes to quota or billing conditions, and waiting a few seconds will not restore access.[4]
3. Transient errors can be retried, but need limits
For requests that the API explicitly allows to be safely repeated, transient failures are suitable for retries with backoff. AWS engineering guidance recommends exponential backoff, random jitter, and limits on retry counts or total elapsed time to avoid adding pressure to an already overloaded service.[5]
A practical configuration should include:
- Which errors can be retried and which require an immediate stop.
- The timeout for each call.
- The maximum number of attempts and the deadline for the entire task.
- The maximum wait interval and randomization rules.
- How to handle wait guidance from the server.
- What status to return and who takes over when the retry budget is exhausted.
For example, the wait time could be random(0, min(cap, base × 2^n)). This is an implementation choice, not a universal formula that every tool must use. When the server returns a valid Retry-After, schedule the wait according to the API contract and the remaining task time; if the required wait exceeds the task budget, end the current run or schedule recovery for later rather than keep occupying an execution slot.
Also check whether the SDK already retries automatically. The official OpenAI Python SDK documentation states that some connection errors, 408, 409, 429, and server errors are automatically retried by default. This applies to model API requests; it does not imply that your payment tool is also safe to retry.[6]
If the SDK, tool wrapper, and outer Agent layer each allow up to three attempts, three nested layers could produce 27 underlying requests in the worst case. This number is a configuration example, not a framework default. AWS explicitly warns against stacking retries across multiple layers and recommends choosing one appropriate layer to control them centrally.[5]
4. Let the model fix invalid arguments; stop on permission errors
Deterministic errors have one thing in common: if the conditions remain unchanged, retrying the same request usually will not help.
But “return it to the model” does not mean “the model can fix everything.”
For a missing date or an incorrect enum value, provide the field requirements so the model can correct it. Insufficient balance, an account without permission, or missing user authorization requires stopping the affected operation or asking the user to take action. The model should not switch accounts, expand permissions, or bypass approval on its own to complete a task.
Anthropic recommends writing tool errors as actionable feedback that identifies the specific problem and allowed next steps, rather than returning only failed or a large stack trace.[7]
For example, after a failed calendar event creation, you could return: “The end time is earlier than the start time; no event was created. Please check both time fields.” This helps the model recover more effectively than “invalid arguments.”
For input formats, OpenAI recommends using strict mode to constrain tool arguments so calls match the declared schema. But schema compliance does not imply business validity: an amount can have the correct type but exceed the allowed limit; an order ID can have the correct format but belong to another user. The application still needs to validate business conditions and permissions.[3]
Error markers are not standardized across platforms either: Claude client tool results can use is_error: true, MCP tool execution errors use isError: true, and OpenAI function results are returned according to its API format. These fields are not interchangeable.[8][9][3]
5. After a timeout, check execution status first
A timeout only means the waiting side did not receive a result in time; it does not mean the operation did not happen. AWS retry guidance specifically warns that side effects may already have occurred when a call times out or fails.[10]
Local commands and remote APIs need different handling.
Local commands: canceling the wait does not necessarily end the process
In Python, for example, a timeout in Popen.communicate(timeout=...) does not automatically kill the child process; the official documentation requires the application to clean up the process and finish collecting output after the exception. Another interface, subprocess.run(timeout=...), behaves differently, so the behavior of one interface cannot be generalized to all executors.[11]
For short-lived commands, a clear termination procedure can be used: request a graceful exit, force termination after a grace period, then collect output and execution status. When shells, child processes, or process trees are involved, the executor also needs to implement appropriate cleanup for the operating system.
But terminating a process does not undo changes already made. The command may have partially written a file or sent a deployment request to a remote service. The actual state still needs to be checked before recovery.[10]
Long-running tasks: use a queryable task system from the start
Time-consuming tasks such as builds, batch tests, and data exports can be managed by a task manager from the start, returning a task ID for subsequent status, log, and result queries.
At minimum, distinguish “accepted,” “running,” “succeeded,” “failed,” and “canceled”; when the final outcome cannot be confirmed, retain “outcome unknown” rather than misrepresenting it as failure or success. This lets the application know what it can safely do next.
Background handoff does not mean casually adding an & after a timeout. The execution environment must already support task persistence, status queries, cancellation, and resource cleanup. Otherwise, you have merely hidden the process from the user without solving task management.
In its engineering write-up on a multi-Agent research system, Anthropic notes that long-running systems need to save state and support recovery from failures rather than restart from scratch after every failure.[12]
Remote operations: query business status first
After an API for creating orders, sending emails, or submitting deployments times out, first query its status using the existing order number, task ID, or business request identifier. A query that temporarily returns no result does not necessarily prove that the original request never arrived; if the API has no duplicate-prevention contract, pause and verify when the outcome is unknown.[13]
6. For tools with side effects, prepare idempotency guarantees before the first execution
Sending emails, creating tickets, issuing refunds, and publishing posts can all be duplicated by retries, just like payments.
Retries of the same business operation must reuse the same idempotency key; only a new business operation should use a new key. AWS guidance on idempotent APIs explains that the request identifier should remain consistent across retries, and the server must recognize requests carrying the same identifier.[13]
If the model initiates another call after a timeout and the application generates a new key each time, the server may treat them as two separate operations.
For order creation, the recommended recovery process is:
- The application first creates a record for the business operation, storing its identifier and stable request parameters.
- On the first submission, use that identifier with the server’s supported idempotency mechanism.
- After a timeout, record the status as “outcome unknown” and retain the original identifier.
- Verify through business status queries; if resubmission is needed, reuse the original identifier and parameters according to the server’s contract.
- Update local state only after receiving a definitive result or completing verification.
Operation records and idempotency guarantees must be implemented by the application; they cannot rely solely on prompts or the model’s memory.
Several constraints also cannot be skipped:
- The server must actually support idempotency; adding an HTTP header with that name to an arbitrary API will not make it work automatically.
- Parameters associated with the same key usually need to remain consistent. If you change the amount or recipient, you cannot keep treating it as a retry of the original request.
- The server must handle concurrent duplicate requests, not just loosely “check the cache, then execute.”
- The key retention period, API scope, and handling of failed results must all be confirmed against the specific service contract.
Stripe provides a concrete example: it saves the status code and response body once the first request with an idempotency key begins execution. Subsequent requests with the same key can return the same result, including 500; records may be removed once they meet its retention criteria. Idempotency therefore does not mean “retries always succeed,” nor does it provide a permanent deduplication guarantee.[14]
If the underlying service does not support idempotency, adding a cache only on the Agent side cannot fully cover the window where “the remote operation succeeds, but the local process crashes before recording it.” You also need queryable business identifiers and reconciliation mechanisms; when the outcome cannot be confirmed, stopping is safer than submitting again.[13]
7. The Model Adjusts the Plan; the Execution Layer Enforces the Boundaries
Based on OpenAI and Anthropic's guides, responsibilities can be organized into the following engineering split, rather than cramming all recovery logic into the prompt.[1][2][7]
| The execution layer must enforce | The model can decide |
|---|---|
| Input and permission validation | Correct parameters based on clear errors |
| Timeouts, cancellation, and retry budgets | Choose permitted alternative tools |
| Persistence of idempotency keys and business state | Narrow queries and split tasks |
| Confirmation mechanisms for high-risk operations | Explain blockers and ask the user for clarification |
| Operation logs and result queries | Adjust subsequent plans based on verified state |
An application can tell the model “try at most twice,” but the execution layer should also actually block the third attempt. In particular, switching tool names, sessions, or sub-Agents must not grant a fresh, unlimited budget.
OpenAI's Agent guide lists exceeding failure thresholds and high-risk operations as typical triggers for human intervention; Anthropic's guide requires Agents to continuously obtain real feedback from tool results and set stopping conditions such as a maximum number of iterations.[1][2]
8. How Does Research Evaluate Tool-Calling Reliability?
For Agent reliability, the independent research benchmark τ-bench offers a more direct evaluation approach: it simulates multi-turn interactions between users and tool-using Agents, determines task completion by comparing the final database state with the target state, and examines consistency across repeated runs of the same task.[15]
In the configurations tested in 2024, the paper found substantial shortcomings in both completion rates and consistency across repeated runs. These were experimental results for specific models, prompts, and tasks at the time, not an upper bound on the capabilities of all models today.
Anthropic's “think” tool experiment studied the effect of adding intermediate thinking steps to Claude's complex tool use. In its Claude 3.7 Sonnet airline-domain configuration, the single-run pass metric increased from 0.370 to 0.570 with an optimized prompt, a result specific to those experimental conditions.[16]
This experiment shows that how a model handles tool results and business rules affects completion rates; it did not validate an idempotency protocol, nor does it prove that letting a model think longer makes repeated writes safe.
Practical evaluations should not just check whether the Agent ultimately says “done.” They should also verify that the final data is correct, that no extra emails were sent or duplicate objects created, and that execution stopped within budget.[15][7]
9. Test at Least These Failures Before Going Live
Before going live, use the following checklist for fault-injection testing:
| Test scenario | Expected result to verify |
|---|---|
| Recovery after temporary rate limiting | Bounded retries after waiting, with no request storm |
| Exhausted quota or insufficient permissions | Calls stop, with no repeated unchanged submissions |
| Missing fields or invalid field values | Errors help the model correct inputs, with no writes performed |
| A local command keeps running | Terminated or managed according to executor policy, with reclaimable resources |
| A remote write succeeds, but the response is lost | Success is confirmed through reconciliation, with no second object created |
| Two Agents submit the same business operation simultaneously | The server still meets its promised duplicate-prevention guarantees |
| The application crashes midway through execution | Business identifiers are preserved during recovery, with no blind resubmission |
| The user cancels midway | No new actions are started, and completed and pending operations are reported |
Logs should at least correlate tasks, tool calls, business operation identifiers, duration, error types, retry counts, and final confirmed status. Secrets and personal information must also be handled appropriately when logging; troubleshooting is not a reason to write complete credentials into logs.
Anthropic recommends collecting task accuracy, call duration, tool call counts, token usage, and tool errors; the MCP specification also calls for consideration of input validation, confirmation for sensitive operations, timeouts, and audit records.[7][9]
Finally, a practical failure-handling workflow can be condensed to:
First, confirm whether an operation occurred; for transient errors that are safe to repeat, retry with bounded backoff; for parameter errors, give the model specific feedback; for permission and quota issues, stop and hand off to external handling; for writes with unknown outcomes, query their status or recover under existing idempotency guarantees; when the budget is exceeded, exit explicitly.
References
- OpenAI: A practical guide to building agents: Failure thresholds, high-risk operations, and human takeover.
- Anthropic: Building effective agents: Environmental feedback, stopping conditions, and Agent design.
- OpenAI: Function calling: Application-side tool execution and strict parameter schemas.
- OpenAI: Error codes: Rate limits, quotas, and billing errors.
- AWS: Control and limit retry calls: Retry limits, backoff, jitter, and retries across multiple layers.
- Official OpenAI Python SDK: Automatic retries and timeout configuration.
- Anthropic: Writing effective tools for agents: Tool design, actionable error feedback, and evaluation metrics.
- Claude: Handle tool calls: Tool results and
is_error. - MCP tool specification, version 2025-06-18: Protocol errors, execution errors, and security requirements.
- AWS: Timeouts, retries, and backoff with jitter: A timeout does not mean side effects did not occur.
- Python: subprocess: Timeout and process cleanup behavior across execution interfaces.
- Anthropic: How we built our multi-agent research system: Failure recovery and state persistence in production systems.
- AWS: Making retries safe with idempotent APIs: Request identifiers, parameter consistency, and reconciliation.
- Stripe: Idempotent requests: Specific service guarantees for idempotency keys.
- τ-bench paper, Shunyu Yao et al., 2024; later published at ICLR 2025: Final-state verification and reliability across repeated runs.
- Anthropic: The “think” tool: Experiments with Claude 3.7 Sonnet on complex tool-use tasks.
Originally published on the SunAI forum.
For additional terminology and implementation details, see the comments on the original post.
Last updated 2026-10-09
Comments 0