AI agent retry logic decides whether a job recovers cleanly or turns one small outage into duplicate work. This test examines the operating rules around a mid-workflow failure: a run completes some steps, hits a timeout or unavailable service, and must choose between retrying one step, resuming from a checkpoint, or stopping for a person.
This is not a model benchmark. The original post names no Wiro model, and its two attached images contain no prompt, model identifier, seed, run record, latency, or price. No model was run for this update. That matters: it would be misleading to attach a runtime or cost to diagrams whose generation details were never recorded. The useful test here is the workflow policy itself.
- What the retry test checks
- What the two outputs show
- Six rules for AI agent retry logic
- Parameters and observability
- When to use each recovery path
What this AI agent retry logic test checks
A workflow is rarely all-or-nothing. It may fetch a customer record, classify the request, create a CRM note, schedule a follow-up, and then send a message. A network timeout after the CRM call does not prove that the CRM call failed. It may have succeeded while the acknowledgement was lost. Replaying the entire run can create a second note, a second task, or a second message.
The test therefore asks six practical questions. Can the step safely run twice? Is there durable evidence that it already finished? Does the downstream system accept an idempotency key? How many attempts are allowed? How long should the job wait between attempts? At what point does the job hand a clear case to an operator instead of continuing on its own?
Microsoft’s AI agent orchestration patterns documentation makes the same distinction between agent reasoning and orchestration. The model can propose an action; the workflow still needs explicit control over state, execution, and recovery. The underlying reference is also available in the Microsoft Architecture Center repository.
What the existing outputs actually show

The first image is useful as an operations view, not as evidence of a particular model’s ability. Its visible design shows a sequence of typed workflow events and a detail panel for the selected item. That is the right mental model for a retry system: each attempt needs a status, a timestamp, an owner, and enough context to explain what happened. The image is a 1536 by 864 JPEG hosted in this blog’s media library. The attachment has no embedded model metadata, so no generation parameters, Wiro runtime, or cost can be reported honestly.

The second image makes a different point. A retry is not a single global switch. The board separates steps and exposes their status independently. That supports a checkpointed resume: if enrichment and record creation completed, but the final notification timed out, the next attempt should inspect the notification state rather than replay enrichment and creation. Like the first image, it is a 1536 by 864 JPEG with no stored generation settings. It is retained as inline blog media, but it should not be presented as a timed or priced Wiro model output.
Six rules for jobs that fail mid-workflow
1. Retry reads more freely than writes
Fetching a record, checking availability, or reading a status endpoint is usually safe to retry. Creating a payment, sending an email, or posting a CRM note is different. Before automatically retrying a write, query for the expected result or send the same idempotency key. If neither option exists, stop and surface the ambiguity.
2. Checkpoint after every durable success
Store a run ID, step ID, input hash, completion timestamp, and returned external ID after a successful step. A checkpoint turns recovery into a targeted resume. Without it, the system only knows that the full job is incomplete, which pushes it toward unsafe replay.
3. Bound the attempt budget
Use a fixed maximum for a transient class, such as three attempts, rather than retrying until the queue drains. The exact number should match the service and the job deadline. A five-minute appointment lookup may deserve a short budget. A nightly enrichment job may tolerate a longer queue. The key rule is that the budget is explicit and visible.
4. Back off, then add jitter
Immediate retries often strike the same overloaded dependency at the same time as every other worker. Exponential backoff spaces attempts out. Jitter prevents a group of jobs from waking up together. Keep the delay parameters in the workflow definition, not inside a natural-language prompt, so operators can review and change them safely.
5. Use dedupe keys at every external boundary
A run ID is not enough when the job can restart. Derive a stable key from the business action, such as account ID plus campaign ID plus scheduled date. Send it where an API supports idempotency, and store it locally where it does not. The next attempt can then ask whether that exact action already happened.
6. Escalate an actionable packet
A failed retry should not produce only “job failed.” Include the last completed checkpoint, the failing step, attempt count, error class, external IDs, next retry time, and the safest manual action. That gives an operator a decision, not a scavenger hunt.
Parameters, runtime, and cost: what can be stated
For this workflow test, the parameters are policy parameters rather than model inference settings: step idempotency, checkpoint storage, dedupe key format, maximum attempts, backoff curve, jitter range, timeout, and escalation destination. Those settings determine the output of the recovery decision. The original images do not document their own creation settings, so they do not supply a prompt, seed, model, resolution setting beyond the stored 1536 by 864 files, execution time, or Wiro price.
No cost-per-output figure is included because there is no named Wiro model or recorded Wiro run to support one. That is not a missing optimization; it is a boundary on the evidence. If a future revision adds model-generated samples, it should name the model, link its Wiro model page and documentation, preserve the run parameters, and report only the runtime and cost returned by that run.
When to pick each recovery path
Pick an automatic retry for a read or other idempotent operation that failed with a known transient error. Pick a checkpointed resume when earlier steps have durable completion records and only a later step failed. Pick a reconciliation check before retrying a write when the response may have been lost after the downstream system accepted it. Pick human review when the job affects money, external communication, legal records, or customer commitments and the final state cannot be proved.
For the surrounding operating model, see AI Agent Audit Trails, Multi-Agent Workflows for Business, and AI Agent Analytics. They cover the evidence, routing, and metrics that make a retry policy inspectable after the fact.
AI agent retry logic earns trust by refusing to guess. A safe system records what completed, retries only what it can prove is safe, and stops with useful context when it cannot. That is less flashy than an endless self-healing loop. It is also how an agent stays dependable in live work.