Durable AI agents: what should happen after a crash
How Pi Durable and Rivet restore agent progress, why external actions need their own retry contracts, and what to test before allowing unattended writes.
By Clairevue · · 5 min read

Suppose an AI assistant handles an approved customer refund. The payment service accepts the request, but the connection drops before the assistant receives confirmation. Then the process crashes.
When it restarts, repeating the refund could move money twice. Marking it complete without checking could leave your records wrong. The agent needs a way to identify that particular approved action and find out what happened to it.
Durable AI agents save progress so a failed process can continue unfinished work. For business workflows, that recovery design must extend to the systems the agent changes. A saved conversation alone can’t tell you whether another service accepted a request whose reply never arrived.
What Pi Durable restores
Pi Durable is an experimental agent runtime that commits conversations and task state to storage. Its tasks save checkpoints, which let a new process continue from saved progress rather than ask a model to reconstruct the job from memory.
The package’s in-memory backend saves nothing across restarts. Its Node SQLite adapter documents survival of process crashes while warning that the newest commits may be lost on power or host failure. “Durable” needs a defined failure case and a suitable storage configuration.
Rivet’s Pi integration runs that runtime inside an Actor, a unit of running application code with its own persisted state. It stores the agent’s progress in the Actor’s SQLite database. Clients can resend a prompt with the same requestId to find the existing submission instead of starting another run.
That input identifier protects admission of the job. A refund inside the job still needs its own operation identity at the payment service.
According to the Rivet recovery table, an interrupted model call starts again; a tool marked replay: "safe" runs again. Other interrupted tools return a result telling the model that execution stopped partway, instead of automatically rerunning. A task phase can run again too.
Declaring a tool replay-safe is a developer promise about its behavior. It doesn’t make a refund API safe to repeat. And declining an automatic rerun doesn’t undo a refund the provider already accepted; the application still has to resolve that uncertainty before allowing another write.
The gap between doing and recording
An agent can save “refund complete” after making the request, but a crash between the external action and that save leaves a missing receipt. Saving “complete” before the request creates the opposite risk: the process can die before sending anything, and recovery skips unfinished work.
Idempotency means retries of the same operation don't create additional effects. AWS’s guidance on idempotent APIs describes this problem through a resource-creation request whose response disappears. AWS uses a caller-provided identifier to recognize repeated requests for the same intended operation. The receiving service must handle that identifier and the resource change together under its contract.
In Rivet’s deployment example, a task creates a preview at an address derived from the task’s ID, so a rerun can address the same preview. In its next phase, it posts a GitHub comment and then saves completion.
If GitHub accepts the comment but the process stops before that checkpoint, recovery can re-enter the comment phase. The example’s comment request doesn’t show a duplicate-prevention mechanism. This is a failure window visible in the sample code, not a duplicate we reproduced. Moving work into a durable task preserves its phase; the external action still needs a retry contract or a way to reconcile its result.
Give the approved action a stable identity
For the fictional refund, the application should create and save an operation key before sending the request. Reuse that key for retries of the same approved refund, with the same parameters. A new key on every attempt tells the receiving service that each attempt is a new operation.
Store the approved amount and target payment with that key. When the provider returns a refund identifier, retain it with the result. Recovery can then use the saved identity to retrieve or reconcile the provider’s record instead of asking the model whether it probably succeeded.
Stripe’s API reference gives a specific contract to examine. For API version 1, once endpoint execution begins, Stripe saves the first request’s status and body for an idempotency key; repeats return that result, including a saved server error. It checks that parameters match. Keys can be pruned after they’re at least 24 hours old, and reuse after pruning creates a new request.
A key in your database doesn’t grant indefinite duplicate protection at the provider. Check the exact endpoint and API version, then decide how to handle delayed recovery or a result that remains ambiguous. If the service offers neither a suitable retry contract nor a reliable lookup, stop for operator review before repeating a financial write.
Ask for the failure test before unattended access
Use a fake provider or the service’s test mode to exercise the refund workflow. Stop the agent before it sends a request, then repeat the test after the provider accepts it but before the agent records the reply. Also restart after the receipt is saved and resend the original input using the same requestId.
Check external records, not just the agent’s final message. Verify the intended refund and its local receipt. When the system can’t determine the outcome, it should leave the job visibly unresolved. Keep the original human approval attached to the exact action through recovery; Pi’s Rivet integration doesn’t supply a built-in approval flow.
Repeat the exercise after a code deployment. Rivet’s version documentation covers worker upgrades and persisted-state migrations; your application must still understand the older workflow’s saved state and approval. A successful restart under unchanged code doesn’t settle that question.
Set a retry budget and identify who handles exhausted or ambiguous runs. Repeated model calls can consume tokens, while unresolved jobs can still require staff work.
Before granting unattended write access, ask the developer to demonstrate the accepted-refund, lost-reply case. The system should recover the approved action’s outcome or leave it visibly unresolved without issuing a second refund.