← Back to Blog
Backend & Architecture

Durable AI Agents: A Runtime Selection Checklist for Production Workflows in 2026

Durable AI agents are not just better prompts attached to stronger models. They are runtime decisions about state, retries, approvals, tool permissions, recovery, and observability across workflows that may outlive a single request.

Written by Hamza Diaz
October 4, 202610 min read22 views

Why durable AI agents are a runtime decision, not just a model decision

Durable AI agents are not stronger prompts with a few tools attached. They are production workflows that can survive delay, partial failure, duplicate events, human review, provider changes, and worker restarts. That distinction matters because most agent demos hide the hardest part. The model answers, a tool runs, and the page looks convincing. Production asks a less flattering question: what happens when the approval comes back tomorrow, the worker dies halfway through, or the same webhook lands twice?

My opinion: many teams choose agent infrastructure too late. They pick a model, build a chat loop, wire in a few internal tools, then discover that the real product is a workflow engine wearing an AI interface. By then, state is scattered across prompts, logs, temporary memory, and side effects. The runtime has already been chosen by accident.

For 2026 planning, the safer question is not which model should run the agent. It is which runtime should own state, retries, permissions, approvals, and recovery. Temporal's Durable AI documentation starts from workflow orchestration concepts. Cloudflare Agents and Durable Objects are relevant for stateful agents and coordination in Cloudflare-hosted applications. Convex agents fit teams already building reactive app backends. Val Town fits lightweight programmable automation. A custom stack can work when control matters more than speed, but it also means the team owns the awkward parts.

The Durable Agent Fit Matrix

A practical selection process needs fewer vendor adjectives and more failure testing. The Durable Agent Fit Matrix compares runtimes across five axes.

AxisWhat to inspectWhy it matters
Workflow durationSeconds, minutes, hours, days, or longerLong tasks need persistence, resumption, and clear timeout behavior.
State ownershipRuntime state, database state, or app stateThe source of truth should be obvious during replay and incident review.
Tool riskRead-only, write actions, approvals, external messagesHigher-risk tools need tighter permissions and audit trails.
Developer workflowApp team, platform team, automation teamThe runtime should match the team that will debug and maintain it.
ObservabilityHistory, logs, traces, replay, evaluation recordsIf the team cannot explain what happened, it cannot operate the agent.

Temporal is a strong candidate when the workflow is long running, stateful, and full of business process edges. Its AI documentation is worth reading if the team already thinks in terms of workflows, activities, signals, and retries: https://docs.temporal.io/ai. The trade-off is that teams need to learn the workflow model and treat agent behavior as part of a larger distributed system.

Cloudflare Agents, paired with Durable Objects, suit teams that want stateful agents in Cloudflare-hosted web applications. The relevant docs are https://developers.cloudflare.com/agents/ and https://developers.cloudflare.com/durable-objects/. This path can be attractive for user-facing agent sessions, coordination, and low-latency state inside the Cloudflare platform. The caveat is platform fit. If the rest of the system lives far from Cloudflare, integration choices matter.

Convex agents fit product teams that already use Convex or want agent state close to reactive app data: https://docs.convex.dev/agents/overview. The advantage is less glue between app state and agent state. The risk is the same as with any app-native runtime: it may be perfect for product workflows and less natural for cross-system orchestration.

Val Town is best treated as a fast automation layer, not a universal agent platform: https://docs.val.town/. It is useful for small internal tools, scripts, scheduled tasks, prototypes, and glue code. It is less appropriate when the workflow needs strict auditability, multi-step approvals, complex permissions, or deep incident response.

A custom runtime, built from queues, databases, workers, schedulers, and model gateways, can be the right answer for regulated environments or unusual architecture constraints. It gives control over storage, networking, model routing, and policy. It also means the team must design idempotency, retries, observability, replay, evaluations, approval flows, and cost controls directly. That is not a weekend project.

A practical evaluation playbook

Do not evaluate durable agent runtimes with a happy-path demo. Pick one real workflow and make it misbehave on purpose. For example, use a support triage workflow that reads a ticket, checks account context, drafts a reply, waits for approval, updates a CRM, and sends a message. Then test the parts that usually break.

Kill the worker after the model call but before the write action. Deliver the same webhook twice. Change a tool response schema. Delay human approval for 24 hours. Return a partial timeout from the CRM. Rotate a model provider key. Ask the agent to resume from the middle. If the runtime cannot make those cases routine, it is not ready for that workflow.

A small proof of durability should answer these questions:

  1. Where is the workflow state stored, and who owns it?
  2. Can the workflow resume after a process restart without redoing unsafe actions?
  3. Are tool calls idempotent, or protected by idempotency keys?
  4. Can a human approve, reject, or edit a proposed action without breaking the run?
  5. Can an operator inspect the history and explain why the agent acted?
  6. Are prompts, tool schemas, model versions, and outputs recorded well enough for evaluation?
  7. What happens when the model gives a bad answer but the runtime behaves correctly?

That last question is easy to skip. It is also where many agent programs become expensive. A durable runtime can make failure recoverable, but it cannot make a weak task design good. Teams still need evaluation sets, policy checks, and clear escalation paths.

Security and governance change the runtime decision

Agent security is not only a prompt issue. It is a runtime architecture issue. Simon Willison's lethal trifecta framing is a useful lens because it connects three conditions: access to private data, exposure to untrusted content, and a way to exfiltrate information: https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/. When all three appear in the same agent flow, prompt injection risk becomes more concrete.

Runtime selection should therefore separate read paths, write paths, and outbound communication paths. A research agent that reads internal documents should not automatically have permission to email arbitrary recipients. A procurement agent that can draft a purchase order should not be able to approve it without policy checks. A developer agent that reads source code should have narrow rules for posting data outside the workspace.

Human approval is still useful, but only if it is specific. Approve this run is vague. Approve sending this exact email to these recipients with these attachments is better. The runtime should preserve the proposed action, the reviewer, the decision, and the final tool call. Otherwise the approval is just weak governance with a timestamp.

What teams get wrong

The first mistake is treating retries as intelligence. Retrying a bad tool call may fix a transient failure. Retrying a bad plan can make the damage repeat faster. Durable runtimes need retry policies, but they also need stop conditions, human escalation, and records that show what was retried and why.

The second mistake is confusing memory with source of truth. Conversation memory is useful context. It should not be the only place where approvals, customer state, financial actions, or compliance decisions live. Durable workflows need a real source of truth, and the runtime should make that boundary clear.

The third mistake is skipping duplicate protection. Webhooks repeat. Users double-click. Schedulers overlap. Providers time out after completing a request. Any workflow that writes to external systems needs idempotency keys, external IDs, or another duplicate-control pattern.

The fourth mistake is building a platform before proving one workflow. Platform work feels productive because it creates abstractions. The better first move is a narrow workflow with painful edges. If the runtime handles that well, abstractions will be based on evidence rather than taste.

The fifth mistake is ignoring cost and latency until launch. Durable agents can call multiple models, tools, databases, and policy checks. Some workflows need that. Others need a simpler path, such as retrieval plus a drafted recommendation. Measure the whole run, not just model latency.

Recommended paths

For product teams adding agents to an existing app, start with the runtime closest to application state. Convex may fit if the product already uses Convex. Cloudflare can fit user-facing stateful sessions inside Cloudflare-hosted applications. The deciding factor is usually operational ownership: the people shipping the feature should also be able to inspect and debug the agent.

For platform teams orchestrating long business processes, Temporal deserves serious evaluation. It makes duration, retries, signals, and workflow history first-class concerns. That is useful when agents become part of onboarding, finance, operations, procurement, or support processes.

For developer teams building internal automations, Val Town can be a good starting point when the task is small and reversible. Keep the scope honest. If the automation begins touching sensitive data, approvals, or irreversible writes, move the workflow into a runtime with stronger controls.

For teams with strict compliance, unusual networking, or custom model routing needs, a custom stack may be justified. The team should budget for more than workers and queues. It will need policy enforcement, audit trails, evaluation storage, replay tooling, incident procedures, and ongoing maintenance.

The planning checklist

Before production work starts, write down the answers in plain language:

DecisionAnswer needed before build
StateWhat system is the source of truth for each step?
RecoveryWhich failures can resume, retry, stop, or escalate?
ToolsWhich actions are read-only, write actions, or externally visible?
ApprovalWhat exactly does a human approve, and where is that recorded?
EvaluationWhat examples prove the agent is safe and useful enough?
ObservabilityHow will an operator reconstruct a run after an incident?
CostWhat budget applies per run, per user, and per failed attempt?

The best runtime is not the one with the most agent branding. It is the one that makes failure understandable, limits unsafe action, and gives operators a clean path back to a known state. Choose it with a real workflow, not a slide deck.

Key Takeaways

  • 1Durable AI agents require runtime guarantees for state, retries, approvals, permissions, and recovery, not only stronger model responses.
  • 2Temporal, Cloudflare Agents SDK with Durable Objects, Convex agents, Val Town, and custom orchestration each fit different workflow shapes.
  • 3The Optijara Durable Agent Fit Matrix compares workflow duration, state ownership, tool risk, developer workflow, and observability before runtime selection.
  • 4Teams should prototype the hard path, including restarts, duplicate events, timeouts, delayed approvals, stale memory, and permission denial.
  • 5Prompt injection and external tool calls make runtime architecture a security decision, especially when private data, untrusted content, and external communication combine.
  • 6Agent memory should support context, not replace authoritative application data, workflow history, permissions, or audit logs.
  • 7The right runtime is the one the owning team can operate, inspect, secure, and repair when production workflows fail.

Conclusion

Durable AI agents are production workflows before they are user experiences. Pick the runtime that makes failure manageable: persistent state, safe retries, specific approvals, scoped tools, inspectable histories, and clear recovery paths. For most teams, the next step is not a broad platform bet. It is a focused evaluation of one real workflow across Temporal, Cloudflare, Convex, Val Town, or a carefully scoped custom stack.

Frequently Asked Questions

What are durable AI agents?

Durable AI agents are agent workflows whose state, tool calls, retries, approvals, and recovery behavior can survive beyond one model response or one server process.

How do I choose an AI agent runtime?

Compare workflow duration, state ownership, tool risk, developer experience, observability, hosting constraints, and failure recovery. Then test the top candidates against one representative workflow with real failure scenarios.

When is Temporal a good fit for AI agents?

Temporal is strongest when the agent workflow is long running, retry heavy, approval based, or part of broader business process orchestration.

When should teams consider Cloudflare Agents SDK or Durable Objects?

They are worth evaluating for web adjacent, stateful agents where coordination, latency, and integration with Cloudflare's platform matter.

How does prompt injection affect runtime selection?

If an agent can read private data, process untrusted content, and act externally, the runtime must support strong permission boundaries, auditing, scoped tools, and approval controls.

Sources

Share this article

Hamza Diaz

Written by

Hamza Diaz

Hamza Diaz is the founder of Optijara, where he builds practical AI agents, automation systems, and Copilot workflows for service businesses. He writes about AI operations, agent strategy, and real-world implementation for teams that want usable systems instead of hype.