1. The Problem: Multi-Step Work That Outlives One Invocation
1.1 Workflows that wait on payments, approvals, or model calls
Many business processes do not finish in a single invocation. An order may wait for a payment provider to confirm a charge, a refund may require a manager's approval, or an AI workflow may pause for human review before publishing a result. On standard Lambda capacity, a single invocation can run for at most 15 minutes, so teams often split longer workflows across queues, polling loops, state tables, or workflow orchestration services such as AWS Step Functions. These approaches work, but they can introduce additional components and state-management logic to build and maintain.
1.2 What durable functions are
AWS Lambda Durable Functions extend Lambda with durable operations that checkpoint workflow progress. You write the workflow as ordinary application code, using operations such as steps, waits, and callbacks to define work that can survive interruptions and long pauses. When a workflow is suspended, it does not need to keep compute running; when execution resumes, previously completed operations can be recovered through checkpointed results rather than being performed again. This lets developers express long-running workflows directly in Lambda code while Lambda manages their durable execution state.
2. How Durable Execution Works
2.1 Checkpoint and replay
A durable function does not keep a process alive while it waits. Instead, each durable operation saves a checkpoint that records its inputs and result. When the function needs to continue, whether after a wait, a retry, or an infrastructure failure, Lambda invokes it again and the SDK runs your code from the beginning but skips completed operations, using stored results instead of running them again. Execution then carries on with the first operation that has not completed.
This differs from resuming a thread or process, where memory and the call stack stay intact. Here, none of that survives. Local variables are rebuilt by running the handler again. Section 3.3 explains why that matters for your code.
2.2 Two clocks
Two separate timers apply. The Lambda function timeout limits each invocation, and on standard Lambda it can be at most 15 minutes. The durable execution timeout applies to the entire execution, not to individual invocations, so it covers every wait and replay. It can be set up to one year.
How you invoke the function decides which limit you actually get. Asynchronous invocations support execution durations up to one year, while synchronous invocations are limited to 15 minutes, because the caller waits for the result. Event source mappings, such as SQS or Kinesis, invoke functions synchronously, so on standard Lambda they carry the same 15-minute ceiling.
In short, Lambda does not run for a year. A workflow can span a year through many short invocations.
2.3 What you pay
Durable functions add their own charges on top of normal Lambda billing. Existing Lambda compute charges still apply, including for the extra invocations caused by replays, and on-demand functions incur no duration charges while a wait is in progress. The durable-specific charges are for operations, data written, and data retained.
In the pricing page example for US East (N. Virginia), operations cost $8.00 per million, data written costs $0.25 per GB, and retained data costs $0.15 per GB-month. Retention after completion is configurable from 1 to 90 days, with a default of 14. Check the pricing page for your Region.
3. Writing Workflows as Ordinary Code
3.1 Durable operations
You build workflows from a small set of operations on the durable context. A step runs and checkpoints a unit of work. A wait pauses execution for a set time. A callback pauses until an external system provides input through the Lambda API. Parallel and map run operations concurrently. Waits and callbacks suspend the function, so nothing runs while you wait. The SDK is available for JavaScript, TypeScript, Python, Java, and .NET.
The example below validates an order, waits for a manager's approval, and then charges the customer.
| typescript import { withDurableExecution, DurableContext } from "@aws/durable-execution-sdk-js"; export const handler = withDurableExecution( async (event: { orderId: string }, context: DurableContext) => { const order = await context.step("validate-order", async () => validateOrder(event.orderId) ); const decision = await context.waitForCallback<string>( "manager-approval", async (callbackId) => sendApprovalRequest(callbackId, order), { timeout: { hours: 24 } } ); if (decision !== "APPROVED") { return { status: "REJECTED" }; } const receipt = await context.step("charge-payment", async () => chargeCustomer(order.id) ); return { status: "COMPLETED", receipt }; } ); |
The functions validateOrder, sendApprovalRequest, and chargeCustomer are your own code.
3.2 Retries and semantics
Durable steps support retry strategies. When a retry is scheduled, Lambda checkpoints the failure and the next attempt begins in a new invocation; if the step exhausts its retry attempts, the final error is returned to the handler. If you don't configure a retry strategy, the SDK applies its default retry policy.
By default, steps use at-least-once-per-retry semantics. If an invocation is interrupted before the step result is checkpointed, the step may run again during replay. For operations with side effects, such as charging a card, make the operation idempotent or consider at-most-once-per-retry semantics, which prevents the step from being re-executed after its start checkpoint has been committed. Neither approach provides exactly-once execution across an entire workflow, so external operations may still require their own idempotency controls.
Checkpointed operation data is limited to 256 KB for step, wait, callback, and context operations. For larger payloads, keep the durable state small and store the underlying data in an external service such as Amazon S3, passing a reference through the workflow instead.
3.3 The determinism rule
A durable function's handler runs from the beginning whenever Lambda replays an execution. Completed durable operations return their checkpointed results without re-running their underlying code, but code outside those operations executes again. That means replayed code must produce the same values and follow the same control-flow path for the same inputs and completed operation results.
This is why operations that depend on wall-clock time, random values, external services, or other changing state should generally be placed inside durable operations. For example, reading the current time or generating a UUID directly in the handler can produce a different value during replay, potentially sending the workflow down a different path. Putting that work inside a step allows its result to be checkpointed and reused during replay. The same principle applies to external API calls, database reads, and other I/O: keep them inside durable operations rather than performing them directly in replayed handler code.
3.4 Deploying safely
Invoke durable functions using a qualified version or alias so that each execution is pinned to a specific function version. If you publish a new version or move an alias, in-progress executions continue using the version on which they started, while new executions use the updated version. AWS recommends numbered versions or aliases for production durable functions rather than $LATEST, because executions started with $LATEST can resume against updated code.
Workflow code also needs to remain compatible with executions already in progress. Avoid renaming durable steps or changing their behavior in ways that prevent the runtime from matching the saved execution state during replay.
4. Where It Fits
4.1 Human approvals via callbacks
Some workflows have to stop and wait for a person. An employee submits an expense, and the manager may reply in ten minutes or ten days. With a callback, the function sends a callback ID to the approver's system and pauses until the decision comes back. You do not need to build a polling loop or a state table to remember where each request stands. The same approach fits onboarding steps and other human-in-the-loop processes.
4.2 Order and payment flows with compensation
Order processing often means reserving inventory, charging a payment, and starting fulfillment. Any step can fail after earlier ones have succeeded. The Saga pattern handles this with compensating actions. If payment succeeds but fulfillment fails, the code runs a refund step. This is an application-level strategy for partial failure. It is not a distributed ACID transaction, and it is not an automatic rollback. You write each compensation yourself and should make it idempotent. Durable functions help by keeping the forward steps and the recovery logic together in one checkpointed function.
4.3 Multi-step AI workflows with human review
AI workflows often chain several model calls, wait on asynchronous jobs, and then pause for a reviewer who can approve the result or send it back. Durable execution keeps that progress safe if something is interrupted along the way. It does not improve model quality or prevent bad outputs. It only makes the workflow around the model durable. AWS lists Amplience, a content platform for retailers, as a customer using durable functions for its Workforce content automation engine, including pipelines with human review steps.
5. Durable Functions and Step Functions: Choosing, Not Replacing
5.1 Different starting points
The two services start from different places. With durable functions, the workflow is application code. You write it in a programming language, test it with your usual tools, and deploy it like any other Lambda function. With Step Functions, the workflow is a separate definition: a state machine written in Amazon States Language or modeled with CDK, often built in a visual designer. It has native integrations with more than 220 AWS services. AWS's own guidance says durable functions are optimized for application development in Lambda, while Step Functions is built for orchestration across AWS services. Neither is a newer version of the other.
5.2 Hybrid architectures
The two can coexist. AWS describes a common pattern where durable functions handle application-level logic inside Lambda, while Step Functions coordinates the broader workflow across services. Many systems will only need one of them, and that is a valid outcome.
5.3 Constraints
Durable Functions simplify long-running workflows, but they come with limits that should influence the design:
- Per-execution limits: A durable execution can perform up to 3,000 durable operations and persist up to 100 MB of durable execution data. These limits apply to the execution as a whole, not to an individual Lambda invocation.
- Invocation method matters: Synchronous invocations are limited by the Lambda invocation timeout, which is 15 minutes or less. Event source mappings, such as SQS, Kinesis, and DynamoDB Streams, also impose a total durable-execution limit of 15 minutes on the default capacity mode. For workflows that need more time, AWS documents an intermediary Lambda that receives the event and starts the durable function asynchronously; asynchronous durable executions can run for up to one year.
- Lambda Managed Instances: When using Lambda Managed Instances, asynchronous and event-source-mapping invocations can support up to 90 minutes per invocation/execution, with Amazon MQ and Amazon DocumentDB event source mappings remaining limited to 15 minutes. Synchronous invocations remain limited to 15 minutes.
- Configuration and invocation: Durable execution must be enabled when the function is created; it cannot be added to an existing function. Durable functions also require a qualified function ARN for invocation, such as a version or alias.
For observability, Lambda provides durable-execution metrics and standard logs through Amazon CloudWatch, supports AWS X-Ray tracing, and exposes durable execution history in the Lambda console. These tools help distinguish individual Lambda invocations and replays from the progress and outcome of the overall durable execution.
6. Takeaway
6.1 A decision rule
Consider durable functions when the workflow logic is closely tied to Lambda code, needs to survive long waits and failures, and is easier to maintain as ordinary code. Step Functions may fit better when the workflow spans many AWS services, benefits from a visual definition, or should exist as a separate artifact that others can read.
6.2 Where to start
Pick one Lambda-centric workflow with a concrete problem, such as an approval wait, a callback, or a multi-step job that needs retries. Build it, watch the limits and costs, and let that result decide where else to use it.

