Photo by Adhitya Sibikumar on Unsplash
A Lambda on-failure destination receives a record of every asynchronous event that fails all retries or expires in the queue. Without one, or a dead-letter queue, Lambda discards the event. To audit it, list functions, call ListFunctionEventInvokeConfigs for each and check DestinationConfig.OnFailure and DeadLetterConfig. Add one with PutFunctionEventInvokeConfig.
When S3, SNS or EventBridge invokes a function, the caller gets a success response as soon as Lambda queues the event. If the function then throws, Lambda retries it twice by default and, after that, drops the event. Nobody is told. The thumbnail never appears, the invoice email never goes out, and the only trace is an Errors datapoint in CloudWatch.
This example is for developers and platform engineers who own many functions. The script reports every function’s asynchronous error handling, including which services invoke it asynchronously, and with --apply adds a Lambda on-failure destination on a standard SQS queue to the unprotected ones. If you’re chasing failures that already happened, start with the guide to investigate Lambda errors with CloudWatch.
What happens to a failed asynchronous Lambda event?
Lambda puts asynchronous events in its own queue. If the function returns an error, Lambda retries twice by default, waiting one minute before the second attempt and two minutes before the third. Throttled events go back to the queue and are retried for up to 6 hours. A timeout counts as an error, so a hung function holds each attempt open for its full timeout; the script to find Lambda functions with a timeout too high for their runtime shows which timeouts are far above real run times. When an event fails every attempt or gets older than the maximum age, Lambda discards it, unless you’ve configured somewhere to send it:
| Setting | Where it’s configured | What it keeps |
|---|---|---|
| On-failure destination | Event invoke config, per function, version or alias | An invocation record: the event, the error, the condition (such as RetriesExhausted) and the attempt count. Targets: standard SQS queue, standard SNS topic, S3 bucket, Lambda function or EventBridge bus |
| Dead-letter queue | Function configuration, function level only | Only the event, plus the error code and first 1 KB of the message as attributes. Targets: standard SQS queue or SNS topic |
MaximumRetryAttempts |
Event invoke config | 0 to 2, default 2 |
MaximumEventAgeInSeconds |
Event invoke config | 60 to 21,600 (6 hours), default 21,600 |
The Lambda guide to capturing records of asynchronous invocations notes that destinations support more targets than dead-letter queues and include details of the function’s response. FIFO queues and FIFO topics don’t work for either: a FIFO dead-letter queue isn’t supported, and a FIFO destination fails delivery.
Which Lambda triggers are asynchronous?
This setting only matters for asynchronous invocations. Per the Lambda Developer Guide, S3, SNS, EventBridge rules and Scheduler, CloudWatch Logs subscriptions, SES, AWS IoT, AWS Config, CodeCommit and CloudFormation invoke asynchronously, as does any Invoke call with InvocationType: "Event", shown in the guide to invoke a Lambda function with AWS SDK v3 in TypeScript.
API Gateway, Application Load Balancers and Cognito invoke synchronously: the caller sees the error. SQS, Kinesis and DynamoDB Streams use event source mappings, which have their own failure handling. For SQS that’s a redrive policy on the queue itself, which the script to find SQS queues without a dead-letter queue audits. The batch setting of an SQS mapping matters as much: the script to find Lambda SQS event source mappings without ReportBatchItemFailures shows which functions retry a whole batch for one failed message. For DynamoDB Streams, a failing batch blocks its shard until the records expire unless you limit retries; the guide to process DynamoDB Streams in Lambda with TypeScript shows mapping settings that avoid that.
To tell which functions are at risk, the script reads each function’s and alias’s resource-based policy and looks for asynchronous service principals such as s3.amazonaws.com. EventBridge Scheduler and direct SDK callers don’t appear there, so treat “no async trigger found” as “check”, not “safe”.
What does the script do?
- Lists functions
paginateListFunctionsreturns each function with itsDeadLetterConfig. - Finds asynchronous triggers
GetPolicyfor the function and each alias fromListAliases; a missing policy returnsResourceNotFoundException, which the script treats as none. - Reads event invoke configs
ListFunctionEventInvokeConfigsreturns the config of the function and of every version or alias that has one, so one call per function is enough. The function-level config is reported as$LATEST. - Classifiesok with an on-failure destination, DLQ only, or UNPROTECTED when an async trigger exists and failed events have nowhere to go.
- Adds a destination with
--applyFor unprotected functions,PutFunctionEventInvokeConfigwhen no function-level config exists, orUpdateFunctionEventInvokeConfigwhen one does.Putreplaces the whole config and removes anything you leave out, so the script only uses it when there’s nothing to lose.
Prerequisites
- Node.js 18 or later,
tsxand@aws-sdk/client-lambda. - For
--apply: a standard SQS queue for failure records, with a retention period long enough for someone to look (the maximum is 14 days). The guide to send and receive SQS messages with AWS SDK v3 shows how to read them back. - Permission for each function’s execution role to send to that queue (next section).
Which IAM permissions does it need?
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ListFunctions",
"Effect": "Allow",
"Action": "lambda:ListFunctions",
"Resource": "*"
},
{
"Sid": "ReadAsyncConfig",
"Effect": "Allow",
"Action": [
"lambda:ListFunctionEventInvokeConfigs",
"lambda:ListAliases",
"lambda:GetPolicy"
],
"Resource": [
"arn:aws:lambda:us-east-1:123456789012:function:*",
"arn:aws:lambda:us-east-1:123456789012:function:*:*"
]
},
{
"Sid": "ApplyOnly",
"Effect": "Allow",
"Action": [
"lambda:PutFunctionEventInvokeConfig",
"lambda:UpdateFunctionEventInvokeConfig"
],
"Resource": "arn:aws:lambda:us-east-1:123456789012:function:*"
}
]
}
Leave ApplyOnly out of the report profile. Separately, each function’s execution role needs sqs:SendMessage on the destination queue, or Lambda can’t deliver the record and publishes a DestinationDeliveryFailures metric instead. The IAM policy generator for TypeScript code derives the caller policy from the script if you change it.
The script to add a Lambda on-failure destination
// find-lambda-functions-without-failure-destination.ts
// Reports every Lambda function's asynchronous error handling: on-failure destination, dead-letter queue,
// retry attempts and maximum event age, per function, version and alias that has a config, plus which
// services allowed by the function's or its aliases' resource-based policies invoke it asynchronously.
// Read-only by default. With --apply --queue-arn ARN it sets that standard SQS queue as the on-failure
// destination of each unprotected function that has an asynchronous trigger.
// Usage:
// npx tsx find-lambda-functions-without-failure-destination.ts [--csv lambda-async.csv]
// npx tsx find-lambda-functions-without-failure-destination.ts --apply --queue-arn arn:aws:sqs:us-east-1:123456789012:lambda-failures
import { writeFileSync } from "node:fs";
import {
GetPolicyCommand,
LambdaClient,
PutFunctionEventInvokeConfigCommand,
ResourceNotFoundException,
UpdateFunctionEventInvokeConfigCommand,
paginateListAliases,
paginateListFunctionEventInvokeConfigs,
paginateListFunctions,
type FunctionEventInvokeConfig,
} from "@aws-sdk/client-lambda";
const args = process.argv.slice(2);
const flag = (name: string): string | undefined => {
const i = args.indexOf(name);
return i >= 0 ? args[i + 1] : undefined;
};
const apply = args.includes("--apply");
const queueArn = flag("--queue-arn");
const csvPath = flag("--csv");
if (apply && !queueArn?.startsWith("arn:aws:sqs:")) {
console.error("--apply needs --queue-arn with the ARN of a standard SQS queue");
process.exit(1);
}
// Service principals of AWS services that invoke Lambda asynchronously (Lambda Developer Guide).
const ASYNC_PRINCIPALS = ["s3", "sns", "events", "logs", "ses", "iot", "config", "codecommit"];
const lambda = new LambdaClient({}); // Region from AWS_REGION or the profile
interface Row {
Function: string;
Qualifier: string;
AsyncTriggers: string;
OnFailure: string;
DeadLetterQueue: string;
Retries: number | string;
MaxAgeSec: number | string;
Verdict: string;
Action: string;
}
/** Async service principals in the resource-based policies of the function and its aliases. */
async function asyncTriggers(name: string): Promise<string[]> {
const qualifiers: (string | undefined)[] = [undefined];
for await (const page of paginateListAliases({ client: lambda }, { FunctionName: name })) {
qualifiers.push(...(page.Aliases ?? []).map((a) => a.Name));
}
const found = new Set<string>();
for (const Qualifier of qualifiers) {
try {
const { Policy = "{}" } = await lambda.send(new GetPolicyCommand({ FunctionName: name, Qualifier }));
const doc = JSON.parse(Policy) as { Statement?: { Principal?: { Service?: string | string[] } }[] };
for (const st of doc.Statement ?? []) {
for (const s of ([] as string[]).concat(st.Principal?.Service ?? [])) {
const prefix = s.split(".")[0];
if (ASYNC_PRINCIPALS.includes(prefix)) found.add(prefix);
}
}
} catch (err) {
if (!(err instanceof ResourceNotFoundException)) throw err; // no policy on this function or alias
}
}
return [...found];
}
async function invokeConfigs(name: string): Promise<FunctionEventInvokeConfig[]> {
const configs: FunctionEventInvokeConfig[] = [];
for await (const page of paginateListFunctionEventInvokeConfigs({ client: lambda }, { FunctionName: name })) {
configs.push(...(page.FunctionEventInvokeConfigs ?? []));
}
return configs;
}
async function addFailureDestination(name: string, existing: FunctionEventInvokeConfig | undefined): Promise<string> {
const OnFailure = { Destination: queueArn };
try {
if (existing) {
// Update keeps the retry and age settings; send OnSuccess back so it isn't dropped.
await lambda.send(new UpdateFunctionEventInvokeConfigCommand({
FunctionName: name,
DestinationConfig: { OnSuccess: existing.DestinationConfig?.OnSuccess, OnFailure },
}));
} else {
// No config yet: Put creates one; retries (2) and maximum age (6 hours) stay at their defaults.
await lambda.send(new PutFunctionEventInvokeConfigCommand({ FunctionName: name, DestinationConfig: { OnFailure } }));
}
return "destination added";
} catch (err) {
return `error: ${err instanceof Error ? `${err.name}: ${err.message}` : String(err)}`;
}
}
function toCsv(rows: Row[]): string {
const cols = Object.keys(rows[0] ?? {}) as (keyof Row)[];
const cell = (v: string | number) => `"${String(v).replace(/"/g, '""')}"`;
return [cols.join(","), ...rows.map((r) => cols.map((c) => cell(r[c])).join(","))].join("\n") + "\n";
}
async function main(): Promise<void> {
const rows: Row[] = [];
for await (const page of paginateListFunctions({ client: lambda }, {})) {
for (const fn of page.Functions ?? []) {
const name = fn.FunctionName ?? "";
const dlq = fn.DeadLetterConfig?.TargetArn ?? "";
const triggers = await asyncTriggers(name);
const configs = await invokeConfigs(name);
// The function-level config is reported with the $LATEST qualifier; versions and aliases can have their own.
const qualifierOf = (c: FunctionEventInvokeConfig) => /:function:[^:]+:(.+)$/.exec(c.FunctionArn ?? "")?.[1] ?? "$LATEST";
const unqualified = configs.find((c) => qualifierOf(c) === "$LATEST");
const entries: [string, FunctionEventInvokeConfig | undefined][] = configs.map((c) => [qualifierOf(c), c]);
if (!unqualified) entries.unshift(["$LATEST", undefined]);
for (const [qualifier, cfg] of entries) {
const onFailure = cfg?.DestinationConfig?.OnFailure?.Destination ?? "";
const protectedBy = onFailure ? "destination" : dlq ? "DLQ only" : "";
const verdict = protectedBy === "destination" ? "ok"
: protectedBy === "DLQ only" ? "DLQ only: no response details"
: triggers.length ? "UNPROTECTED: failed events are dropped"
: "no failure handling (no async trigger found)";
const row: Row = {
Function: name,
Qualifier: qualifier,
AsyncTriggers: triggers.join(" ") || "-",
OnFailure: onFailure.split(":").slice(-1)[0] || "-",
DeadLetterQueue: dlq.split(":").slice(-1)[0] || "-",
Retries: cfg?.MaximumRetryAttempts ?? "2 (default)",
MaxAgeSec: cfg?.MaximumEventAgeInSeconds ?? "21600 (default)",
Verdict: verdict,
Action: "",
};
// Fix only the function-level config of unprotected functions with an async trigger.
if (verdict.startsWith("UNPROTECTED") && qualifier === "$LATEST") {
row.Action = apply ? await addFailureDestination(name, unqualified) : "would add destination";
}
rows.push(row);
}
}
}
if (rows.length === 0) {
console.log("No Lambda functions in this Region.");
return;
}
console.table(rows);
const unprotected = rows.filter((r) => r.Verdict.startsWith("UNPROTECTED"));
console.log(`${unprotected.length} function configs with asynchronous triggers drop failed events without a record.`);
if (csvPath) {
writeFileSync(csvPath, toCsv(rows));
console.log(`Wrote ${rows.length} rows to ${csvPath}`);
}
if (!apply) console.log("Dry run: nothing changed. Add --apply --queue-arn <standard SQS queue ARN> to add on-failure destinations.");
}
main().catch((err) => {
console.error(err);
process.exit(1);
});
How do you run it?
npm install @aws-sdk/client-lambda
npm install --save-dev tsx typescript @types/node
# Report for one Region
AWS_PROFILE=readonly AWS_REGION=us-east-1 npx tsx find-lambda-functions-without-failure-destination.ts --csv lambda-async.csv
# Add the queue as on-failure destination to unprotected functions
AWS_PROFILE=lambda-admin AWS_REGION=us-east-1 npx tsx find-lambda-functions-without-failure-destination.ts \
--apply --queue-arn arn:aws:sqs:us-east-1:123456789012:lambda-failures
The script covers one Region per run; loop over your Regions in the shell if you use several.
Sample output
┌─────────┬───────────────────────┬───────────┬───────────────┬─────────────────────────┬─────────────────┬───────────────┬───────────────────┬────────────────────────────────────────────────┬─────────────────────────┐
│ (index) │ Function │ Qualifier │ AsyncTriggers │ OnFailure │ DeadLetterQueue │ Retries │ MaxAgeSec │ Verdict │ Action │
├─────────┼───────────────────────┼───────────┼───────────────┼─────────────────────────┼─────────────────┼───────────────┼───────────────────┼────────────────────────────────────────────────┼─────────────────────────┤
│ 0 │ 'thumbnail-generator' │ '$LATEST' │ 's3' │ '-' │ '-' │ '2 (default)' │ '21600 (default)' │ 'UNPROTECTED: failed events are dropped' │ 'would add destination' │
│ 1 │ 'order-events' │ '$LATEST' │ 'events' │ 'order-events-failures' │ '-' │ 1 │ 3600 │ 'ok' │ '' │
│ 2 │ 'invoice-mailer' │ '$LATEST' │ 'sns' │ '-' │ 'invoice-dlq' │ '2 (default)' │ '21600 (default)' │ 'DLQ only: no response details' │ '' │
│ 3 │ 'report-builder' │ '$LATEST' │ 'events' │ '-' │ '-' │ '2 (default)' │ '21600 (default)' │ 'UNPROTECTED: failed events are dropped' │ 'would add destination' │
│ 4 │ 'report-builder' │ 'live' │ 'events' │ 'report-failures' │ '-' │ 2 │ 7200 │ 'ok' │ '' │
│ 5 │ 'api-handler' │ '$LATEST' │ '-' │ '-' │ '-' │ '2 (default)' │ '21600 (default)' │ 'no failure handling (no async trigger found)' │ '' │
└─────────┴───────────────────────┴───────────┴───────────────┴─────────────────────────┴─────────────────┴───────────────┴───────────────────┴────────────────────────────────────────────────┴─────────────────────────┘
2 function configs with asynchronous triggers drop failed events without a record.
Wrote 6 rows to lambda-async.csv
Dry run: nothing changed. Add --apply --queue-arn <standard SQS queue ARN> to add on-failure destinations.
Names are illustrative. thumbnail-generator is fed by S3 with no protection: a real gap. invoice-mailer has a DLQ, so events survive but without the error details. report-builder shows why the qualifier column matters: its EventBridge rule targets the live alias, which already has a destination, so the $LATEST row is only at risk if something invokes the unqualified function. Adding a function-level destination anyway does no harm.
What should you do with failure records?
A Lambda on-failure destination nobody reads is only a slower way to lose events. Put a CloudWatch alarm on the queue’s ApproximateNumberOfMessagesVisible, and when it fires, read the records: requestContext.condition says whether retries were exhausted or the event expired, and responsePayload carries the error. The guide to ask AI about Lambda errors in your AWS account helps you find the cause from the terminal. Once fixed, replay the requestPayload with an asynchronous Invoke.
Keep the destination private. An SNS topic or SQS queue with a wide resource policy leaks every failed event; the script to find public SNS topics and SQS queues checks that. For S3 destinations, AWS recommends limiting the role’s s3:PutObject to buckets in your own account with the s3:ResourceAccount condition, so a deleted and re-created bucket name can’t collect your records.
Troubleshooting
- An error on apply. The Action column shows the error name and message. The usual causes are a mistyped queue ARN or a profile without
lambda:PutFunctionEventInvokeConfig. - Records never arrive. Check the function’s
DestinationDeliveryFailuresmetric. Lambda publishes it when the execution role lackssqs:SendMessage, the queue is FIFO, the record is too large, or the queue uses a customer managed KMS key the role can’t use. - Every function says “no async trigger found”. Your functions may be invoked by EventBridge Scheduler or other code, which leaves nothing in the resource-based policy. Compare with the functions that have
Errorsbut no clear caller, using the script to get Lambda invocation counts for the last 24 hours. TooManyRequestsException. Large accounts hit the Lambda control-plane rate. The SDK retries; if it keeps failing, run the report per Region at a quieter time.
Ask ChatWithCloud instead
For a quick look, ask ChatWithCloud “Which Lambda functions have no on-failure destination and no dead-letter queue?” It writes AWS SDK for JavaScript v2 code, runs it on your machine with your profile and explains the result, as how ChatWithCloud works describes. It works in one profile and Region per session and applies changes without a confirmation step, so connect ChatWithCloud to a read-only AWS profile and keep the --apply step in the reviewed script. The AWS practical examples hub has more Lambda audits, such as the one to find Lambda functions on deprecated runtimes.
Frequently asked questions
What is the difference between a Lambda DLQ and an on-failure destination?
A dead-letter queue stores only the event and is set once per function. An on-failure destination stores an invocation record with the event, the error and the attempt count, can be set per version or alias, and supports SQS, SNS, S3, Lambda and EventBridge.
Does a Lambda on-failure destination work for SQS triggers?
No. SQS triggers use an event source mapping, so failed messages go back to the queue. Configure a redrive policy with a dead-letter queue on the source queue instead.
How many times does Lambda retry an asynchronous invocation?
Twice by default, so three attempts in total, one minute and then two minutes apart. You can set MaximumRetryAttempts to 0, 1 or 2.
Can I use a FIFO queue as a Lambda destination?
No. Destinations and dead-letter queues accept standard SQS queues and standard SNS topics only.
Related guides
Ask your AWS account in plain English
Your first 15 runs are free, with no OpenAI key needed.
npx chatwithcloud