
Photo by Galib Rahman Nadim on Pexels
Lambda SQS ReportBatchItemFailures lets a function tell Lambda which messages in a batch failed, so only those return to the queue. Without it, one bad message makes the whole batch visible again, including messages already processed. Find triggers without it by listing event source mappings and checking FunctionResponseTypes on every mapping whose source is an SQS queue.
An SQS trigger that retries a whole batch works fine in testing and badly in production. One malformed message fails ten, the nine good ones run again, and anything that isn’t idempotent, such as an email or a payment, happens twice.
This example is for engineers running queue-driven Lambda functions. The script lists every SQS event source mapping in a Region and reports whether Lambda SQS ReportBatchItemFailures is on, the batch size and window, the function timeout against the queue’s visibility timeout, and whether the queue has a dead-letter queue. With --apply --confirm-handler it turns partial batch responses on for the mappings you name. You’ll also get a TypeScript handler that returns failures in the format Lambda expects.
How does ReportBatchItemFailures change Lambda SQS retries?
By default, if your function hits an error anywhere in a batch, every message in that batch becomes visible in the queue again after the visibility timeout. With ReportBatchItemFailures in the mapping’s FunctionResponseTypes, your function catches errors per message and returns the IDs that failed:
{
"batchItemFailures": [
{ "itemIdentifier": "id2" },
{ "itemIdentifier": "id4" }
]
}
Lambda deletes the rest. The exact rules decide whether this helps or hurts:
| Function returns | Lambda treats the batch as |
|---|---|
An empty or null batchItemFailures list, or an empty or null response |
Complete success: every message is deleted |
| A list of real message IDs | Partial success: only those messages return to the queue |
Invalid JSON, an empty or null itemIdentifier, a wrong key name, or an ID not in the batch |
Complete failure: every message returns |
| A thrown exception | Complete failure |
The mapping setting and the handler have to change together. A handler that returns batchItemFailures on a mapping without the setting gets its list ignored, and a normal return deletes the whole batch, failed messages included. That’s the worst case this audit looks for. The reverse, turning the setting on for a handler that doesn’t return the list, gains nothing. The Lambda guide to handling errors for an SQS event source has the full rules.
Which other settings does the script check?
- Visibility timeout. AWS recommends at least six times the function timeout, plus
MaximumBatchingWindowInSeconds, so throttled batches can be retried before messages reappear. Lambda rejects a mapping whose function timeout exceeds the visibility timeout, but anything between 1× and 6× passes silently. A function timeout far above what the code needs inflates that requirement too; the script to find Lambda functions with a timeout too high for their runtime suggests a tighter value. - Dead-letter queue. A redrive policy on the source queue moves a message out after
maxReceiveCountreceives; AWS suggests at least 5. Without one, a poison message retries until it expires. The script to find SQS queues without a dead-letter queue covers queues with no Lambda trigger too, and the audit of public SNS topics and SQS queues checks who else can send to them. - Batch size and window. Standard queues allow up to 10,000 messages per batch and FIFO queues 10; above 10, the window must be at least 1 second. Larger batches make whole-batch retries more expensive.
One more reason to turn it on: with ReportBatchItemFailures active, Lambda doesn’t scale down polling when invocations fail, so a few bad messages don’t slow the whole queue.
What does the script do?
- Lists mappings
paginateListEventSourceMappingsand keeps those whoseEventSourceArnis an SQS queue. - Reads the function timeout
GetFunctionConfigurationper function, cached. - Reads the queue
GetQueueUrlfrom the queue name and owner account in the ARN, thenGetQueueAttributesforVisibilityTimeoutandRedrivePolicy. - ReportsPartial responses, batch size and window, timeout versus the recommended visibility timeout, DLQ and mapping state.
- Updates, if asked
--apply --uuids ... --confirm-handlercallsUpdateEventSourceMappingwithFunctionResponseTypes: ["ReportBatchItemFailures"], thenGetEventSourceMappingto confirm. Without--confirm-handlerit refuses.
Prerequisites
- Node.js 18 or later with
tsx, plus@aws-sdk/client-lambdaand@aws-sdk/client-sqs. - A read-only profile for the report. The guide to SDK v3 credential providers covers profiles and SSO.
- Before
--apply: the handler below (or the Batch Processor utility from Powertools for AWS Lambda, available for TypeScript, Python, Java and .NET) deployed to the function.
Which IAM permissions does it need?
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ListMappings",
"Effect": "Allow",
"Action": "lambda:ListEventSourceMappings",
"Resource": "*"
},
{
"Sid": "ReadFunctionsAndQueues",
"Effect": "Allow",
"Action": ["lambda:GetFunctionConfiguration", "sqs:GetQueueUrl", "sqs:GetQueueAttributes"],
"Resource": [
"arn:aws:lambda:*:111122223333:function:*",
"arn:aws:sqs:*:111122223333:*"
]
},
{
"Sid": "UpdateOnlyWithApply",
"Effect": "Allow",
"Action": ["lambda:UpdateEventSourceMapping", "lambda:GetEventSourceMapping"],
"Resource": "arn:aws:lambda:*:111122223333:event-source-mapping:*"
}
]
}
For queues owned by another account, that account’s queue policy must also allow the two SQS read actions, or the DLQ column shows unknown.
The script to find Lambda SQS triggers without ReportBatchItemFailures
// find-sqs-triggers-without-partial-batch.ts
// Lists every Lambda event source mapping that reads from SQS in a Region and checks: ReportBatchItemFailures,
// batch size and window, queue visibility timeout against 6x the function timeout (plus the batch window),
// and whether the queue has a dead-letter queue. --apply turns on ReportBatchItemFailures for the mappings
// you name, and only with --confirm-handler (the function code must already return batchItemFailures).
// Usage:
// npx tsx find-sqs-triggers-without-partial-batch.ts [--region us-east-1]
// npx tsx find-sqs-triggers-without-partial-batch.ts --region us-east-1 --apply --uuids <uuid> --confirm-handler
import {
LambdaClient,
GetEventSourceMappingCommand,
GetFunctionConfigurationCommand,
UpdateEventSourceMappingCommand,
paginateListEventSourceMappings,
} from "@aws-sdk/client-lambda";
import { SQSClient, GetQueueAttributesCommand, GetQueueUrlCommand } from "@aws-sdk/client-sqs";
const args = process.argv.slice(2);
const flag = (name: string): string | undefined => {
const i = args.indexOf(name);
return i >= 0 ? args[i + 1] : undefined;
};
const region = flag("--region") ?? process.env.AWS_REGION ?? "us-east-1";
const apply = args.includes("--apply");
const confirmHandler = args.includes("--confirm-handler");
const uuids = (flag("--uuids") ?? "").split(",").map((u) => u.trim()).filter(Boolean);
const lambda = new LambdaClient({ region });
const sqs = new SQSClient({ region });
interface Row {
UUID: string;
Function: string;
Queue: string;
Partial: string;
Batch: string;
FnTimeout: number | string;
Visibility: number | string;
NeedVisibility: number | string;
DLQ: string;
Findings: string;
}
const errText = (err: unknown): string => (err instanceof Error ? `${err.name}: ${err.message}` : String(err));
const timeouts = new Map<string, number | undefined>();
async function functionTimeout(functionArn: string): Promise<number | undefined> {
if (!timeouts.has(functionArn)) {
try {
const fn = await lambda.send(new GetFunctionConfigurationCommand({ FunctionName: functionArn }));
timeouts.set(functionArn, fn.Timeout);
} catch {
timeouts.set(functionArn, undefined);
}
}
return timeouts.get(functionArn);
}
async function queueSettings(queueArn: string): Promise<{ visibility?: number; dlq: string }> {
const [, , , , account, name] = queueArn.split(":");
try {
const { QueueUrl } = await sqs.send(new GetQueueUrlCommand({ QueueName: name, QueueOwnerAWSAccountId: account }));
const out = await sqs.send(
new GetQueueAttributesCommand({ QueueUrl, AttributeNames: ["VisibilityTimeout", "RedrivePolicy"] }),
);
const redrive = out.Attributes?.RedrivePolicy;
const maxReceive = redrive ? (JSON.parse(redrive) as { maxReceiveCount?: number | string }).maxReceiveCount : undefined;
return {
visibility: Number(out.Attributes?.VisibilityTimeout ?? NaN),
dlq: redrive ? `yes (maxReceiveCount ${maxReceive})` : "none",
};
} catch (err) {
return { dlq: `unknown (${err instanceof Error ? err.name : "error"})` };
}
}
async function main(): Promise<void> {
const rows: Row[] = [];
for await (const page of paginateListEventSourceMappings({ client: lambda }, {})) {
for (const m of page.EventSourceMappings ?? []) {
if (!m.UUID || !m.EventSourceArn || m.EventSourceArn.split(":")[2] !== "sqs" || !m.FunctionArn) continue;
const partial = (m.FunctionResponseTypes ?? []).includes("ReportBatchItemFailures");
const window = m.MaximumBatchingWindowInSeconds ?? 0;
const timeout = await functionTimeout(m.FunctionArn);
const queue = await queueSettings(m.EventSourceArn);
const need = timeout !== undefined ? timeout * 6 + window : undefined;
const findings: string[] = [];
if (!partial) findings.push("no partial batch response");
if (need !== undefined && queue.visibility !== undefined && !Number.isNaN(queue.visibility) && queue.visibility < need)
findings.push(`visibility < ${need}s`);
if (queue.dlq === "none") findings.push("no DLQ");
if (m.State && m.State !== "Enabled") findings.push(`mapping ${m.State}`);
rows.push({
UUID: m.UUID,
Function: m.FunctionArn.split(":function:")[1] ?? m.FunctionArn,
Queue: m.EventSourceArn.split(":")[5] ?? m.EventSourceArn,
Partial: partial ? "yes" : "NO",
Batch: `${m.BatchSize ?? "?"} / ${window}s`,
FnTimeout: timeout ?? "?",
Visibility: queue.visibility ?? "?",
NeedVisibility: need ?? "?",
DLQ: queue.dlq,
Findings: findings.join("; ") || "ok",
});
}
}
console.table(rows);
const missing = rows.filter((r) => r.Partial === "NO");
console.log(`${rows.length} SQS triggers in ${region}: ${missing.length} without ReportBatchItemFailures.`);
if (!apply) {
console.log("Report only: nothing was changed. Use --apply --uuids <uuid,uuid> --confirm-handler to turn it on.");
return;
}
if (!confirmHandler) {
console.error("Refusing to change mappings without --confirm-handler: update the function to return batchItemFailures first.");
process.exit(1);
}
for (const uuid of uuids) {
if (!missing.some((r) => r.UUID === uuid)) {
console.error(`Skipping ${uuid}: not an SQS trigger without partial batch responses in ${region}`);
continue;
}
try {
await lambda.send(new UpdateEventSourceMappingCommand({ UUID: uuid, FunctionResponseTypes: ["ReportBatchItemFailures"] }));
const after = await lambda.send(new GetEventSourceMappingCommand({ UUID: uuid }));
console.log(`Updated ${uuid}: FunctionResponseTypes ${JSON.stringify(after.FunctionResponseTypes ?? [])}, state ${after.State}`);
} catch (err) {
console.error(`Could not update ${uuid}: ${errText(err)}`);
process.exitCode = 1;
}
}
}
main().catch((err) => {
console.error(errText(err));
process.exit(1);
});
What should the handler return?
Catch errors per message, collect the failed messageId values and return them. Don’t throw from the handler: an exception fails the whole batch. Types come from @types/aws-lambda.
// handler.ts: an SQS-triggered Lambda handler that reports partial batch failures.
// Requires ReportBatchItemFailures on the event source mapping. Types: npm install -D @types/aws-lambda
import type { SQSBatchItemFailure, SQSBatchResponse, SQSEvent, SQSRecord } from "aws-lambda";
async function processMessage(record: SQSRecord): Promise<void> {
const order = JSON.parse(record.body) as { orderId?: string };
if (!order.orderId) throw new Error(`Message ${record.messageId} has no orderId`);
// Do the real work here, and make it idempotent: SQS delivers at least once.
}
export const handler = async (event: SQSEvent): Promise<SQSBatchResponse> => {
const batchItemFailures: SQSBatchItemFailure[] = [];
for (const record of event.Records) {
try {
await processMessage(record);
} catch (err) {
console.error(`Failed ${record.messageId}:`, err);
batchItemFailures.push({ itemIdentifier: record.messageId });
// On a FIFO queue, stop here and report every remaining message as failed to keep the order.
}
}
return { batchItemFailures };
};
On a FIFO queue, stop at the first failure and return that message and every one after it, so ordering is kept. Unit test the failure path before you deploy; the guide to mocking AWS SDK v3 clients in unit tests shows how to fake the calls inside processMessage.
How do you run it?
npm install @aws-sdk/client-lambda @aws-sdk/client-sqs
npm install --save-dev tsx typescript @types/node @types/aws-lambda
# Report only
AWS_PROFILE=readonly npx tsx find-sqs-triggers-without-partial-batch.ts --region us-east-1
# After deploying the new handler to process-orders
AWS_PROFILE=lambda-admin npx tsx find-sqs-triggers-without-partial-batch.ts --region us-east-1 \
--apply --uuids 6f1c2d3e-0000-4a1b-9c2d-111111111111 --confirm-handler
Sample output
┌─────────┬────────────────────────────────────────┬──────────────────┬──────────┬─────────┬────────────┬───────────┬────────────┬────────────────┬───────────────────────────┬────────────────────────────────────────────────────────┐
│ (index) │ UUID │ Function │ Queue │ Partial │ Batch │ FnTimeout │ Visibility │ NeedVisibility │ DLQ │ Findings │
├─────────┼────────────────────────────────────────┼──────────────────┼──────────┼─────────┼────────────┼───────────┼────────────┼────────────────┼───────────────────────────┼────────────────────────────────────────────────────────┤
│ 0 │ '6f1c2d3e-0000-4a1b-9c2d-111111111111' │ 'process-orders' │ 'orders' │ 'NO' │ '10 / 0s' │ 30 │ 30 │ 180 │ 'none' │ 'no partial batch response; visibility < 180s; no DLQ' │
│ 1 │ '8a9b0c1d-0000-4e2f-8a3b-222222222222' │ 'send-emails' │ 'emails' │ 'yes' │ '100 / 5s' │ 60 │ 900 │ 365 │ 'yes (maxReceiveCount 5)' │ 'ok' │
└─────────┴────────────────────────────────────────┴──────────────────┴──────────┴─────────┴────────────┴───────────┴────────────┴────────────────┴───────────────────────────┴────────────────────────────────────────────────────────┘
2 SQS triggers in us-east-1: 1 without ReportBatchItemFailures.
Report only: nothing was changed. Use --apply --uuids <uuid,uuid> --confirm-handler to turn it on.
Names and UUIDs are illustrative. process-orders has all three problems: whole-batch retries, a 30-second visibility timeout that equals the 30-second function timeout (the recommendation is 180), and no DLQ. Fix the queue first by raising the visibility timeout and adding a redrive policy, then deploy the handler, then run --apply. send-emails is configured well: 6 × 60 + 5 = 365 seconds needed, 900 set.
Troubleshooting
- Failed messages don’t come back after enabling it. The handler is returning an empty list or nothing, which Lambda treats as success. Log the returned IDs and check them against
record.messageId; the guide to investigating Lambda errors with CloudWatch shows where those logs land. - Every message keeps coming back. The response is invalid (wrong key, empty ID, an ID from another batch) or the handler throws. Both count as a complete failure.
- The mapping shows
Updating. Changes take a moment to apply. Run the report again after a minute. - The DLQ column shows
unknown. The queue is in another account or the policy lackssqs:GetQueueUrl. The guide to troubleshooting IAM access denied errors helps narrow it down.
Ask ChatWithCloud instead
For a quick read, ask ChatWithCloud “Which Lambda SQS triggers in us-east-1 don’t have ReportBatchItemFailures, and what are their queues’ visibility timeouts?” It writes AWS SDK for JavaScript v2 code, runs it on your machine with your profile and summarizes the answer. When a trigger is already failing, the walkthrough on how to ask AI about Lambda errors in your AWS account covers that case. Use a read-only profile: changes run without a confirmation step.
Frequently asked questions
How do I enable ReportBatchItemFailures for a Lambda SQS trigger?
Update the event source mapping with FunctionResponseTypes set to ["ReportBatchItemFailures"], for example aws lambda update-event-source-mapping --uuid ... --function-response-types ReportBatchItemFailures, and deploy a handler that returns batchItemFailures.
What happens if my Lambda throws with ReportBatchItemFailures on?
Lambda treats the whole batch as failed, and every message returns to the queue. Catch errors per message and return their IDs instead.
What should the SQS visibility timeout be for a Lambda trigger?
At least six times the function timeout, plus the batching window if you use one. Lambda rejects mappings where the function timeout is longer than the visibility timeout.
Does ReportBatchItemFailures work with FIFO queues?
Yes. Stop processing at the first failure and return that message and all unprocessed ones in batchItemFailures, so the queue keeps its order.
Related guides
Ask your AWS account in plain English
Your first 15 runs are free, with no OpenAI key needed.
npx chatwithcloud