
Photo by William Warby on Pexels
A Lambda timeout is too high when it sits far above what the function actually needs. Compare each function’s Timeout from ListFunctions with its CloudWatch Duration p99 and Maximum over the last 14 days, fetched with GetMetricData. A timeout of 900 seconds on a function whose p99 is under a second means a hung call can run, and bill, for 15 minutes.
Timeouts get raised in a hurry and rarely lowered. A function timed out once during an incident, someone set it to the maximum, and it stayed there. Nothing breaks, so nobody notices, until a downstream service hangs and every invocation waits the full 15 minutes instead of failing fast.
This example is for teams that run many Lambda functions. The script compares every function’s timeout with its real runtime, flags any Lambda timeout too high for what the function does, flags functions whose longest run reaches the timeout, and checks the SQS visibility timeout for queue-triggered functions. It suggests a new value and changes nothing.
When is a Lambda timeout too high?
The timeout is the most time Lambda lets one invocation run: 3 seconds by default, adjustable in 1-second steps up to 900 seconds. AWS’s guidance cuts one way: a timeout close to the average duration risks unexpected timeouts. The opposite mistake has its own costs:
- Hung calls bill longer. You pay for
Duration, so an invocation stuck on a slow dependency is billed until it finishes or times out. - Hung calls hold concurrency. Each stuck invocation occupies an execution environment for its whole run, which can throttle other invocations in the same account and Region.
- Callers stopped waiting long ago. An API Gateway REST API integration times out at 29 seconds by default (Regional and private APIs can request more), and HTTP APIs have a fixed 30-second maximum. A synchronous function behind them gains nothing from a 900-second timeout.
- Failures surface late. A retry, a dead-letter queue or an on-failure destination only kicks in after the timeout.
The script’s rule of thumb: suggest the larger of the worst daily p99 times --headroom (default 2) and the observed maximum plus 20%, and flag a timeout that’s at least 3 times that suggestion and at least 10 seconds more. Tune both for your workloads.
What does a too-high timeout cost?
| Lambda price, US East (N. Virginia) | x86 | Arm |
|---|---|---|
| Duration, first pricing tier (per GB-second) | $0.0000166667 | $0.0000133334 |
| Requests (per 1 million) | $0.20 | $0.20 |
Prices as of September 2026 from the AWS Price List API (publication dated 19 September 2026); the first tier covers the first 6 billion GB-seconds a month on x86 and 7.5 billion on Arm. Check current rates on the AWS Lambda pricing page.
Worked example. orders-api has 1,024 MB of memory (1 GB) and a 900-second timeout. A payment provider hangs for half an hour and 2,000 invocations wait until they time out:
- 900-second timeout: 2,000 × 900 s × 1 GB = 1,800,000 GB-s × $0.0000166667 = $30.00
- 10-second timeout: 2,000 × 10 s × 1 GB = 20,000 GB-s × $0.0000166667 = $0.33
The money is small; the concurrency isn’t. Those 2,000 hung invocations could hold up to 2,000 execution environments for 15 minutes each. With a 10-second timeout they’d release them in seconds, and callers would get a fast error they can retry. If memory is also oversized, the example to find Lambda functions with too much memory cuts the GB side of the same calculation, and moving to Arm cuts the price; the example to find Lambda functions not running on Graviton lists candidates.
What does the script do?
- Lists functions
paginateListFunctionsreturns each function’sTimeout. - Reads Duration
paginateGetMetricDatafetches dailyp99andMaximumofDuration, plusInvocationsandErrorssums, four queries per function in batches of 100 functions. Lambda’sDurationmetric supports percentile statistics. - ClassifiesToo few invocations to judge, at the timeout (
Maximumat 95% or more of it), far above runtime, or fine. - Checks SQS triggersFor flagged functions,
ListEventSourceMappingsfinds SQS sources andGetQueueAttributesreadsVisibilityTimeout.
The Errors metric counts timeouts together with other function errors, so an “AT TIMEOUT” row tells you the limit was reached, not how many errors were timeouts. Check the function’s logs for that.
Prerequisites
- Node.js 18 or later with
tsx, plus@aws-sdk/client-lambda,@aws-sdk/client-cloudwatchand@aws-sdk/client-sqs. The guide to AWS SDK v3 paginators coverspaginateGetMetricData. - At least two weeks of traffic. Monthly jobs and rarely called functions need a longer
--daysor a manual look. - Know which functions sit behind API Gateway, since their useful timeout is capped by the API’s integration timeout.
Which IAM permissions does it need?
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ReadFunctionsAndMetrics",
"Effect": "Allow",
"Action": ["lambda:ListFunctions", "lambda:ListEventSourceMappings", "cloudwatch:GetMetricData"],
"Resource": "*"
},
{
"Sid": "ReadQueueVisibilityTimeout",
"Effect": "Allow",
"Action": ["sqs:GetQueueUrl", "sqs:GetQueueAttributes"],
"Resource": "arn:aws:sqs:*:111122223333:*"
}
]
}
All read-only. The IAM policy generator for TypeScript code builds a starting policy if you extend the script.
The script to find a Lambda timeout too high for its runtime
// find-lambda-functions-with-excessive-timeouts.ts
// Compares each Lambda function's configured timeout with how long it actually runs: the worst daily p99
// and the Maximum of the CloudWatch Duration metric over the last N days. Flags functions whose timeout is
// far above their runtime, and functions whose Maximum reaches the timeout (they are probably timing out).
// For SQS-triggered functions it also checks the queue's visibility timeout. Report only.
// Usage: npx tsx find-lambda-functions-with-excessive-timeouts.ts [--region us-east-1] [--days 14] [--headroom 2]
import {
LambdaClient,
paginateListEventSourceMappings,
paginateListFunctions,
type FunctionConfiguration,
} from "@aws-sdk/client-lambda";
import { CloudWatchClient, paginateGetMetricData, type MetricDataQuery } from "@aws-sdk/client-cloudwatch";
import { SQSClient, GetQueueAttributesCommand, GetQueueUrlCommand } from "@aws-sdk/client-sqs";
const args = process.argv.slice(2);
const flag = (name: string): string | undefined => {
const i = args.indexOf(name);
return i >= 0 ? args[i + 1] : undefined;
};
const region = flag("--region") ?? process.env.AWS_REGION ?? "us-east-1";
const days = Number(flag("--days") ?? "14");
const headroom = Number(flag("--headroom") ?? "2");
const minInvocations = Number(flag("--min-invocations") ?? "10");
if (!(days >= 1 && days <= 30) || !(headroom >= 1)) {
console.error("--days must be 1-30 and --headroom at least 1");
process.exit(1);
}
const lambda = new LambdaClient({ region });
const cloudwatch = new CloudWatchClient({ region });
const sqs = new SQSClient({ region });
const errText = (err: unknown): string => (err instanceof Error ? `${err.name}: ${err.message}` : String(err));
interface Stats {
p99: number; // worst daily p99 Duration, ms
max: number; // highest Maximum Duration, ms
invocations: number;
errors: number;
}
async function metrics(fns: FunctionConfiguration[]): Promise<Map<string, Stats>> {
const end = new Date();
const start = new Date(end.getTime() - days * 86_400_000);
const stats = new Map<string, Stats>();
const metric = (fn: string, name: string, stat: string, id: string): MetricDataQuery => ({
Id: id,
MetricStat: {
Metric: { Namespace: "AWS/Lambda", MetricName: name, Dimensions: [{ Name: "FunctionName", Value: fn }] },
Period: 86_400,
Stat: stat,
},
});
for (let i = 0; i < fns.length; i += 100) {
const batch = fns.slice(i, i + 100); // 4 queries per function, GetMetricData allows 500 per call
const queries = batch.flatMap((f, j) => [
metric(f.FunctionName ?? "", "Duration", "p99", `p${j}`),
metric(f.FunctionName ?? "", "Duration", "Maximum", `m${j}`),
metric(f.FunctionName ?? "", "Invocations", "Sum", `i${j}`),
metric(f.FunctionName ?? "", "Errors", "Sum", `e${j}`),
]);
const values = new Map<string, number[]>();
for await (const page of paginateGetMetricData({ client: cloudwatch }, { MetricDataQueries: queries, StartTime: start, EndTime: end })) {
for (const r of page.MetricDataResults ?? []) {
if (r.Id) values.set(r.Id, [...(values.get(r.Id) ?? []), ...(r.Values ?? [])]);
}
}
const top = (id: string): number => Math.max(0, ...(values.get(id) ?? []));
const sum = (id: string): number => (values.get(id) ?? []).reduce((a, b) => a + b, 0);
batch.forEach((f, j) => {
stats.set(f.FunctionName ?? "", { p99: top(`p${j}`), max: top(`m${j}`), invocations: sum(`i${j}`), errors: sum(`e${j}`) });
});
}
return stats;
}
// For SQS triggers: the queue visibility timeout, in seconds, plus the batching window
async function sqsTriggers(fn: string): Promise<{ queue: string; visibility: number; window: number }[]> {
const out: { queue: string; visibility: number; window: number }[] = [];
for await (const page of paginateListEventSourceMappings({ client: lambda }, { FunctionName: fn })) {
for (const esm of page.EventSourceMappings ?? []) {
const arn = esm.EventSourceArn ?? "";
if (!arn.startsWith("arn:aws:sqs:")) continue;
const [, , , , account, queueName] = arn.split(":");
const { QueueUrl } = await sqs.send(new GetQueueUrlCommand({ QueueName: queueName, QueueOwnerAWSAccountId: account }));
const attrs = await sqs.send(new GetQueueAttributesCommand({ QueueUrl, AttributeNames: ["VisibilityTimeout"] }));
out.push({ queue: queueName, visibility: Number(attrs.Attributes?.VisibilityTimeout ?? 0), window: esm.MaximumBatchingWindowInSeconds ?? 0 });
}
}
return out;
}
async function main(): Promise<void> {
const fns: FunctionConfiguration[] = [];
for await (const page of paginateListFunctions({ client: lambda }, {})) fns.push(...(page.Functions ?? []));
const stats = await metrics(fns);
const rows: Record<string, string | number>[] = [];
for (const f of fns) {
const name = f.FunctionName ?? "?";
const timeout = f.Timeout ?? 3;
const s = stats.get(name) ?? { p99: 0, max: 0, invocations: 0, errors: 0 };
const suggested = Math.min(900, Math.max(1, Math.ceil((s.p99 * headroom) / 1000), Math.ceil((s.max * 1.2) / 1000)));
let finding = "ok";
if (s.invocations < minInvocations) finding = `too few invocations (${s.invocations})`;
else if (s.max >= timeout * 1000 * 0.95) finding = `AT TIMEOUT: Maximum reaches the limit (${s.errors} errors)`;
else if (timeout >= suggested * 3 && timeout - suggested >= 10) finding = "TIMEOUT FAR ABOVE RUNTIME";
let sqsNote = "-";
if (finding !== "ok" && !finding.startsWith("too few")) {
try {
const triggers = await sqsTriggers(name);
sqsNote = triggers
.map((t) => `${t.queue}: ${t.visibility}s${t.visibility < 6 * timeout + t.window ? " (<6x timeout)" : ""}`)
.join("; ") || "-";
} catch (err) {
sqsNote = `error: ${errText(err)}`;
}
}
rows.push({
Function: name,
TimeoutS: timeout,
P99S: Math.round(s.p99) / 1000,
MaxS: Math.round(s.max) / 1000,
Invocations: s.invocations,
SuggestedS: finding === "ok" || finding.startsWith("too few") ? "-" : suggested,
SQS: sqsNote,
Finding: finding,
});
}
rows.sort((a, b) => Number(b.TimeoutS) - Number(a.TimeoutS));
console.table(rows);
const high = rows.filter((r) => r.Finding === "TIMEOUT FAR ABOVE RUNTIME").length;
const atLimit = rows.filter((r) => String(r.Finding).startsWith("AT TIMEOUT")).length;
console.log(`${fns.length} functions in ${region} over ${days} days: ${high} with a timeout far above runtime, ${atLimit} hitting their timeout.`);
console.log("Report only: nothing was changed.");
}
main().catch((err) => {
console.error(errText(err));
process.exit(1);
});
Using the worst daily p99 rather than one 14-day p99 keeps a single bad day from being averaged away, and the maximum-plus-20% floor means the suggestion never cuts below the longest run the function has actually had.
How do you run it?
npm install @aws-sdk/client-lambda @aws-sdk/client-cloudwatch @aws-sdk/client-sqs
npm install --save-dev tsx typescript @types/node
AWS_PROFILE=readonly npx tsx find-lambda-functions-with-excessive-timeouts.ts --region us-east-1 --days 14 --headroom 2
Sample output
┌─────────┬───────────────────┬──────────┬───────┬──────┬─────────────┬────────────┬────────────────────────────────┬─────────────────────────────────────────────────────┐
│ (index) │ Function │ TimeoutS │ P99S │ MaxS │ Invocations │ SuggestedS │ SQS │ Finding │
├─────────┼───────────────────┼──────────┼───────┼──────┼─────────────┼────────────┼────────────────────────────────┼─────────────────────────────────────────────────────┤
│ 0 │ 'orders-api' │ 900 │ 0.82 │ 2.4 │ 1260000 │ 3 │ '-' │ 'TIMEOUT FAR ABOVE RUNTIME' │
│ 1 │ 'nightly-report' │ 600 │ 412 │ 455 │ 14 │ '-' │ '-' │ 'ok' │
│ 2 │ 'thumbnailer' │ 300 │ 5.1 │ 9.8 │ 56000 │ 12 │ 'thumbnail-jobs: 1800s' │ 'TIMEOUT FAR ABOVE RUNTIME' │
│ 3 │ 'old-webhook' │ 60 │ 0.3 │ 0.4 │ 0 │ '-' │ '-' │ 'too few invocations (0)' │
│ 4 │ 'invoice-worker' │ 30 │ 21 │ 30 │ 16800 │ 42 │ 'invoices: 180s (<6x timeout)' │ 'AT TIMEOUT: Maximum reaches the limit (17 errors)' │
│ 5 │ 'auth-authorizer' │ 10 │ 0.095 │ 0.31 │ 700000 │ '-' │ '-' │ 'ok' │
└─────────┴───────────────────┴──────────┴───────┴──────┴─────────────┴────────────┴────────────────────────────────┴─────────────────────────────────────────────────────┘
6 functions in us-east-1 over 14 days: 2 with a timeout far above runtime, 1 hitting their timeout.
Report only: nothing was changed.
Function names and numbers are illustrative. orders-api is the worked example: 900 seconds of timeout for a p99 under a second, behind an API that stops waiting at 29. thumbnailer reads from SQS and could drop from 300 to 12 seconds; its queue’s visibility timeout could then come down too. invoice-worker is the opposite problem: its longest run hits the 30-second limit, and its queue’s 180-second visibility timeout is already below six times the timeout plus the batching window. nightly-report runs once a day for about 7 minutes and is fine at 600 seconds.
How do you change a timeout safely?
- Lower it in steps. Set the suggestion, deploy, and watch
ErrorsandDurationfor a few days. Change the value in your template (SAMTimeout, CDKtimeout) rather than the console so the next deploy doesn’t undo it. - Set SDK timeouts below the function timeout. A lower Lambda timeout only helps if calls inside the function give up first and handle the error. The guide to configure retry and timeout settings in AWS SDK for JavaScript v3 shows
requestTimeoutandmaxAttempts. - Keep SQS in step. AWS recommends a queue visibility timeout of at least six times the function timeout, plus
MaximumBatchingWindowInSeconds. Lambda also rejects an event source mapping whose function timeout is longer than the queue’s visibility timeout. Raise the queue first when you raise a function’s timeout. Pair this with the example to find Lambda SQS triggers without ReportBatchItemFailures, so one slow message doesn’t send the whole batch back. - Make failures go somewhere. Timed-out asynchronous invocations are retried and then dropped unless you configure a destination; the example to find Lambda functions without an async failure destination finds the gaps, and the one to find SQS queues without a dead-letter queue covers queue sources.
Troubleshooting
- Every function shows “too few invocations”. Check the Region and that the profile can call
cloudwatch:GetMetricData. Metrics are per Region, like functions. - “AT TIMEOUT” on a function you thought was fine. Look at its errors around the dates of the longest runs. The walkthrough to investigate Lambda errors with CloudWatch shows how to find the matching log lines.
- Suggested value seems too low for a batch job. The script only knows what has happened in the window. Raise
--headroomor--daysfor jobs whose input grows over time. - Aliases and versions. The script reads metrics by function name, which aggregates all versions. Timeouts are per version, so if old versions still take traffic, check them too.
Ask ChatWithCloud instead
For a quick check, ask ChatWithCloud “Which Lambda functions in us-east-1 have a timeout over 60 seconds but a maximum duration under 5 seconds this week?” It writes AWS SDK for JavaScript v2 code, runs it on your machine with your profile and answers from the metrics. The guide to ask AI about Lambda errors in your AWS account covers related questions. It can be wrong and runs changes without a confirmation step, so don’t ask it to update timeouts; use a read-only profile, as the ChatWithCloud security page recommends.
Frequently asked questions
What is the maximum Lambda timeout?
900 seconds (15 minutes). The default is 3 seconds, and you set it in 1-second steps.
Does a high Lambda timeout cost more if the function finishes quickly?
No. You pay for the time the function actually runs. A high timeout only costs more when an invocation hangs or runs long, because it’s allowed to keep running and billing.
What timeout should a Lambda function behind API Gateway have?
No more than the API’s integration timeout: 29 seconds by default for REST APIs and at most 30 seconds for HTTP APIs. Usually well below that, based on the function’s p99.
How should the SQS visibility timeout relate to the Lambda timeout?
AWS recommends at least six times the function timeout, plus the batching window. The function timeout must not be longer than the visibility timeout.
Related guides
Ask your AWS account in plain English
Your first 15 runs are free, with no OpenAI key needed.
npx chatwithcloud