Photo by Nathana Rebouças on Unsplash
To run CloudWatch Logs Insights with AWS SDK v3, send StartQueryCommand from @aws-sdk/client-cloudwatch-logs with the log groups, a query string and a time range in epoch seconds. It returns a queryId. Poll GetQueryResultsCommand with backoff until status is Complete, then map each row’s field/value pairs to an object. Call StopQueryCommand if you give up.
Logs Insights is asynchronous. There’s no single call that returns results: you start a query, wait, and fetch. That’s easy to get almost right: poll too fast and you hit the rate limit, pass milliseconds and get nothing back, stop waiting without stopping the query and it keeps scanning.
This guide is for TypeScript and Node.js developers who want CloudWatch Logs Insights in AWS SDK v3 code: runbooks, scheduled reports, CI checks. You’ll build a reusable runInsightsQuery helper with backoff, cancellation and pagination, see which quotas apply, and work out what each query costs. The code was type-checked with strict tsc and run against aws-sdk-client-mock with @aws-sdk/client-cloudwatch-logs 3.1141.0 in September 2026. If you’re here to debug a function, the guide to investigate Lambda errors with CloudWatch has the queries; this one is about running them from code.
How does a Logs Insights query run through the API?
StartQuerytakes the log groups (or aSOURCEcommand in the query),startTime,endTime,queryStringand an optionallimit, and returns aqueryId.GetQueryResultsreturnsstatus:Scheduled,Running,Complete,Failed,Cancelled,TimeoutorUnknown. While it’sRunning, you get partial results.- When it’s
Complete,resultsholds the rows andstatisticsholdsrecordsScanned,recordsMatchedandbytesScanned. StopQuerycancels a query that’s still running. Queries also time out on their own after 60 minutes.
The same pattern (start, poll, read) is how running an Athena query with AWS SDK v3 works, so the helper below will look familiar if you’ve written that one.
Prerequisites
- Node.js 18 or later and TypeScript, with
@aws-sdk/client-cloudwatch-logs. - Credentials the SDK can resolve; see AWS SDK v3 credential providers: fromIni, fromSSO and assume role. The client’s Region must be the log groups’ Region.
- Log groups with data in the time range you query. Empty ones cost nothing to scan but also return nothing; the script to find and delete empty or abandoned CloudWatch log groups cleans those up.
How to run a CloudWatch Logs Insights query with AWS SDK v3, step by step
- Convert times to epoch seconds
Math.floor(date.getTime() / 1000). The API reads the number as seconds, so a millisecond value describes a range far in the future. - Pick one way to name log groupsPass exactly one of
logGroupName,logGroupNamesorlogGroupIdentifiers. Identifiers accept names or ARNs, up to 50. - Start the queryKeep the
queryId. IfStartQuerythrowsLimitExceededException, wait and retry. - Poll with backoffStart at about 500 ms and grow to a few seconds. Continue while the status is
ScheduledorRunning. - Read and flatten rowsEach row is an array of
{ field, value }. Turn it into an object and drop@ptrunless you need the full record. - Stop what you abandonOn timeout, cancellation or error, call
StopQueryso the query doesn’t keep running.
Example: a reusable Logs Insights helper
// logs-insights.ts: run CloudWatch Logs Insights queries with AWS SDK for JavaScript v3
import {
CloudWatchLogsClient,
GetQueryResultsCommand,
LimitExceededException,
StartQueryCommand,
StopQueryCommand,
type QueryStatistics,
} from "@aws-sdk/client-cloudwatch-logs";
const logs = new CloudWatchLogsClient({});
export interface InsightsQuery {
logGroups: string[]; // names or ARNs, up to 50
query: string;
start: Date;
end: Date;
limit?: number; // rows to return, 1 to 10,000 per page (default 1,000)
maxWaitMs?: number; // give up (and StopQuery) after this long
signal?: AbortSignal; // cancel from the caller, e.g. on Ctrl+C
}
export interface InsightsResult {
rows: Record<string, string>[];
statistics?: QueryStatistics;
status: string;
}
const sleep = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));
/** StartQuery, retrying LimitExceededException (too many concurrent queries) with backoff. */
async function start(q: InsightsQuery): Promise<string> {
for (let attempt = 1; ; attempt++) {
try {
const { queryId } = await logs.send(new StartQueryCommand({
logGroupIdentifiers: q.logGroups,
queryString: q.query,
startTime: Math.floor(q.start.getTime() / 1000), // epoch seconds, not milliseconds
endTime: Math.floor(q.end.getTime() / 1000),
limit: q.limit ?? 1000,
}));
if (!queryId) throw new Error("StartQuery returned no queryId");
return queryId;
} catch (err) {
if (!(err instanceof LimitExceededException) || attempt >= 5) throw err;
await sleep(Math.min(30_000, 1000 * 2 ** attempt) * (0.5 + Math.random() / 2));
}
}
}
/** Each result row is a list of { field, value }; turn it into an object and drop the @ptr pointer. */
function toRow(fields: { field?: string; value?: string }[]): Record<string, string> {
const row: Record<string, string> = {};
for (const { field, value } of fields) {
if (field && field !== "@ptr") row[field] = value ?? "";
}
return row;
}
export async function runInsightsQuery(q: InsightsQuery): Promise<InsightsResult> {
const queryId = await start(q);
const deadline = Date.now() + (q.maxWaitMs ?? 5 * 60_000);
let delay = 500;
try {
for (;;) {
if (q.signal?.aborted) throw new Error("Query aborted by caller");
if (Date.now() > deadline) throw new Error(`Query ${queryId} still running after ${q.maxWaitMs ?? 300_000} ms`);
await sleep(delay);
delay = Math.min(delay * 1.5, 5000); // 0.5 s, 0.75 s, 1.1 s ... capped at 5 s
const res = await logs.send(new GetQueryResultsCommand({ queryId }));
if (res.status === "Scheduled" || res.status === "Running") continue;
if (res.status !== "Complete") throw new Error(`Query ${queryId} ended with status ${res.status}`);
const rows = (res.results ?? []).map(toRow);
// More than 10,000 rows (Logs Insights QL only): page through with nextToken.
for (let token = res.nextToken; token; ) {
const next = await logs.send(new GetQueryResultsCommand({ queryId, nextToken: token }));
rows.push(...(next.results ?? []).map(toRow));
token = next.nextToken;
}
return { rows, statistics: res.statistics, status: res.status };
}
} catch (err) {
// Don't leave a query running (and scanning) after we stop waiting for it.
await logs.send(new StopQueryCommand({ queryId })).catch(() => undefined);
throw err;
}
}
/** us-east-1 Logs Insights price, $0.005 per GB scanned (AWS Price List, September 2026). */
export const scanCostUsd = (bytes = 0) => (bytes / 1024 ** 3) * 0.005;
async function main(): Promise<void> {
const [fn = "checkout-handler", hours = "24"] = process.argv.slice(2);
const controller = new AbortController();
process.once("SIGINT", () => controller.abort());
const end = new Date();
const start = new Date(end.getTime() - Number(hours) * 3_600_000);
const errors = await runInsightsQuery({
logGroups: [`/aws/lambda/${fn}`],
query: `fields @timestamp, @requestId, @message
| filter @message like /(?i)(error|exception|task timed out)/
| sort @timestamp desc
| limit 20`,
start,
end,
signal: controller.signal,
});
for (const r of errors.rows) console.log(`${r["@timestamp"]} ${r["@requestId"] ?? "-"} ${(r["@message"] ?? "").trim()}`);
const latency = await runInsightsQuery({
logGroups: [`/aws/lambda/${fn}`],
query: `filter @type = "REPORT"
| stats count(*) as invocations, avg(@duration) as avgMs, pct(@duration, 99) as p99Ms, max(@duration) as maxMs by bin(1h)`,
start,
end,
signal: controller.signal,
});
console.table(latency.rows);
const scanned = (errors.statistics?.bytesScanned ?? 0) + (latency.statistics?.bytesScanned ?? 0);
console.log(`Scanned ${(scanned / 1024 ** 2).toFixed(1)} MB, about $${scanCostUsd(scanned).toFixed(4)}`);
}
if (process.argv[1]?.endsWith("logs-insights.ts")) {
main().catch((err) => {
console.error(err);
process.exit(1);
});
}
Run it with AWS_REGION=eu-west-1 npx tsx logs-insights.ts checkout-handler 24. The main function runs two queries against one Lambda log group: the 20 newest error lines, and hourly invocation count, average, p99 and maximum duration from the REPORT lines. Ctrl+C aborts through an AbortSignal, which the MDN reference for AbortController describes, and the helper stops the query on its way out.
The polling loop uses GetQueryResultsCommand directly rather than a paginator, because the thing you wait for is a status change, not a next page. Once the query is Complete, nextToken pages work like any other API; the guide to paginate any AWS API with AWS SDK v3 paginators covers the general pattern.
Reading results: fields, @ptr and statistics
- Values are strings.
statsresults such asp99Mscome back as"812.4". Convert withNumber()before comparing. - Only requested fields are returned, plus
@ptr, a pointer you can pass toGetLogRecordto fetch the whole event. - Rows can differ in shape. A field missing from an event is missing from its row, so read with a default.
bytesScannedis what you pay for. Log it next to each run, and alert when a scheduled query starts scanning much more than usual.
Which limits and quotas apply?
| Limit | Value |
|---|---|
| Log groups per query | 50 with logGroupNames or logGroupIdentifiers |
| Query runtime | Times out after 60 minutes |
| Concurrent queries | Up to 100, including dashboard queries and scheduled query runs |
| Rows returned | Up to 10,000 per GetQueryResults call; up to 100,000 per query through nextToken (Logs Insights QL only) |
nextToken lifetime |
1 hour |
| Query string length | 10,000 characters |
StartQuery and GetQueryResults rate |
10 requests per second each, per account and Region; not adjustable |
Values are from the CloudWatch Logs API reference and quotas page, checked in September 2026. The request-rate quotas are why the helper backs off: a single loop polling every 100 ms would use the whole GetQueryResults budget for the account and Region. The guide to monitor AWS service quota usage and get alerts helps if you run many scheduled queries, and the SDK’s own throttling retries are covered in configure retry and timeout settings in AWS SDK for JavaScript v3.
What does a Logs Insights query cost?
You pay per GB of log data scanned. As of September 2026, the AWS Price List shows $0.005 per GB in US East (N. Virginia), and the CloudWatch free tier includes 5 GB of log data per month shared across ingestion, archive storage and Logs Insights scanning. The guide on what Lambda log queries cost has the full CloudWatch Logs price table.
Worked example: a scheduled report scans 7 days of logs from 20 log groups, about 60 GB, once an hour. That’s 60 × $0.005 = $0.30 a run, 24 × 30 = 720 runs a month, and 720 × $0.30 = $216 a month. Narrow it to the last hour of logs and each run scans about 60 ÷ 168 ≈ 0.36 GB, or $0.0018; the month drops to about $1.29. The time range is the biggest lever. Fewer log groups and shorter retention help too; the script to set CloudWatch log retention for all log groups handles the latter. To check what you’re spending, a Cost Explorer script can get this month’s CloudWatch cost.
Useful queries to run from code
fields @log
| filter @message like /(?i)(error|exception)/
| stats count(*) as errors by @log
| sort errors desc
filter @type = "REPORT"
| fields @requestId, @duration, @maxMemoryUsed
| sort @duration desc
| limit 10
The first works across up to 50 log groups at once, which is what you want when a request passes through several services. The second pairs well with the helper’s p99 query: once you know the hour that was slow, it names the requests. For turning a count into an alarm, the guide to publish CloudWatch custom metrics with PutMetricData shows how to push the number back to CloudWatch.
Permissions needed
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "StartQueriesOnLambdaLogs",
"Effect": "Allow",
"Action": "logs:StartQuery",
"Resource": "arn:aws:logs:us-east-1:123456789012:log-group:/aws/lambda/*"
},
{
"Sid": "ReadAndStopQueries",
"Effect": "Allow",
"Action": ["logs:GetQueryResults", "logs:StopQuery"],
"Resource": "*"
}
]
}
logs:StartQuery is scoped to the log groups you query. GetQueryResults and StopQuery take only a query ID, so this policy grants them on "*". The IAM policy generator for TypeScript code lists the actions from the module, and the guide to find the IAM actions your AWS SDK for JavaScript code needs explains how commands map to actions.
Troubleshooting and common mistakes
- Zero rows, or an error about the time range. Times passed in milliseconds, a range before the logs existed, or the client in a different Region from the log group.
MalformedQueryException. The query doesn’t parse. Test it in the console’s Logs Insights editor first, then copy it into code with a template literal so the pipes and newlines survive.LimitExceededExceptionfromStartQuery. A limit was reached; when you start many queries at once, the concurrent-query quota is the likely one. The helper retries with backoff; if it persists, run fewer queries in parallel.ResourceNotFoundException. A log group name is wrong or in another Region.- Only 1,000 rows. That’s the helper’s default
limit. Raise it, up to 10,000 per page, and let thenextTokenloop fetch the rest. - Status
Timeout. The query hit 60 minutes. Split the time range into smaller queries and combine the results.
Testing the helper doesn’t need AWS: mock StartQueryCommand and a Running then Complete response for GetQueryResultsCommand, as the guide to mock AWS SDK v3 clients in unit tests shows.
Limits of this approach
- It isn’t real time. For a live feed, use Live Tail in the console; for alerts, use metric filters and alarms.
- Polling holds a process open. For long queries on a schedule, CloudWatch scheduled queries run them for you and share the same concurrency quota.
- The 100,000-row pagination applies to Logs Insights QL only, not the OpenSearch PPL or SQL languages.
For a one-off question you don’t want to write code for, ChatWithCloud can run the query for you; see how to ask AI about Lambda errors in your AWS account. It generates SDK v2 code under the hood, so it answers questions rather than producing a v3 helper like this one.
Frequently asked questions
Why does my CloudWatch Logs Insights query return no results from the SDK?
Most often, startTime and endTime were passed in milliseconds. The API expects epoch seconds. Also check the client’s Region.
How long does a Logs Insights query take?
It depends on how much data it scans. Small ranges finish in seconds; any query times out after 60 minutes. Poll with backoff rather than waiting a fixed time.
How many log groups can one Logs Insights query search?
Up to 50 when you list them with logGroupNames or logGroupIdentifiers. A SOURCE command in the query can select log groups by prefix instead.
Do I pay for a query I stop?
You pay for data scanned. Stopping a query that’s still running prevents further scanning, which is why the helper calls StopQuery when it gives up.
Related guides
Ask your AWS account in plain English
Your first 15 runs are free, with no OpenAI key needed.
npx chatwithcloud