Find Stuck SQS Queues by Approximate Age of Oldest Message

A long line of cars with red tail lights stopped on a highway at night

Photo by Musa Haef on Unsplash

The SQS approximate age of oldest message (ApproximateAgeOfOldestMessage in CloudWatch) is how long the oldest unprocessed message has waited, in seconds. Compare its maximum over the last 24 hours with the queue’s MessageRetentionPeriod: a value close to retention means messages are about to be deleted unprocessed. Read queue depth with GetQueueAttributes and consumer activity from NumberOfMessagesReceived.

A queue doesn’t fail loudly. Producers keep sending, SendMessage keeps succeeding, and the backlog grows quietly until messages reach the retention period and SQS deletes them. By the time someone notices missing orders or emails, the evidence has expired.

This example is for engineers who own SQS-based pipelines and want a sweep of every queue in a Region. The script combines the current queue depth with the SQS approximate age of oldest message and receive counts from CloudWatch, then flags messages close to expiry, backlogs no consumer is reading, old messages and dead-letter queues that hold messages. It only reports. If your queues have no dead-letter queues yet, fix that first with the example to find SQS queues without a dead-letter queue.

What does the SQS approximate age of oldest message measure?

It’s the age of the oldest message still waiting in the queue, reported only while the queue holds at least one message. In the terms of Google’s SRE book, which lists latency and saturation among the four golden signals of monitoring, it’s the queue’s latency: depth tells you how much work is waiting, age tells you how long. AWS documents several quirks worth knowing before you alarm on it:

  • Poison messages drop out. On a standard queue, a message received three or more times without being deleted moves to the back of the queue, and the metric then reports the next message instead. A message that fails forever doesn’t keep the age high.
  • FIFO queues don’t reorder. A failing message blocks its message group until it’s deleted, expires or moves to the dead-letter queue.
  • Dead-letter queues reset the clock, retention doesn’t. On a standard queue’s DLQ, the metric shows the time since the message moved, but expiry still counts from the original send. A message that waited 1 day in the source queue and then moved to a DLQ with 4-day retention is deleted after 3 more days, while the DLQ’s metric says 3 days. FIFO queues reset the enqueue time on the move.

That last point is why AWS recommends a DLQ retention period longer than the source queue’s. Retention itself ranges from 60 seconds to 1,209,600 seconds (14 days); the default is 345,600 seconds (4 days).

What does the script do?

  1. Lists queues and reads attributespaginateListQueues (1,000 per page), then GetQueueAttributes for ApproximateNumberOfMessages, ApproximateNumberOfMessagesNotVisible, MessageRetentionPeriod and RedrivePolicy. Queues named in another queue’s redrive policy are marked as DLQs.
  2. Reads three metrics per queuepaginateGetMetricData fetches the hourly Maximum of ApproximateAgeOfOldestMessage and the Sum of NumberOfMessagesReceived and NumberOfMessagesDeleted over --hours (default 24), for up to 160 queues per call.
  3. ClassifiesEXPIRING when the oldest message reached --expire-ratio (default 80%) of retention; OLD MESSAGES above --max-age (default 1 hour); NO CONSUMER when messages wait but nothing was received; RECEIVED BUT NEVER DELETED when consumers fail every message.
  4. Checks DLQsFor a DLQ holding messages, ListMessageMoveTasks shows the last redrive task. It also flags DLQs whose retention isn’t longer than their source queue’s.

Prerequisites

  • Node.js 18 or later, tsx, @aws-sdk/client-sqs and @aws-sdk/client-cloudwatch.
  • An idea of normal latency per queue. A nightly batch queue that holds messages for 6 hours is fine; a payment queue that does is not. Tune --max-age or run it per queue prefix.
  • The guide to AWS SDK v3 paginators explains how paginateGetMetricData follows NextToken when a long window exceeds 100,800 data points.

Which IAM permissions does it need?

sqs-stuck-messages-policy.json

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "ListQueuesAndMetrics",
      "Effect": "Allow",
      "Action": ["sqs:ListQueues", "cloudwatch:GetMetricData"],
      "Resource": "*"
    },
    {
      "Sid": "ReadQueues",
      "Effect": "Allow",
      "Action": ["sqs:GetQueueAttributes", "sqs:ListMessageMoveTasks"],
      "Resource": "arn:aws:sqs:*:111122223333:*"
    }
  ]
}

All read-only; the script never receives or deletes messages. GetMetricData costs $0.01 per 1,000 metrics requested in US East (N. Virginia) as of September 2026 (AWS Price List API), so a sweep of 300 queues requests 900 metrics and costs about one cent.

The script to check the SQS approximate age of oldest message

find-sqs-queues-with-stuck-messages.ts

// find-sqs-queues-with-stuck-messages.ts
// For every SQS queue in a Region: current depth from GetQueueAttributes, plus the maximum
// ApproximateAgeOfOldestMessage and the NumberOfMessagesReceived/Deleted sums from CloudWatch over the last
// N hours. Flags messages close to their retention period, backlogs nobody is reading, old messages and
// dead-letter queues that hold messages. Report only.
// Usage: npx tsx find-sqs-queues-with-stuck-messages.ts [--region us-east-1] [--hours 24] [--max-age 3600]
import {
  GetQueueAttributesCommand,
  ListMessageMoveTasksCommand,
  SQSClient,
  paginateListQueues,
} from "@aws-sdk/client-sqs";
import { CloudWatchClient, paginateGetMetricData, type MetricDataQuery } from "@aws-sdk/client-cloudwatch";

const args = process.argv.slice(2);
const flag = (name: string): string | undefined => {
  const i = args.indexOf(name);
  return i >= 0 ? args[i + 1] : undefined;
};
const region = flag("--region") ?? process.env.AWS_REGION ?? "us-east-1";
const hours = Number(flag("--hours") ?? "24");
const maxAge = Number(flag("--max-age") ?? "3600"); // seconds before a message counts as old
const expireRatio = Number(flag("--expire-ratio") ?? "0.8"); // share of retention used by the oldest message
if (!(hours >= 1 && hours <= 336) || !(maxAge > 0) || !(expireRatio > 0 && expireRatio <= 1)) {
  console.error("--hours must be 1-336, --max-age above 0 and --expire-ratio between 0 and 1");
  process.exit(1);
}

const sqs = new SQSClient({ region });
const cloudwatch = new CloudWatchClient({ region });
const errText = (err: unknown): string => (err instanceof Error ? `${err.name}: ${err.message}` : String(err));

interface Queue {
  url: string;
  name: string;
  arn: string;
  visible: number;
  inFlight: number;
  retention: number; // seconds
  dlqArn?: string; // this queue's own dead-letter queue, if any
}
interface Stats {
  oldest: number; // max ApproximateAgeOfOldestMessage, seconds
  received: number;
  deleted: number;
}

async function describe(url: string): Promise<Queue> {
  const { Attributes: a = {} } = await sqs.send(
    new GetQueueAttributesCommand({
      QueueUrl: url,
      AttributeNames: [
        "QueueArn",
        "ApproximateNumberOfMessages",
        "ApproximateNumberOfMessagesNotVisible",
        "MessageRetentionPeriod",
        "RedrivePolicy",
      ],
    }),
  );
  let dlqArn: string | undefined;
  if (a.RedrivePolicy) dlqArn = (JSON.parse(a.RedrivePolicy) as { deadLetterTargetArn?: string }).deadLetterTargetArn;
  return {
    url,
    name: url.split("/").pop() ?? url,
    arn: a.QueueArn ?? "",
    visible: Number(a.ApproximateNumberOfMessages ?? 0),
    inFlight: Number(a.ApproximateNumberOfMessagesNotVisible ?? 0),
    retention: Number(a.MessageRetentionPeriod ?? 345600),
    dlqArn,
  };
}

async function metrics(queues: Queue[]): Promise<Map<string, Stats>> {
  const end = new Date();
  const start = new Date(end.getTime() - hours * 3_600_000);
  const q = (queue: string, name: string, stat: string, id: string): MetricDataQuery => ({
    Id: id,
    MetricStat: {
      Metric: { Namespace: "AWS/SQS", MetricName: name, Dimensions: [{ Name: "QueueName", Value: queue }] },
      Period: 3600,
      Stat: stat,
    },
  });
  const stats = new Map<string, Stats>();
  for (let i = 0; i < queues.length; i += 160) {
    const batch = queues.slice(i, i + 160); // 3 queries per queue, 500 per GetMetricData call
    const queries = batch.flatMap((x, j) => [
      q(x.name, "ApproximateAgeOfOldestMessage", "Maximum", `a${j}`),
      q(x.name, "NumberOfMessagesReceived", "Sum", `r${j}`),
      q(x.name, "NumberOfMessagesDeleted", "Sum", `d${j}`),
    ]);
    const values = new Map<string, number[]>();
    for await (const page of paginateGetMetricData({ client: cloudwatch }, { MetricDataQueries: queries, StartTime: start, EndTime: end })) {
      for (const r of page.MetricDataResults ?? []) {
        if (r.Id) values.set(r.Id, [...(values.get(r.Id) ?? []), ...(r.Values ?? [])]);
      }
    }
    const max = (id: string): number => Math.max(0, ...(values.get(id) ?? []));
    const sum = (id: string): number => (values.get(id) ?? []).reduce((a, b) => a + b, 0);
    batch.forEach((x, j) => stats.set(x.name, { oldest: max(`a${j}`), received: sum(`r${j}`), deleted: sum(`d${j}`) }));
  }
  return stats;
}

// Latest redrive task on a DLQ, if one was ever started
async function lastMove(dlqArn: string): Promise<string> {
  const out = await sqs.send(new ListMessageMoveTasksCommand({ SourceArn: dlqArn, MaxResults: 1 }));
  const t = out.Results?.[0];
  return t ? `${t.Status}${t.ApproximateNumberOfMessagesMoved !== undefined ? ` (${t.ApproximateNumberOfMessagesMoved} moved)` : ""}` : "none";
}

const human = (s: number): string => (s >= 86_400 ? `${(s / 86_400).toFixed(1)}d` : s >= 3600 ? `${(s / 3600).toFixed(1)}h` : `${Math.round(s)}s`);

async function main(): Promise<void> {
  const urls: string[] = [];
  for await (const page of paginateListQueues({ client: sqs }, { MaxResults: 1000 })) urls.push(...(page.QueueUrls ?? []));
  const queues: Queue[] = [];
  for (const url of urls) {
    try {
      queues.push(await describe(url));
    } catch (err) {
      console.error(`${url}: ${errText(err)}`); // e.g. deleted between ListQueues and GetQueueAttributes
    }
  }
  const byArn = new Map(queues.map((x) => [x.arn, x]));
  const dlqs = new Set(queues.map((x) => x.dlqArn).filter((a): a is string => Boolean(a)));
  const stats = await metrics(queues);

  const rows: Record<string, string | number>[] = [];
  for (const x of queues) {
    const s = stats.get(x.name) ?? { oldest: 0, received: 0, deleted: 0 };
    const isDlq = dlqs.has(x.arn);
    const findings: string[] = [];
    if (s.oldest >= x.retention * expireRatio) findings.push(`EXPIRING: oldest at ${Math.round((s.oldest / x.retention) * 100)}% of retention`);
    else if (!isDlq && s.oldest >= maxAge) findings.push("OLD MESSAGES");
    if (!isDlq && x.visible > 0 && s.received === 0) findings.push(`NO CONSUMER in ${hours}h`);
    if (!isDlq && s.received > 0 && s.deleted === 0) findings.push("RECEIVED BUT NEVER DELETED");
    if (isDlq && x.visible > 0) {
      let move = "?";
      try {
        move = await lastMove(x.arn);
      } catch (err) {
        move = errText(err);
      }
      findings.push(`DLQ HOLDS MESSAGES (last redrive: ${move})`);
    }
    if (x.dlqArn) {
      const dlq = byArn.get(x.dlqArn);
      if (dlq && dlq.retention <= x.retention) findings.push("DLQ retention not longer than source");
    }
    rows.push({
      Queue: x.name,
      Role: isDlq ? "DLQ" : "queue",
      Visible: x.visible,
      InFlight: x.inFlight,
      OldestMax: human(s.oldest),
      Retention: human(x.retention),
      Received: s.received,
      Findings: findings.join("; ") || "ok",
    });
  }
  rows.sort((a, b) => Number(b.Findings !== "ok") - Number(a.Findings !== "ok") || Number(b.Visible) - Number(a.Visible));
  console.table(rows);
  const stuck = rows.filter((r) => r.Findings !== "ok").length;
  console.log(`${queues.length} queues in ${region} over ${hours}h: ${stuck} to look at. Report only: nothing was changed.`);
}

main().catch((err) => {
  console.error(errText(err));
  process.exit(1);
});

The script uses the worst hour in the window rather than the current value, so a backlog that built up overnight and drained by morning still shows. GetQueueAttributes counts are approximate and can take a minute to settle after producers stop.

How do you run it?

Terminal

npm install @aws-sdk/client-sqs @aws-sdk/client-cloudwatch
npm install --save-dev tsx typescript @types/node

AWS_PROFILE=readonly npx tsx find-sqs-queues-with-stuck-messages.ts --region us-east-1 --hours 24 --max-age 3600

Sample output

Output

┌─────────┬──────────────┬─────────┬─────────┬──────────┬───────────┬───────────┬──────────┬──────────────────────────────────────────────────────────────────────────────────────────────────┐
│ (index) │ Queue        │ Role    │ Visible │ InFlight │ OldestMax │ Retention │ Received │ Findings                                                                                         │
├─────────┼──────────────┼─────────┼─────────┼──────────┼───────────┼───────────┼──────────┼──────────────────────────────────────────────────────────────────────────────────────────────────┤
│ 0       │ 'orders'     │ 'queue' │ 5200    │ 40       │ '2.5h'    │ '4.0d'    │ 300      │ 'OLD MESSAGES; DLQ retention not longer than source'                                             │
│ 1       │ 'emails'     │ 'queue' │ 310     │ 0        │ '3.4d'    │ '4.0d'    │ 0        │ 'EXPIRING: oldest at 84% of retention; NO CONSUMER in 24h'                                       │
│ 2       │ 'orders-dlq' │ 'DLQ'   │ 37      │ 0        │ '3.6d'    │ '4.0d'    │ 0        │ 'EXPIRING: oldest at 90% of retention; DLQ HOLDS MESSAGES (last redrive: COMPLETED (120 moved))' │
│ 3       │ 'reports'    │ 'queue' │ 0       │ 0        │ '0s'      │ '14.0d'   │ 0        │ 'ok'                                                                                             │
└─────────┴──────────────┴─────────┴─────────┴──────────┴───────────┴───────────┴──────────┴──────────────────────────────────────────────────────────────────────────────────────────────────┘
4 queues in us-east-1 over 24h: 3 to look at. Report only: nothing was changed.

Queue names and numbers are illustrative; the output came from a run against mocked SDK clients. emails is the urgent one: 310 messages, nothing received in 24 hours, and the oldest message has used 84% of its 4-day retention. Its consumer is down, and in less than a day those messages are gone. orders-dlq holds 37 failed orders, and because it’s a standard queue with the same 4-day retention as orders, they expire sooner than the 3.6 days shown suggests. orders itself is busy but running 2.5 hours behind.

What should you do with a stuck queue?

  • No consumer: check the consumer first. For a Lambda trigger, look for a disabled event source mapping or a function that errors on every batch; the guide to investigate Lambda errors with CloudWatch walks through it.
  • Received but never deleted: consumers fail and messages return after the visibility timeout. If a Lambda function fails a whole batch because of one message, the example to find Lambda SQS triggers without ReportBatchItemFailures fixes that; if it times out, the one to find Lambda functions with a timeout too high for their runtime also checks the visibility timeout.
  • Expiring: buy time by raising MessageRetentionPeriod with SetQueueAttributes, up to 14 days, while you fix the consumer. Never lower it on a queue with a backlog: AWS warns that existing messages older than the new period can be deleted.
  • DLQ holds messages: fix the cause, then redrive with StartMessageMoveTask, which moves messages from a DLQ back to their source queue (or another queue) at up to 500 messages per second. It only works for DLQs of SQS queues, and one task can run per queue at a time.

To receive and delete a few messages by hand while you debug, the guide to send and receive SQS messages with AWS SDK v3 has the code. Then add an alarm on the metric so the next backlog pages someone; alarms that never fire are covered by the example to find CloudWatch alarms stuck in INSUFFICIENT_DATA.

Troubleshooting

  • OldestMax is 0 for a queue with messages. Check the Region first. SQS emits these metrics only while a queue is active, and a queue whose only messages are poison messages can also report a low age.
  • QueueDoesNotExist errors. A queue was deleted between ListQueues and GetQueueAttributes. The script logs it and continues.
  • Delayed messages. Messages sent with DelaySeconds aren’t counted in ApproximateNumberOfMessages; the ApproximateNumberOfMessagesDelayed attribute holds them.
  • Slow on large accounts. GetQueueAttributes runs once per queue. Filter with QueueNamePrefix in the paginateListQueues call if you only care about one system.

Ask ChatWithCloud instead

To check one pipeline, ask ChatWithCloud “Which SQS queues in us-east-1 have messages older than an hour?” It writes AWS SDK for JavaScript v2 code, runs it on your machine with your profile and answers from the metrics; the guide to troubleshoot AWS infrastructure with an AI CLI shows how to follow up on a finding. It can be wrong and doesn’t ask before making changes, so keep it on a read-only profile, as the ChatWithCloud security page recommends.

Frequently asked questions

What is a good alarm threshold for ApproximateAgeOfOldestMessage?

Base it on the queue’s job. For interactive work, minutes; for batch queues, a few hours. Always alarm well before the retention period, for example at half of it, so there’s time to react.

Why does ApproximateAgeOfOldestMessage drop while messages are still failing?

On standard queues, a message received three or more times without deletion moves to the back of the queue, and the metric reports the next message’s age instead.

What happens when SQS messages reach the retention period?

SQS deletes them. The default retention is 4 days and the maximum is 14 days.

How do I move messages out of a dead-letter queue?

Use DLQ redrive: StartMessageMoveTask in the API, or the dead-letter queue redrive option in the SQS console. It moves messages back to the source queue or to a queue you choose.

Related guides

Ask your AWS account in plain English

Your first 15 runs are free, with no OpenAI key needed.

npx chatwithcloud