Find Unused Lambda Provisioned Concurrency

Close-up of cooling fans and status lights on the front of a server rack

Photo by Daniel Miksha on Unsplash

Lambda provisioned concurrency cost is charged per GB-second for every pre-initialized environment, from the moment you enable it until you remove it, whether or not requests arrive. Find waste with ListProvisionedConcurrencyConfigs per function, then read ProvisionedConcurrencyInvocations and the Maximum of ProvisionedConcurrencyUtilization in CloudWatch. No provisioned invocations for two weeks means the config is idle.

Provisioned concurrency keeps execution environments initialized so a function answers without a cold start. It’s set on an alias or a version, never on $LATEST, and it’s easy to forget: an alias that stopped receiving traffic after a routing change, a version left behind by a deployment, a launch-day setting nobody lowered afterwards.

This example is for serverless and platform engineers who want to see what that costs and fix it. The script lists every provisioned concurrency config in the Regions you pass, shows peak utilization and the monthly charge, suggests a smaller setting for oversized ones, and with --apply deletes configs that served no provisioned invocations in the window.

How is Lambda provisioned concurrency billed?

You pay for the memory of each provisioned environment for as long as the config exists, plus requests and a lower duration rate for invocations that run on it. As of September 2026, the AWS Price List (published 19 September 2026) shows these rates in US East (N. Virginia):

Charge x86_64 arm64
Provisioned concurrency, per GB-second enabled $0.0000041667 $0.0000033334
Duration on provisioned concurrency, per GB-second $0.0000097222 $0.0000077778
Duration on demand (first tier), per GB-second $0.0000166667 $0.0000133334
Requests, per million $0.20 $0.20

Worked example: 10 environments of a 1,024 MB x86 function cost 10 × 1 GB × 2,628,000 seconds (730 hours) × $0.0000041667 = $109.50 a month before a single request. The same 10 on a 512 MB function cost $54.75.

The duration discount means busy provisioned concurrency can pay for itself. For both architectures, the enabled charge roughly equals the duration saving at 60% utilization: $0.0000041667 ÷ ($0.0000166667 − $0.0000097222) = 0.6. Below that, you pay a premium for lower latency, which may still be the right call for a checkout API; the script shows where you pay it.

The Lambda provisioned concurrency documentation adds two costs that don’t show on this line: Lambda bills initialization even for environments that never serve a request, and provisioned concurrency counts against the account’s concurrency pool, leaving less for other functions.

Which metrics show unused provisioned concurrency?

  • ProvisionedConcurrencyInvocations (view with Sum): invocations that ran on provisioned environments. Zero over two weeks means nothing uses the alias or version.
  • ProvisionedConcurrencyUtilization (view with Maximum): the share of provisioned environments busy at once. The Lambda docs call consistently low values a sign of over-allocation.
  • ProvisionedConcurrencySpilloverInvocations (Sum): invocations that overflowed to standard concurrency. A high value means you have too little, not too much.

Lambda emits these only while the function receives requests. An idle config therefore has no data points at all, which is why the script treats “no series in ListMetrics” as zero invocations. ListMetrics only sees metrics that reported in the past two weeks, so --days is capped at 14.

For oversized configs the script suggests peak concurrent use plus the 10% buffer the Lambda docs recommend: ceil(peak utilization × allocated × 1.1).

What does the script do?

  1. Lists functionspaginateListFunctions for name, memory size and architecture.
  2. Lists configspaginateListProvisionedConcurrencyConfigs per function: qualifier, allocated environments, status and last change.
  3. Checks auto scalingpaginateDescribeScalableTargets for the lambda namespace, so configs managed by Application Auto Scaling are never deleted behind its back.
  4. Reads usageListMetrics finds the alias or version series, then GetMetricData reads peak utilization and invocation sums.
  5. Deletes only with --applyDeleteProvisionedConcurrencyConfig for idle, READY configs not changed within the window and not managed by auto scaling.

Prerequisites

  • Node.js 18 or later with tsx, plus @aws-sdk/client-lambda, @aws-sdk/client-application-auto-scaling and @aws-sdk/client-cloudwatch.
  • An AWS profile set up as in the guide to AWS SDK v3 credential providers and assumed roles. Use a read-only one unless you pass --apply.

Which IAM permissions does it need?

unused-provisioned-concurrency-policy.json

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "ReadProvisionedConcurrency",
      "Effect": "Allow",
      "Action": [
        "lambda:ListFunctions",
        "lambda:ListProvisionedConcurrencyConfigs",
        "application-autoscaling:DescribeScalableTargets",
        "cloudwatch:ListMetrics",
        "cloudwatch:GetMetricData"
      ],
      "Resource": "*"
    },
    {
      "Sid": "DeleteIdleConfigsWithApply",
      "Effect": "Allow",
      "Action": "lambda:DeleteProvisionedConcurrencyConfig",
      "Resource": [
        "arn:aws:lambda:*:123456789012:function:*",
        "arn:aws:lambda:*:123456789012:function:*:*"
      ]
    }
  ]
}

Replace the account ID, and drop the second statement for a report-only role. The IAM policy generator that reads TypeScript SDK code lists the actions if you extend the script.

The script to find unused Lambda provisioned concurrency

find-unused-lambda-provisioned-concurrency.ts

// find-unused-lambda-provisioned-concurrency.ts
// Lists every Lambda provisioned concurrency config (per alias or version) with its peak utilization over
// the last N days and its idle monthly charge, and suggests a smaller setting for over-provisioned ones.
// Report only by default. --apply deletes configs with no provisioned invocations in the window, unless
// Application Auto Scaling manages them.
// Usage: npx tsx find-unused-lambda-provisioned-concurrency.ts [--regions us-east-1] [--days 14] [--low 0.3] [--csv pc.csv] [--apply]
import { writeFileSync } from "node:fs";
import {
  DeleteProvisionedConcurrencyConfigCommand,
  LambdaClient,
  paginateListFunctions,
  paginateListProvisionedConcurrencyConfigs,
  type FunctionConfiguration,
} from "@aws-sdk/client-lambda";
import { ApplicationAutoScalingClient, paginateDescribeScalableTargets } from "@aws-sdk/client-application-auto-scaling";
import { CloudWatchClient, paginateGetMetricData, paginateListMetrics, type Dimension } from "@aws-sdk/client-cloudwatch";

const args = process.argv.slice(2);
const flag = (name: string): string | undefined => {
  const i = args.indexOf(name);
  return i >= 0 ? args[i + 1] : undefined;
};
const regions = (flag("--regions") ?? process.env.AWS_REGION ?? "us-east-1").split(",").map((r) => r.trim()).filter(Boolean);
const days = Number(flag("--days") ?? 14);
const low = Number(flag("--low") ?? 0.3); // peak utilization below this is over-provisioned
const csvPath = flag("--csv");
const apply = args.includes("--apply");

// us-east-1 USD per GB-second of provisioned concurrency while it's enabled. AWS Price List (AWSLambda), published 19 September 2026.
const PC_GB_SECOND = { x86_64: 0.0000041667, arm64: 0.0000033334 };
const SECONDS_PER_MONTH = 730 * 3600;

interface Row {
  Region: string;
  Function: string;
  Qualifier: string;
  MemoryMB: number;
  Arch: string;
  Allocated: number;
  Status: string;
  PeakUtil: string;
  PcInvocations: number;
  Spillover: number;
  AutoScaling: string;
  PerMonth: string;
  Verdict: string;
  Action: string;
}

interface Usage {
  peak?: number;
  invocations: number;
  spillover: number;
}

/** Finds the alias/version series via ListMetrics, then reads peak utilization and invocation sums. */
async function usage(cw: CloudWatchClient, fn: string, qualifier: string): Promise<Usage> {
  let dims: Dimension[] | undefined;
  for await (const page of paginateListMetrics({ client: cw }, {
    Namespace: "AWS/Lambda",
    MetricName: "ProvisionedConcurrencyInvocations",
    Dimensions: [{ Name: "FunctionName", Value: fn }],
  })) {
    for (const m of page.Metrics ?? []) {
      const resource = m.Dimensions?.find((d) => d.Name === "Resource")?.Value;
      if (resource?.endsWith(`:${qualifier}`) && m.Dimensions?.length === 2) dims = m.Dimensions;
    }
  }
  if (!dims) return { invocations: 0, spillover: 0 }; // no data points in the last two weeks
  const end = new Date();
  const start = new Date(end.getTime() - days * 86_400_000);
  const stat = (id: string, metric: string, s: string) => ({
    Id: id,
    MetricStat: { Metric: { Namespace: "AWS/Lambda", MetricName: metric, Dimensions: dims }, Period: 3600, Stat: s },
  });
  const values: Record<string, number[]> = { util: [], inv: [], spill: [] };
  for await (const page of paginateGetMetricData({ client: cw }, {
    StartTime: start,
    EndTime: end,
    MetricDataQueries: [
      stat("util", "ProvisionedConcurrencyUtilization", "Maximum"),
      stat("inv", "ProvisionedConcurrencyInvocations", "Sum"),
      stat("spill", "ProvisionedConcurrencySpilloverInvocations", "Sum"),
    ],
  })) {
    for (const r of page.MetricDataResults ?? []) values[r.Id ?? ""]?.push(...(r.Values ?? []));
  }
  const sum = (xs: number[] | undefined) => (xs ?? []).reduce((a, b) => a + b, 0);
  const util = values.util ?? [];
  return { peak: util.length ? Math.max(...util) : undefined, invocations: sum(values.inv), spillover: sum(values.spill) };
}

async function scalableTargets(aas: ApplicationAutoScalingClient): Promise<Set<string>> {
  const ids = new Set<string>();
  for await (const page of paginateDescribeScalableTargets({ client: aas }, { ServiceNamespace: "lambda" })) {
    for (const t of page.ScalableTargets ?? []) if (t.ResourceId) ids.add(t.ResourceId); // "function:<name>:<qualifier>"
  }
  return ids;
}

async function scanRegion(region: string): Promise<Row[]> {
  const lambda = new LambdaClient({ region });
  const cw = new CloudWatchClient({ region });
  const managed = await scalableTargets(new ApplicationAutoScalingClient({ region }));
  const functions: FunctionConfiguration[] = [];
  for await (const page of paginateListFunctions({ client: lambda }, {})) functions.push(...(page.Functions ?? []));
  const rows: Row[] = [];
  for (const fn of functions) {
    const name = fn.FunctionName ?? "";
    for await (const page of paginateListProvisionedConcurrencyConfigs({ client: lambda }, { FunctionName: name })) {
      for (const pc of page.ProvisionedConcurrencyConfigs ?? []) {
        const qualifier = pc.FunctionArn?.split(":").pop() ?? "";
        const allocated = pc.AllocatedProvisionedConcurrentExecutions ?? 0;
        const memoryMb = fn.MemorySize ?? 128;
        const arch = fn.Architectures?.[0] === "arm64" ? "arm64" : "x86_64";
        const perMonth = allocated * (memoryMb / 1024) * SECONDS_PER_MONTH * PC_GB_SECOND[arch];
        const u = await usage(cw, name, qualifier);
        const aasManaged = managed.has(`function:${name}:${qualifier}`);
        const modifiedDaysAgo = pc.LastModified ? (Date.now() - Date.parse(pc.LastModified)) / 86_400_000 : Infinity;
        let verdict = "in use";
        if (pc.Status !== "READY") verdict = `status ${pc.Status ?? "unknown"}`;
        else if (modifiedDaysAgo < days) verdict = "changed recently: re-check later";
        else if (u.invocations === 0) verdict = "IDLE: no provisioned invocations";
        else if (u.peak !== undefined && u.peak < low) {
          const suggested = Math.max(1, Math.ceil(u.peak * allocated * 1.1)); // peak concurrent use + 10% buffer
          verdict = `LOW: try ${suggested}`;
        }
        let action = verdict.startsWith("IDLE") ? (aasManaged ? "managed by auto scaling" : "would delete (--apply)") : "-";
        if (apply && verdict.startsWith("IDLE") && !aasManaged) {
          try {
            await lambda.send(new DeleteProvisionedConcurrencyConfigCommand({ FunctionName: name, Qualifier: qualifier }));
            action = "deleted";
          } catch (err) {
            action = `failed: ${err instanceof Error ? err.name : String(err)}`;
          }
        }
        rows.push({
          Region: region,
          Function: name,
          Qualifier: qualifier,
          MemoryMB: memoryMb,
          Arch: arch,
          Allocated: allocated,
          Status: pc.Status ?? "",
          PeakUtil: u.peak === undefined ? "-" : `${Math.round(u.peak * 100)}%`,
          PcInvocations: Math.round(u.invocations),
          Spillover: Math.round(u.spillover),
          AutoScaling: aasManaged ? "yes" : "no",
          PerMonth: `$${(Math.round(perMonth * 100) / 100).toFixed(2)}`,
          Verdict: verdict,
          Action: action,
        });
      }
    }
  }
  return rows;
}

function toCsv(rows: Row[]): string {
  const cols = Object.keys(rows[0] ?? {}) as (keyof Row)[];
  const cell = (v: string | number) => `"${String(v).replace(/"/g, '""')}"`;
  return [cols.join(","), ...rows.map((r) => cols.map((c) => cell(r[c])).join(","))].join("\n") + "\n";
}

async function main(): Promise<void> {
  if (!Number.isInteger(days) || days < 1 || days > 14) throw new Error("--days must be 1 to 14: ListMetrics only sees metrics with data in the last two weeks");
  if (!(low > 0 && low <= 1)) throw new Error("--low must be between 0 and 1");
  const rows: Row[] = [];
  for (const region of regions) {
    try {
      rows.push(...(await scanRegion(region)));
    } catch (err) {
      console.error(`${region}: ${err instanceof Error ? `${err.name}: ${err.message}` : String(err)}`);
    }
  }
  if (rows.length === 0) {
    console.log(`No provisioned concurrency configured in ${regions.join(", ")}.`);
    return;
  }
  console.table(rows);
  const idle = rows.filter((r) => r.Verdict.startsWith("IDLE"));
  const monthly = idle.reduce((s, r) => s + Number(r.PerMonth.replace("$", "")), 0);
  console.log(`${idle.length} of ${rows.length} configs had no provisioned invocations in ${days} days: $${monthly.toFixed(2)} a month of idle charge.`);
  if (csvPath) {
    writeFileSync(csvPath, toCsv(rows));
    console.log(`Wrote ${rows.length} rows to ${csvPath}`);
  }
}

main().catch((err) => {
  console.error(err);
  process.exit(1);
});

How do you run it?

Terminal

npm install @aws-sdk/client-lambda @aws-sdk/client-application-auto-scaling @aws-sdk/client-cloudwatch
npm install --save-dev tsx typescript @types/node

# Report for two Regions, flag anything under 30% peak utilization, write a CSV
AWS_PROFILE=readonly npx tsx find-unused-lambda-provisioned-concurrency.ts --regions us-east-1,eu-west-1 --csv pc.csv

# Delete idle configs that auto scaling doesn't manage
AWS_PROFILE=lambda-admin npx tsx find-unused-lambda-provisioned-concurrency.ts --regions us-east-1 --apply

Sample output

Output

┌─────────┬─────────────┬──────────────────┬───────────┬──────────┬──────────┬───────────┬─────────┬──────────┬───────────────┬───────────┬─────────────┬───────────┬────────────────────────────────────┬───────────────────────────┐
│ (index) │ Region      │ Function         │ Qualifier │ MemoryMB │ Arch     │ Allocated │ Status  │ PeakUtil │ PcInvocations │ Spillover │ AutoScaling │ PerMonth  │ Verdict                            │ Action                    │
├─────────┼─────────────┼──────────────────┼───────────┼──────────┼──────────┼───────────┼─────────┼──────────┼───────────────┼───────────┼─────────────┼───────────┼────────────────────────────────────┼───────────────────────────┤
│ 0       │ 'us-east-1' │ 'checkout-api'   │ 'live'    │ 1024     │ 'x86_64' │ 20        │ 'READY' │ '85%'    │ 1800000       │ 3100      │ 'no'        │ '$219.00' │ 'in use'                           │ '-'                       │
│ 1       │ 'us-east-1' │ 'search-api'     │ 'prod'    │ 2048     │ 'arm64'  │ 50        │ 'READY' │ '20%'    │ 240000        │ 0         │ 'no'        │ '$876.02' │ 'LOW: try 11'                      │ '-'                       │
│ 2       │ 'us-east-1' │ 'legacy-webhook' │ '7'       │ 512      │ 'x86_64' │ 10        │ 'READY' │ '-'      │ 0             │ 0         │ 'no'        │ '$54.75'  │ 'IDLE: no provisioned invocations' │ 'would delete (--apply)'  │
│ 3       │ 'us-east-1' │ 'report-worker'  │ 'prod'    │ 1024     │ 'x86_64' │ 5         │ 'READY' │ '-'      │ 0             │ 0         │ 'yes'       │ '$54.75'  │ 'IDLE: no provisioned invocations' │ 'managed by auto scaling' │
└─────────┴─────────────┴──────────────────┴───────────┴──────────┴──────────┴───────────┴─────────┴──────────┴───────────────┴───────────┴─────────────┴───────────┴────────────────────────────────────┴───────────────────────────┘
2 of 4 configs had no provisioned invocations in 14 days: $109.50 a month of idle charge.

This is a run against mocked AWS responses, so names and numbers are illustrative. checkout-api:live peaks at 85% with some spillover, so it’s sized about right. search-api:prod never passes 20% of its 50 environments; 11 would cover the peak with a 10% buffer and cut its $876.02 bill by about 78%. legacy-webhook has 10 environments on version 7 that nothing calls. report-worker is also idle but belongs to an auto scaling target, so the script leaves it for you to fix at the target.

Delete it, shrink it or schedule it?

Idle configs on old versions are the easy win: delete the config, then see whether the version itself can go with deleting old and unused Lambda function versions. For traffic that follows office hours, a schedule keeps the latency benefit only when it matters. With 20 environments at 1 GB from 08:00 to 20:00 on 22 weekdays and 1 environment otherwise, the enabled charge falls from $219.00 to about $86.19 a month.

schedule-provisioned-concurrency.ts

// schedule-provisioned-concurrency.ts
// Keeps 20 provisioned environments on the "live" alias during business hours and 1 outside them.
// The alias must already have a provisioned concurrency config before Application Auto Scaling manages it.
// Usage: npx tsx schedule-provisioned-concurrency.ts checkout-api live
import {
  ApplicationAutoScalingClient,
  PutScheduledActionCommand,
  RegisterScalableTargetCommand,
} from "@aws-sdk/client-application-auto-scaling";

const [fn = "checkout-api", alias = "live"] = process.argv.slice(2);
const aas = new ApplicationAutoScalingClient({ region: process.env.AWS_REGION ?? "us-east-1" });
const target = {
  ServiceNamespace: "lambda" as const,
  ResourceId: `function:${fn}:${alias}`,
  ScalableDimension: "lambda:function:ProvisionedConcurrency" as const,
};

async function main(): Promise<void> {
  await aas.send(new RegisterScalableTargetCommand({ ...target, MinCapacity: 1, MaxCapacity: 20 }));
  await aas.send(new PutScheduledActionCommand({
    ...target,
    ScheduledActionName: `${fn}-${alias}-business-hours`,
    Schedule: "cron(0 8 ? * MON-FRI *)",
    Timezone: "America/New_York",
    ScalableTargetAction: { MinCapacity: 20, MaxCapacity: 20 },
  }));
  await aas.send(new PutScheduledActionCommand({
    ...target,
    ScheduledActionName: `${fn}-${alias}-after-hours`,
    Schedule: "cron(0 20 ? * MON-FRI *)",
    Timezone: "America/New_York",
    ScalableTargetAction: { MinCapacity: 1, MaxCapacity: 1 },
  }));
  console.log(`Scheduled ${target.ResourceId}: 20 from 08:00 to 20:00 on weekdays, 1 otherwise (America/New_York).`);
}

main().catch((err) => {
  console.error(err);
  process.exit(1);
});

The docs say to configure an initial provisioned concurrency value before Application Auto Scaling takes over. Target tracking is the other option, but it depends on ProvisionedConcurrencyUtilization; when a function goes quiet the metric stops, its alarms go to INSUFFICIENT_DATA and it can’t scale in. The script to find CloudWatch alarms stuck in INSUFFICIENT_DATA surfaces those alarms.

Before shrinking, right-size the function itself: memory drives the GB-second price, so finding Lambda functions with too much memory lowers the provisioned concurrency cost too, and moving Lambda functions to Graviton (arm64) cuts the rate by 20%. Shrinking the cold start itself is often the cheaper fix; the guide to reduce AWS SDK v3 bundle size and cold starts in Lambda shows how much a smaller bundle saves before you pay for pre-warmed environments.

Troubleshooting

  • A busy alias shows “IDLE”. Callers may invoke $LATEST or another alias instead of the one with provisioned concurrency. Compare with the script to get Lambda invocation counts for the last 24 hours; if the function is busy but the alias isn’t, fix the event source or caller.
  • failed: ResourceConflictException in the Action column. The SDK describes this as another operation being in progress on the resource. Wait and re-run.
  • Status is FAILED. Lambda couldn’t allocate the environments; the StatusReason field in the same response says why. The docs cap provisioned concurrency at the account’s unreserved concurrency minus 100.
  • Spillover is high. The alias needs more provisioned concurrency, not less. Investigating Lambda errors with CloudWatch shows whether cold starts cause timeouts.

Ask ChatWithCloud instead

For a one-off look, ask ChatWithCloud “Which Lambda aliases have provisioned concurrency, and how many environments does each have?” It writes AWS SDK for JavaScript v2 code, runs it locally with your profile and explains the result; how ChatWithCloud turns questions into AWS SDK calls covers the loop, and asking AI about Lambda errors in your account shows the same approach for failures. It runs changes without a confirmation step, so connect ChatWithCloud to a read-only profile and delete configs with the script above. More reports are in the AWS practical examples hub.

Frequently asked questions

Do you pay for Lambda provisioned concurrency when the function isn’t invoked?

Yes. The enabled charge runs per GB-second for every allocated environment from the time you configure it until you remove it, and Lambda also bills initialization for environments that never serve a request.

How do I turn off provisioned concurrency?

Call DeleteProvisionedConcurrencyConfig with the function name and the alias or version as Qualifier, or remove it in the console under Configuration, Concurrency. If Application Auto Scaling manages it, deregister the scalable target first or it may set it again.

What utilization makes provisioned concurrency worth it on cost alone?

Around 60%. At US East (N. Virginia) prices as of September 2026, that’s where the lower duration rate on provisioned environments offsets the enabled charge, for both x86_64 and arm64.

Can I set provisioned concurrency on $LATEST?

No. It has to be on a published version or an alias, and callers must invoke that qualifier to use it.

Related guides

Ask your AWS account in plain English

Your first 15 runs are free, with no OpenAI key needed.

npx chatwithcloud