Find Idle Amazon MSK Clusters

Bundles of blue and yellow network cables plugged into a patch panel in a data center rack

Photo by Wonderlane on Unsplash

Idle MSK cluster cost is the broker-hours and storage you pay for while no producer writes and no consumer reads. A provisioned cluster bills per broker per hour until you delete it, because MSK has no stop action. Find idle ones with ListClustersV2 from @aws-sdk/client-kafka, then read the free BytesInPerSec and BytesOutPerSec metrics in the AWS/Kafka namespace for each broker or topic.

Amazon Managed Streaming for Apache Kafka (MSK) clusters tend to outlive the project that needed them: a proof of concept for event streaming, a change data capture pipeline that was replaced, a test cluster per team. Three kafka.m5.large brokers with 1,000 GiB each cost about $760 a month in US East (N. Virginia) whether they carry traffic or not.

This example is for platform and data engineers who want a list of those clusters and a number for the idle MSK cluster cost. The script covers provisioned (Standard and Express brokers) and serverless clusters in the Regions you pass. It only reports; it never changes a cluster.

What does an idle MSK cluster cost?

Provisioned clusters bill for broker instance hours and for the EBS storage attached to each broker. Serverless clusters bill per cluster-hour, per partition-hour and per GB in and out. As of September 2026, the AWS Price List (published 11 September 2026) shows these on-demand rates in US East (N. Virginia):

Item Rate Per 730-hour month
kafka.t3.small broker $0.0456 per hour $33.29
kafka.m5.large broker $0.21 per hour $153.30
kafka.m7g.large broker $0.204 per hour $148.92
express.m7g.large broker $0.408 per hour $297.84
Broker storage (Standard and Express) $0.10 per GB-month –
Serverless cluster $0.75 per cluster-hour $547.50
Serverless partition $0.0015 per partition-hour $1.095

Check the Amazon MSK pricing page for your Region and for tiered storage, provisioned storage throughput and data transfer, which the script leaves out. Express brokers bill for the storage you use rather than a fixed volume, so the script counts only their broker-hours.

Worked example: a cluster of 3 kafka.m5.large brokers with 1,000 GiB each costs 3 × $0.21 × 730 = $459.90 in broker-hours plus 3 × 1,000 × $0.10 = $300.00 in storage, $759.90 a month. An idle serverless cluster with 50 partitions costs $0.75 × 730 = $547.50 plus 50 × $0.0015 × 730 = $54.75, or $602.25, before any data moves.

How do you tell that an MSK cluster is idle?

Client traffic is the clearest signal. At the free DEFAULT monitoring level, provisioned clusters publish BytesInPerSec (bytes received from clients) and BytesOutPerSec (bytes sent to clients) with the Cluster Name and Broker ID dimensions. Serverless clusters publish the same two metrics per topic, with Cluster Name and Topic. The script uses the highest hourly Maximum across all brokers or topics, so one burst is enough to mark a cluster as used.

It labels each cluster:

  • IDLE: no traffic metrics when CloudWatch has no BytesInPerSec or BytesOutPerSec series for it at all. The MSK docs say these metrics appear only after you create a topic, so a cluster that never had one shows up here.
  • IDLE: traffic below threshold when both peaks stay under --min-bytes (default 1,024 bytes per second).
  • no producers: check consumers when nothing is written but something is still read, for example a consumer replaying old data. Confirm before calling it dead.
  • too new to judge for clusters younger than the window.

ConnectionCount isn’t used for the verdict because it includes inter-broker connections, so it’s never zero on a healthy cluster.

What does the script do?

  1. Lists clusterspaginateListClustersV2 returns provisioned and serverless clusters with broker type, broker count and per-broker EBS volume size, so no separate DescribeClusterV2 call is needed.
  2. Finds the traffic seriespaginateListMetrics on AWS/Kafka filtered by Cluster Name, keeping per-broker series (provisioned) or per-topic series (serverless).
  3. Reads peaksGetMetricData with hourly Maximum, in batches of up to 500 queries.
  4. Prices itBroker-hours plus storage for provisioned clusters, the cluster-hour charge for serverless ones.
  5. ReportsA table, a monthly total for idle clusters and an optional CSV.

ListMetrics only returns metrics that reported data in the past two weeks, which is why --days is capped at 14.

Prerequisites

Which IAM permissions does it need?

idle-msk-report-policy.json

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "ReadMskClustersAndMetrics",
      "Effect": "Allow",
      "Action": [
        "kafka:ListClustersV2",
        "cloudwatch:ListMetrics",
        "cloudwatch:GetMetricData"
      ],
      "Resource": "*"
    }
  ]
}

All three actions are read-only and work across the account, so the policy uses "Resource": "*". If you extend the script, the IAM policy generator for TypeScript AWS SDK code lists any new actions.

The script to find idle MSK clusters

find-idle-msk-clusters.ts

// find-idle-msk-clusters.ts
// Lists Amazon MSK clusters (provisioned and serverless) with their size, peak client traffic over
// the last N days and an estimated monthly cost, and flags clusters that nobody writes to or reads from.
// Report only: it never changes or deletes a cluster.
// Usage: npx tsx find-idle-msk-clusters.ts [--regions us-east-1,eu-west-1] [--days 14] [--min-bytes 1024] [--csv msk.csv]
import { writeFileSync } from "node:fs";
import { KafkaClient, paginateListClustersV2, type Cluster } from "@aws-sdk/client-kafka";
import {
  CloudWatchClient,
  paginateGetMetricData,
  paginateListMetrics,
  type Dimension,
  type MetricDataQuery,
} from "@aws-sdk/client-cloudwatch";

const args = process.argv.slice(2);
const flag = (name: string): string | undefined => {
  const i = args.indexOf(name);
  return i >= 0 ? args[i + 1] : undefined;
};
const regions = (flag("--regions") ?? process.env.AWS_REGION ?? "us-east-1").split(",").map((r) => r.trim()).filter(Boolean);
const days = Number(flag("--days") ?? 14);
const minBytes = Number(flag("--min-bytes") ?? 1024); // peak bytes per second below this counts as no traffic
const csvPath = flag("--csv");

// us-east-1 on-demand USD per broker-hour. AWS Price List (AmazonMSK), published 11 September 2026.
const BROKER_HOUR: Record<string, number> = {
  "kafka.t3.small": 0.0456,
  "kafka.m5.large": 0.21,
  "kafka.m5.xlarge": 0.42,
  "kafka.m5.2xlarge": 0.84,
  "kafka.m5.4xlarge": 1.68,
  "kafka.m7g.large": 0.204,
  "kafka.m7g.xlarge": 0.408,
  "kafka.m7g.2xlarge": 0.816,
  "kafka.m7g.4xlarge": 1.632,
  "express.m7g.large": 0.408,
  "express.m7g.xlarge": 0.816,
  "express.m7g.2xlarge": 1.632,
};
const STORAGE_GB_MONTH = 0.1; // provisioned broker storage
const SERVERLESS_CLUSTER_HOUR = 0.75; // plus partition-hours and data in/out, not counted here
const HOURS_PER_MONTH = 730;

interface Row {
  Region: string;
  Cluster: string;
  Type: string;
  State: string;
  Brokers: number | string;
  BrokerType: string;
  StorageGiB: number | string;
  AgeDays: number;
  PeakInBps: number | string;
  PeakOutBps: number | string;
  PerMonth: string;
  Verdict: string;
}

/** Monthly estimate: broker-hours plus provisioned storage, or the serverless cluster-hour charge. */
function monthlyCost(c: Cluster): string {
  if (c.ClusterType === "SERVERLESS") return `$${(SERVERLESS_CLUSTER_HOUR * HOURS_PER_MONTH).toFixed(2)}+`;
  const p = c.Provisioned;
  const type = p?.BrokerNodeGroupInfo?.InstanceType ?? "";
  const brokers = p?.NumberOfBrokerNodes ?? 0;
  const rate = BROKER_HOUR[type];
  if (rate === undefined) return "unpriced type";
  const gib = p?.BrokerNodeGroupInfo?.StorageInfo?.EbsStorageInfo?.VolumeSize ?? 0; // per broker; Express has none
  const total = brokers * rate * HOURS_PER_MONTH + brokers * gib * STORAGE_GB_MONTH;
  return `$${(Math.round(total * 100) / 100).toFixed(2)}`;
}

/** Every BytesInPerSec/BytesOutPerSec series for the cluster: per broker (provisioned) or per topic (serverless). */
async function trafficSeries(cw: CloudWatchClient, name: string, serverless: boolean): Promise<{ metric: string; dims: Dimension[] }[]> {
  const series: { metric: string; dims: Dimension[] }[] = [];
  for (const metric of ["BytesInPerSec", "BytesOutPerSec"]) {
    for await (const page of paginateListMetrics({ client: cw }, {
      Namespace: "AWS/Kafka",
      MetricName: metric,
      Dimensions: [{ Name: "Cluster Name", Value: name }],
    })) {
      for (const m of page.Metrics ?? []) {
        const keys = (m.Dimensions ?? []).map((d) => d.Name).sort().join("|");
        const wanted = serverless ? "Cluster Name|Topic" : "Broker ID|Cluster Name";
        if (keys === wanted) series.push({ metric, dims: m.Dimensions ?? [] });
      }
    }
  }
  return series;
}

/** Highest hourly Maximum across all series of each metric, in bytes per second. */
async function peaks(cw: CloudWatchClient, series: { metric: string; dims: Dimension[] }[]): Promise<Record<string, number>> {
  const end = new Date();
  const start = new Date(end.getTime() - days * 86_400_000);
  const peak: Record<string, number> = { BytesInPerSec: 0, BytesOutPerSec: 0 };
  for (let i = 0; i < series.length; i += 500) { // GetMetricData takes up to 500 queries per request
    const chunk = series.slice(i, i + 500);
    const queries: MetricDataQuery[] = chunk.map((s, j) => ({
      Id: `m${j}`,
      MetricStat: { Metric: { Namespace: "AWS/Kafka", MetricName: s.metric, Dimensions: s.dims }, Period: 3600, Stat: "Maximum" },
    }));
    for await (const page of paginateGetMetricData({ client: cw }, { StartTime: start, EndTime: end, MetricDataQueries: queries })) {
      for (const r of page.MetricDataResults ?? []) {
        const metric = chunk[Number((r.Id ?? "m0").slice(1))]?.metric ?? "";
        for (const v of r.Values ?? []) peak[metric] = Math.max(peak[metric] ?? 0, v);
      }
    }
  }
  return peak;
}

async function scanRegion(region: string): Promise<Row[]> {
  const kafka = new KafkaClient({ region });
  const cw = new CloudWatchClient({ region });
  const rows: Row[] = [];
  for await (const page of paginateListClustersV2({ client: kafka }, {})) {
    for (const c of page.ClusterInfoList ?? []) {
      const name = c.ClusterName ?? "";
      const serverless = c.ClusterType === "SERVERLESS";
      const ageDays = c.CreationTime ? Math.floor((Date.now() - c.CreationTime.getTime()) / 86_400_000) : 0;
      const series = await trafficSeries(cw, name, serverless);
      const peak = series.length ? await peaks(cw, series) : undefined;
      const inBps = peak?.BytesInPerSec ?? 0;
      const outBps = peak?.BytesOutPerSec ?? 0;
      let verdict = "in use";
      if (c.State !== "ACTIVE") verdict = `skipped: ${c.State ?? "unknown"}`;
      else if (ageDays < days) verdict = "too new to judge";
      else if (!peak) verdict = "IDLE: no traffic metrics";
      else if (inBps < minBytes && outBps < minBytes) verdict = "IDLE: traffic below threshold";
      else if (inBps < minBytes) verdict = "no producers: check consumers";
      const p = c.Provisioned;
      rows.push({
        Region: region,
        Cluster: name,
        Type: serverless ? "serverless" : "provisioned",
        State: c.State ?? "",
        Brokers: serverless ? "-" : p?.NumberOfBrokerNodes ?? 0,
        BrokerType: serverless ? "-" : p?.BrokerNodeGroupInfo?.InstanceType ?? "",
        StorageGiB: serverless ? "-" : p?.BrokerNodeGroupInfo?.StorageInfo?.EbsStorageInfo?.VolumeSize ?? "-",
        AgeDays: ageDays,
        PeakInBps: peak ? Math.round(inBps) : "-",
        PeakOutBps: peak ? Math.round(outBps) : "-",
        PerMonth: monthlyCost(c),
        Verdict: verdict,
      });
    }
  }
  return rows;
}

function toCsv(rows: Row[]): string {
  const cols = Object.keys(rows[0] ?? {}) as (keyof Row)[];
  const cell = (v: string | number) => `"${String(v).replace(/"/g, '""')}"`;
  return [cols.join(","), ...rows.map((r) => cols.map((c) => cell(r[c])).join(","))].join("\n") + "\n";
}

async function main(): Promise<void> {
  if (!Number.isInteger(days) || days < 1 || days > 14) throw new Error("--days must be 1 to 14: ListMetrics only sees metrics with data in the last two weeks");
  if (!Number.isFinite(minBytes) || minBytes < 0) throw new Error("--min-bytes must be 0 or more");
  const rows: Row[] = [];
  for (const region of regions) {
    try {
      rows.push(...(await scanRegion(region)));
    } catch (err) {
      console.error(`${region}: ${err instanceof Error ? `${err.name}: ${err.message}` : String(err)}`);
    }
  }
  if (rows.length === 0) {
    console.log(`No MSK clusters in ${regions.join(", ")}.`);
    return;
  }
  console.table(rows);
  const idle = rows.filter((r) => r.Verdict.startsWith("IDLE"));
  const monthly = idle.reduce((s, r) => s + (Number(r.PerMonth.replace(/[$+]/g, "")) || 0), 0);
  console.log(`${idle.length} of ${rows.length} MSK clusters idle for ${days} days: at least $${monthly.toFixed(2)} a month (us-east-1 on-demand prices).`);
  if (csvPath) {
    writeFileSync(csvPath, toCsv(rows));
    console.log(`Wrote ${rows.length} rows to ${csvPath}`);
  }
}

main().catch((err) => {
  console.error(err);
  process.exit(1);
});

How do you run it?

Terminal

npm install @aws-sdk/client-kafka @aws-sdk/client-cloudwatch
npm install --save-dev tsx typescript @types/node

# Two Regions, 14-day window, CSV for the platform team
AWS_PROFILE=readonly npx tsx find-idle-msk-clusters.ts --regions us-east-1,eu-west-1 --csv idle-msk.csv

# Treat anything under 10 KB/s at peak as idle
AWS_PROFILE=readonly npx tsx find-idle-msk-clusters.ts --min-bytes 10240

Sample output

Output

┌─────────┬─────────────┬───────────────────┬───────────────┬──────────┬─────────┬───────────────────┬────────────┬─────────┬───────────┬────────────┬────────────┬─────────────────────────────────┐
│ (index) │ Region      │ Cluster           │ Type          │ State    │ Brokers │ BrokerType        │ StorageGiB │ AgeDays │ PeakInBps │ PeakOutBps │ PerMonth   │ Verdict                         │
├─────────┼─────────────┼───────────────────┼───────────────┼──────────┼─────────┼───────────────────┼────────────┼─────────┼───────────┼────────────┼────────────┼─────────────────────────────────┤
│ 0       │ 'us-east-1' │ 'orders-events'   │ 'provisioned' │ 'ACTIVE' │ 3       │ 'kafka.m5.large'  │ 1000       │ 200     │ 250000    │ 250000     │ '$759.90'  │ 'in use'                        │
│ 1       │ 'us-east-1' │ 'poc-clickstream' │ 'provisioned' │ 'ACTIVE' │ 3       │ 'kafka.m7g.large' │ 100        │ 200     │ '-'       │ '-'        │ '$476.76'  │ 'IDLE: no traffic metrics'      │
│ 2       │ 'us-east-1' │ 'legacy-cdc'      │ 'provisioned' │ 'ACTIVE' │ 2       │ 'kafka.m5.large'  │ 500        │ 200     │ 12        │ 40000      │ '$406.60'  │ 'no producers: check consumers' │
│ 3       │ 'us-east-1' │ 'team-sandbox'    │ 'serverless'  │ 'ACTIVE' │ '-'     │ '-'               │ '-'        │ 200     │ 3         │ 3          │ '$547.50+' │ 'IDLE: traffic below threshold' │
└─────────┴─────────────┴───────────────────┴───────────────┴──────────┴─────────┴───────────────────┴────────────┴─────────┴───────────┴────────────┴────────────┴─────────────────────────────────┘
2 of 4 MSK clusters idle for 14 days: at least $1024.26 a month (us-east-1 on-demand prices).

This output comes from a test run against mocked AWS responses, so names and numbers are illustrative. poc-clickstream never had a topic with traffic and costs $476.76 a month. team-sandbox is a serverless cluster with a trickle of test messages; its $547.50 is a floor, because partition-hours and data charges come on top. legacy-cdc gets almost no writes but a consumer still reads 40 KB/s, so find that consumer before you act.

Delete, downsize or move to serverless?

Situation What to do
No topics or no traffic for weeks Confirm with the owner, save the cluster configuration and topic settings you want to keep, then delete it with DeleteCluster. There’s no way to pause it.
Consumers only List consumer groups and their lag with the Apache Kafka kafka-consumer-groups tool, then find out who still reads.
Steady but tiny traffic A smaller broker type through UpdateBrokerType, or a serverless cluster if the partition count is low.
Serverless cluster with test traffic Delete it; the $0.75 cluster-hour charge runs even when nothing is sent.

Clusters rarely go idle alone. The scripts to find idle Amazon EMR clusters and find idle Amazon OpenSearch Service domains catch the processing and search layers that often sat downstream of the same pipeline, and finding idle DMS replication instances covers the change data capture side. If the stream was feeding Kinesis-style consumers, the guide to write to Kinesis Data Streams with PutRecords in SDK v3 shows the managed alternative for small volumes. If that alternative is already running, the script to find idle Kinesis Data Streams checks it for the same kind of waste.

Troubleshooting

  • Every cluster shows “IDLE: no traffic metrics”. Check that the profile has cloudwatch:ListMetrics and that you scan the Region the cluster runs in; metrics are Regional.
  • PerMonth shows “unpriced type”. Add the broker type to BROKER_HOUR with its rate from the Price List. Rates vary by Region.
  • A cluster in MAINTENANCE or UPDATING is skipped. The script only judges ACTIVE clusters; run it again later.
  • An access-denied error for one Region. An SCP may block that Region; the steps to troubleshoot AWS IAM access denied errors show how to read the message.

Ask ChatWithCloud instead

For a quick answer, ask ChatWithCloud “Which MSK clusters had no incoming bytes in the last two weeks, and how many brokers do they have?” It writes AWS SDK for JavaScript v2 code, runs it on your machine with your profile and explains the result; how ChatWithCloud runs AWS SDK calls locally has the details. It uses one profile and Region per session and runs changes without a confirmation step, so connect ChatWithCloud to a read-only AWS profile and delete clusters yourself. Services launched after SDK v2’s end of support may be missing from its answers. More cleanup scripts are in the AWS practical examples library, and an AWS budget alert created with SDK v3 flags the next forgotten cluster. Once you’ve cut the idle MSK cluster cost, find unattached elastic network interfaces left in the cluster’s subnets.

Frequently asked questions

Can you stop an Amazon MSK cluster to save money?

No. The MSK API has actions to create, update, reboot a broker and delete a cluster, but none to stop one. A provisioned cluster bills for broker-hours and storage until you delete it.

Are the MSK CloudWatch metrics this script reads free?

Yes. The MSK docs say DEFAULT-level metrics are free; BytesInPerSec and BytesOutPerSec are in that level. You still pay the normal CloudWatch API charges for GetMetricData.

Does an idle MSK Serverless cluster cost anything?

Yes. As of September 2026 it bills $0.75 per cluster-hour in US East (N. Virginia) plus $0.0015 per partition-hour, even with no data flowing.

What should I save before deleting an MSK cluster?

Topic names, partition counts and retention settings, any custom broker configuration, and the consumer group offsets if someone might resume. Deleting the cluster deletes its data.

Related guides

Ask your AWS account in plain English

Your first 15 runs are free, with no OpenAI key needed.

npx chatwithcloud