Find Idle Amazon EMR Clusters

Large industrial machines standing still on an empty factory floor under overhead lights

Photo by Shavr IK on Unsplash

To find idle EMR clusters, list clusters in the RUNNING and WAITING states with ListClusters, then read the IsIdle metric in the AWS/ElasticMapReduce CloudWatch namespace for the last 24 hours. A cluster that reports 1 for almost every 5-minute check and has no pending or running steps is idle. Price its instances, then attach an auto-termination policy or terminate it.

An Amazon EMR cluster on EC2 bills for every instance, every second it runs, whether Spark is busy or the cluster has sat in WAITING since last Tuesday. Long-running clusters created for an experiment, a notebook session or a one-off backfill are easy to forget, because the console shows them as healthy.

This example is for data platform and FinOps engineers who want a list with numbers. The script finds idle EMR clusters in every Region you pass, shows how much of the window each one was idle, its most recent step, its auto-termination setting and an hourly cost. It never terminates anything. With --apply it attaches an auto-termination policy to idle clusters that have none, so the cluster shuts itself down the next time it goes quiet.

What does an idle EMR cluster cost?

You pay two rates for each node: the Amazon EMR fee and the Amazon EC2 instance price, per second with a one-minute minimum, plus any EBS volumes attached to the nodes. As of September 2026, the AWS Price List shows these on-demand rates in US East (N. Virginia):

Instance type EMR fee per hour EC2 per hour Total per hour Per 730-hour month
m5.xlarge $0.048 $0.192 $0.240 $175.20
m5.2xlarge $0.096 $0.384 $0.480 $350.40
r5.2xlarge $0.126 $0.504 $0.630 $459.90
m6g.xlarge $0.039 $0.154 $0.193 $140.89
r6g.xlarge $0.0504 $0.2016 $0.252 $183.96

Worked example: a proof-of-concept cluster with one m5.2xlarge primary node, three on-demand r5.2xlarge core nodes and four Spot r5.2xlarge task nodes costs $0.48 + 3 × $0.63 + 4 × $0.126 = $2.874 an hour before the Spot EC2 price. Left in WAITING for a month, that’s $2.874 × 730 = $2,098 plus the Spot instances. Spot prices change, so the script counts only the EMR fee for Spot nodes; the script to compare EC2 Spot price history across Availability Zones gives you the rest.

How does EMR decide a cluster is idle?

There are two signals, and they measure different things:

  • IsIdle is published by every cluster every five minutes. It’s 1 when no tasks or jobs are running and 0 otherwise. The EMR Management Guide warns that a 1 means the cluster was idle at that check, not for all five minutes, and suggests alarming only after it has been 1 for 30 minutes or more. The script uses the share of idle checks over the window and the unbroken idle time at the end of it.
  • Auto-termination uses a wider test on Amazon EMR 5.34.0 and later and 6.4.0 and later: no active YARN applications, HDFS utilization below 10%, no EMR notebook or EMR Studio connections, no on-cluster application UIs in use and no pending steps. Older supported releases only check YARN applications and Spark jobs.

Neither signal sees an engineer who is SSH’d in and running shell scripts, or non-YARN engines such as Presto, Trino and HBase. On 6.4.0 and later, a job can touch /emr/metricscollector/isbusy on the primary node to tell EMR it’s still busy.

What does the script do?

  1. Lists live clusterspaginateListClusters with ClusterStates: ["RUNNING", "WAITING"], then DescribeCluster for the release label and name.
  2. Counts running nodespaginateListInstances with InstanceStates: ["RUNNING"] works for instance groups and instance fleets alike and reports each node’s type and whether it’s ON_DEMAND or SPOT.
  3. Reads the newest stepListSteps returns steps newest first, so the first page shows the last step and any that are pending or running.
  4. Measures idlenesspaginateGetMetricData fetches IsIdle at 5-minute resolution for --hours (default 24), newest first.
  5. Checks auto-terminationGetAutoTerminationPolicy on releases that support it (5.30.0+ and 6.1.0+).
  6. Acts only with --applyFor IDLE and MOSTLY IDLE clusters without a policy, PutAutoTerminationPolicy with --idle-timeout seconds (default 3600). It never calls TerminateJobFlows.

Prerequisites

Which IAM permissions does it need?

idle-emr-report-policy.json

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "ReadEmrClusters",
      "Effect": "Allow",
      "Action": [
        "elasticmapreduce:ListClusters",
        "elasticmapreduce:DescribeCluster",
        "elasticmapreduce:ListInstances",
        "elasticmapreduce:ListSteps",
        "elasticmapreduce:GetAutoTerminationPolicy",
        "cloudwatch:GetMetricData"
      ],
      "Resource": "*"
    },
    {
      "Sid": "AttachAutoTerminationWithApply",
      "Effect": "Allow",
      "Action": "elasticmapreduce:PutAutoTerminationPolicy",
      "Resource": "arn:aws:elasticmapreduce:*:123456789012:cluster/*"
    }
  ]
}

Replace the account ID, and drop the second statement for a report-only role. If you extend the script, the IAM policy generator for TypeScript code lists the actions the new version calls.

The script to find idle EMR clusters

find-idle-emr-clusters.ts

// find-idle-emr-clusters.ts
// Lists running and waiting Amazon EMR clusters with how much of the last N hours they were idle (IsIdle metric),
// their last step, their auto-termination policy and an on-demand hourly cost estimate.
// Report only by default. --apply attaches an auto-termination policy to idle clusters that have none.
// It never terminates a cluster.
// Usage: npx tsx find-idle-emr-clusters.ts [--regions us-east-1,eu-west-1] [--hours 24] [--csv emr.csv] [--apply] [--idle-timeout 3600]
import { writeFileSync } from "node:fs";
import {
  DescribeClusterCommand,
  EMRClient,
  GetAutoTerminationPolicyCommand,
  ListStepsCommand,
  PutAutoTerminationPolicyCommand,
  paginateListClusters,
  paginateListInstances,
  type Cluster,
} from "@aws-sdk/client-emr";
import { CloudWatchClient, paginateGetMetricData } from "@aws-sdk/client-cloudwatch";

const args = process.argv.slice(2);
const flag = (name: string): string | undefined => {
  const i = args.indexOf(name);
  return i >= 0 ? args[i + 1] : undefined;
};
const regions = (flag("--regions") ?? process.env.AWS_REGION ?? "us-east-1").split(",").map((r) => r.trim()).filter(Boolean);
const hours = Number(flag("--hours") ?? 24);
const csvPath = flag("--csv");
const apply = args.includes("--apply");
const idleTimeout = Number(flag("--idle-timeout") ?? 3600); // seconds; EMR accepts 60 to 604800

// On-demand USD per instance-hour in us-east-1: [Amazon EMR fee, Amazon EC2 price]. AWS Price List, September 2026.
const PRICE: Record<string, [number, number]> = {
  "m5.xlarge": [0.048, 0.192],
  "m5.2xlarge": [0.096, 0.384],
  "m5.4xlarge": [0.192, 0.768],
  "m6g.xlarge": [0.039, 0.154],
  "m7g.xlarge": [0.0408, 0.1632],
  "r5.xlarge": [0.063, 0.252],
  "r5.2xlarge": [0.126, 0.504],
  "r6g.xlarge": [0.0504, 0.2016],
  "c5.xlarge": [0.043, 0.17],
};
const HOURS_PER_MONTH = 730;

interface Row {
  Region: string;
  ClusterId: string;
  Name: string;
  State: string;
  Release: string;
  Nodes: string;
  IdlePct: number | string;
  IdleForHours: number | string;
  LastStep: string;
  AutoTermination: string;
  PerHour: string;
  Verdict: string;
  Action: string;
}

/** Auto-termination needs Amazon EMR 5.30.0+ or 6.1.0+ (6.14.0+ in some newer Regions). */
function supportsAutoTermination(release: string): boolean {
  const m = /^emr-(\d+)\.(\d+)\./.exec(release);
  if (!m) return false;
  const [major, minor] = [Number(m[1]), Number(m[2])];
  return major >= 7 || (major === 6 && minor >= 1) || (major === 5 && minor >= 30);
}

async function nodesAndCost(emr: EMRClient, clusterId: string): Promise<{ nodes: string; perHour: number; unpriced: boolean }> {
  const byType = new Map<string, number>();
  let perHour = 0;
  let unpriced = false;
  for await (const page of paginateListInstances({ client: emr }, { ClusterId: clusterId, InstanceStates: ["RUNNING"] })) {
    for (const inst of page.Instances ?? []) {
      const type = inst.InstanceType ?? "unknown";
      byType.set(type, (byType.get(type) ?? 0) + 1);
      const price = PRICE[type];
      if (!price) unpriced = true;
      else perHour += inst.Market === "SPOT" ? price[0] : price[0] + price[1]; // Spot: EMR fee only, EC2 Spot price varies
    }
  }
  const nodes = [...byType].map(([t, n]) => `${n}x ${t}`).join(", ") || "none running";
  return { nodes, perHour, unpriced };
}

async function lastStep(emr: EMRClient, clusterId: string): Promise<{ label: string; active: boolean }> {
  const res = await emr.send(new ListStepsCommand({ ClusterId: clusterId })); // newest steps first
  const steps = res.Steps ?? [];
  const active = steps.some((s) => ["PENDING", "RUNNING", "CANCEL_PENDING"].includes(s.Status?.State ?? ""));
  const newest = steps[0];
  if (!newest) return { label: "no steps", active };
  const t = newest.Status?.Timeline;
  const when = t?.EndDateTime ?? t?.StartDateTime ?? t?.CreationDateTime;
  return { label: `${newest.Status?.State ?? "?"} ${when ? when.toISOString().slice(0, 16).replace("T", " ") : ""}`.trim(), active };
}

async function idleStats(cw: CloudWatchClient, clusterId: string): Promise<{ pct: number; trailingHours: number } | undefined> {
  const end = new Date();
  const start = new Date(end.getTime() - hours * 3_600_000);
  const values: number[] = [];
  for await (const page of paginateGetMetricData({ client: cw }, {
    StartTime: start,
    EndTime: end,
    ScanBy: "TimestampDescending", // newest datapoint first
    MetricDataQueries: [{
      Id: "idle",
      MetricStat: {
        Metric: { Namespace: "AWS/ElasticMapReduce", MetricName: "IsIdle", Dimensions: [{ Name: "JobFlowId", Value: clusterId }] },
        Period: 300,
        Stat: "Maximum", // EMR checks every 5 minutes: 1 = idle at that check, 0 = work was running
      },
    }],
  })) {
    for (const r of page.MetricDataResults ?? []) values.push(...(r.Values ?? []));
  }
  if (values.length === 0) return undefined;
  const idle = values.filter((v) => v >= 1).length;
  const trailing = values.findIndex((v) => v < 1);
  const trailingPoints = trailing === -1 ? values.length : trailing;
  return { pct: Math.round((idle / values.length) * 100), trailingHours: Math.round((trailingPoints * 5) / 60 * 10) / 10 };
}

async function autoTermination(emr: EMRClient, c: Cluster): Promise<string> {
  if (!supportsAutoTermination(c.ReleaseLabel ?? "")) return "not supported";
  try {
    const res = await emr.send(new GetAutoTerminationPolicyCommand({ ClusterId: c.Id }));
    const seconds = res.AutoTerminationPolicy?.IdleTimeout;
    return seconds === undefined ? "none" : `${Math.round(seconds / 60)} min`;
  } catch (err) {
    return `unknown (${err instanceof Error ? err.name : "error"})`;
  }
}

async function scanRegion(region: string): Promise<Row[]> {
  const emr = new EMRClient({ region });
  const cw = new CloudWatchClient({ region });
  const rows: Row[] = [];
  for await (const page of paginateListClusters({ client: emr }, { ClusterStates: ["RUNNING", "WAITING"] })) {
    for (const summary of page.Clusters ?? []) {
      const id = summary.Id ?? "";
      const { Cluster: c } = await emr.send(new DescribeClusterCommand({ ClusterId: id }));
      if (!c) continue;
      const [cost, step, idle, policy] = await Promise.all([
        nodesAndCost(emr, id), lastStep(emr, id), idleStats(cw, id), autoTermination(emr, c),
      ]);
      let verdict = "in use";
      if (!idle) verdict = "no IsIdle data: check the cluster age";
      else if (step.active) verdict = "in use: steps pending or running";
      else if (idle.pct >= 95) verdict = "IDLE";
      else if (idle.pct >= 80) verdict = "MOSTLY IDLE";
      const flagged = verdict === "IDLE" || verdict === "MOSTLY IDLE";
      let action = flagged && policy === "none" ? "would add auto-termination (--apply)" : "-";
      if (apply && flagged && policy === "none") {
        try {
          await emr.send(new PutAutoTerminationPolicyCommand({ ClusterId: id, AutoTerminationPolicy: { IdleTimeout: idleTimeout } }));
          action = `auto-termination set: ${Math.round(idleTimeout / 60)} min`;
        } catch (err) {
          action = `failed: ${err instanceof Error ? err.name : String(err)}`;
        }
      }
      rows.push({
        Region: region,
        ClusterId: id,
        Name: c.Name ?? "",
        State: c.Status?.State ?? "",
        Release: c.ReleaseLabel ?? "",
        Nodes: cost.nodes,
        IdlePct: idle ? idle.pct : "-",
        IdleForHours: idle ? idle.trailingHours : "-",
        LastStep: step.label,
        AutoTermination: policy,
        PerHour: cost.unpriced ? `>= $${cost.perHour.toFixed(2)} (unpriced type)` : `$${cost.perHour.toFixed(2)}`,
        Verdict: verdict,
        Action: action,
      });
    }
  }
  return rows;
}

function toCsv(rows: Row[]): string {
  const cols = Object.keys(rows[0] ?? {}) as (keyof Row)[];
  const cell = (v: string | number) => `"${String(v).replace(/"/g, '""')}"`;
  return [cols.join(","), ...rows.map((r) => cols.map((c) => cell(r[c])).join(","))].join("\n") + "\n";
}

async function main(): Promise<void> {
  if (!Number.isInteger(idleTimeout) || idleTimeout < 60 || idleTimeout > 604_800) {
    throw new Error("--idle-timeout must be a whole number of seconds from 60 to 604800");
  }
  const rows: Row[] = [];
  for (const region of regions) {
    try {
      rows.push(...(await scanRegion(region)));
    } catch (err) {
      console.error(`${region}: ${err instanceof Error ? `${err.name}: ${err.message}` : String(err)}`);
    }
  }
  if (rows.length === 0) {
    console.log(`No running or waiting EMR clusters in ${regions.join(", ")}.`);
    return;
  }
  console.table(rows);
  const idle = rows.filter((r) => r.Verdict === "IDLE" || r.Verdict === "MOSTLY IDLE");
  const perHour = idle.reduce((sum, r) => sum + Number(/\$([\d.]+)/.exec(r.PerHour)?.[1] ?? 0), 0);
  console.log(`${idle.length} of ${rows.length} clusters idle over the last ${hours} hours: about $${perHour.toFixed(2)} an hour, ` +
    `$${(perHour * HOURS_PER_MONTH).toFixed(0)} a month if left running (us-east-1 on-demand prices).`);
  if (csvPath) {
    writeFileSync(csvPath, toCsv(rows));
    console.log(`Wrote ${rows.length} rows to ${csvPath}`);
  }
}

main().catch((err) => {
  console.error(err);
  process.exit(1);
});

How do you run it?

Terminal

npm install @aws-sdk/client-emr @aws-sdk/client-cloudwatch
npm install --save-dev tsx typescript @types/node

# Report only: last 24 hours, two Regions, CSV for the owners
AWS_PROFILE=readonly npx tsx find-idle-emr-clusters.ts --regions us-east-1,eu-west-1 --csv idle-emr.csv

# Attach a 30-minute auto-termination policy to idle clusters that have none
AWS_PROFILE=emr-admin npx tsx find-idle-emr-clusters.ts --regions us-east-1 --hours 48 --apply --idle-timeout 1800

Sample output

Output

┌─────────┬─────────────┬───────────────────┬────────────────────┬───────────┬──────────────┬────────────────────────────────┬─────────┬──────────────┬──────────────────────────────┬─────────────────┬─────────┬────────────────────────────────────┬────────────────────────────────────────┐
│ (index) │ Region      │ ClusterId         │ Name               │ State     │ Release      │ Nodes                          │ IdlePct │ IdleForHours │ LastStep                     │ AutoTermination │ PerHour │ Verdict                            │ Action                                 │
├─────────┼─────────────┼───────────────────┼────────────────────┼───────────┼──────────────┼────────────────────────────────┼─────────┼──────────────┼──────────────────────────────┼─────────────────┼─────────┼────────────────────────────────────┼────────────────────────────────────────┤
│ 0       │ 'us-east-1' │ 'j-2AXXXXXXGAPLF' │ 'nightly-etl'      │ 'WAITING' │ 'emr-7.5.0'  │ '1x m5.xlarge, 4x r5.2xlarge'  │ 86      │ 3.3          │ 'COMPLETED 2026-09-28 04:41' │ 'none'          │ '$2.76' │ 'MOSTLY IDLE'                      │ 'would add auto-termination (--apply)' │
│ 1       │ 'us-east-1' │ 'j-3BXXXXXXKQ7Z1' │ 'spark-poc-mlteam' │ 'WAITING' │ 'emr-6.15.0' │ '1x m5.2xlarge, 7x r5.2xlarge' │ 100     │ 24           │ 'COMPLETED 2026-09-02 16:20' │ 'none'          │ '$2.87' │ 'IDLE'                             │ 'would add auto-termination (--apply)' │
│ 2       │ 'us-east-1' │ 'j-1CXXXXXXW3M8P' │ 'adhoc-hive-2023'  │ 'WAITING' │ 'emr-5.29.0' │ '3x m5.xlarge'                 │ 100     │ 23.7         │ 'no steps'                   │ 'not supported' │ '$0.72' │ 'IDLE'                             │ '-'                                    │
│ 3       │ 'us-east-1' │ 'j-4DXXXXXXR9TT2' │ 'streaming-ingest' │ 'RUNNING' │ 'emr-7.5.0'  │ '1x m5.xlarge, 2x m5.2xlarge'  │ 0       │ 0            │ 'RUNNING 2026-09-14 08:00'   │ 'none'          │ '$1.20' │ 'in use: steps pending or running' │ '-'                                    │
└─────────┴─────────────┴───────────────────┴────────────────────┴───────────┴──────────────┴────────────────────────────────┴─────────┴──────────────┴──────────────────────────────┴─────────────────┴─────────┴────────────────────────────────────┴────────────────────────────────────────┘
3 of 4 clusters idle over the last 24 hours: about $6.35 an hour, $4636 a month if left running (us-east-1 on-demand prices).

IDs, names and numbers are illustrative. spark-poc-mlteam has been idle for the whole window and its last step finished weeks ago; it’s the one to act on. adhoc-hive-2023 runs Amazon EMR 5.29.0, which predates auto-termination, so the script can only report it: ask the owner and terminate it by hand. nightly-etl is mostly idle because it works a few hours a night and waits the rest of the day. That pattern is cheaper as a transient cluster that terminates after its steps, or with a short idle timeout.

Auto-terminate, resize or terminate?

Option What happens Use it when
Auto-termination policy EMR terminates the cluster after it has been idle for IdleTimeout (60 seconds to 7 days; 1 hour by default) The cluster is still used, but in bursts. You keep it for the next job and stop paying for the gaps after it
Terminate now All instances stop billing; HDFS data on the cluster is gone Nobody claims it. Check that outputs are in S3, not only in HDFS
EMR Serverless Applications auto-start on job submission and, by default, auto-stop after 15 idle minutes Spark or Hive jobs are occasional and you’d rather not size clusters at all

The Amazon EMR auto-termination policy documentation lists what to check before relying on it. It needs release 5.30.0 or 6.1.0 and later (6.14.0 in some newer Regions), it isn’t supported for non-YARN applications such as Presto, Trino or HBase, and the metrics collector must reach the public auto-termination endpoint. An API Gateway interface endpoint with private DNS in the cluster’s VPC breaks it. For SQL on data that’s already in S3, the guide to run an Athena query with AWS SDK v3 is another way to skip a standing cluster.

What should you check before acting?

  • Interactive users. SSH sessions and notebook kernels that aren’t running a job don’t count as activity for IsIdle. Ask in the team channel before a short timeout ends someone’s afternoon.
  • Data only in HDFS. Terminating deletes it. HDFS utilization in the console or the HDFSUtilization metric tells you whether there’s anything there.
  • Longer cycles. A weekly job looks idle in a 24-hour window. Re-run with --hours 168 before terminating anything with a schedule-like name.
  • Owners. Cluster tags tell you whom to ask; the script to find untagged AWS resources lists clusters nobody has claimed.

Idle analytics clusters rarely come alone. The same review usually turns up a warehouse or search domain nobody queries; the scripts to find idle Amazon Redshift clusters and find idle Amazon OpenSearch Service domains cover those, and an AWS budget alert created with SDK v3 warns you when the next forgotten cluster starts adding up. Data migrations can leave compute behind as well; the script to find idle AWS DMS replication instances left after a migration lists replication instances with no running tasks. Streaming ingestion leaves brokers behind too; the script to find idle Amazon MSK clusters flags Kafka clusters with no producers or consumers.

Troubleshooting

  • “no IsIdle data: check the cluster age”. The cluster started after the window began or is unreachable. EMR pulls metrics from the cluster, and an unreachable cluster reports nothing until it’s reachable again.
  • AutoTermination shows “unknown (…)”. GetAutoTerminationPolicy failed; the error name is in brackets. A missing elasticmapreduce:GetAutoTerminationPolicy permission is the usual cause.
  • PerHour shows “unpriced type”. Add the EMR fee and EC2 price for that type to PRICE. Prices differ by Region, so adjust them outside us-east-1.
  • AccessDeniedException in one Region only. An SCP may restrict that Region. The steps to troubleshoot AWS IAM access denied errors show how to read the message.

Ask ChatWithCloud instead

For a one-off check, ask ChatWithCloud “Which EMR clusters are in WAITING state, and when did their last step finish?” It writes AWS SDK for JavaScript v2 code, runs it locally with your profile and explains the answer; how ChatWithCloud runs AWS SDK code locally shows each step. It uses one profile and Region per session and runs changes without a confirmation step, so connect ChatWithCloud to a read-only AWS profile and leave termination to a person. More scripts like this are on the AWS practical examples hub.

Frequently asked questions

How do I know if an EMR cluster is idle?

Check the IsIdle CloudWatch metric for the cluster (dimension JobFlowId). A value of 1 for 30 minutes or more, with no pending steps, means no jobs or tasks ran in that time.

Do you pay for an EMR cluster in WAITING state?

Yes. WAITING means the cluster is up and waiting for work, and every running instance bills the EMR fee and the EC2 price until the cluster terminates.

Can I add auto-termination to a running EMR cluster?

Yes, on release 5.30.0 or 6.1.0 and later. Call PutAutoTerminationPolicy or run aws emr put-auto-termination-policy --cluster-id j-XXXX --auto-termination-policy IdleTimeout=3600.

Can I stop an EMR cluster instead of terminating it?

No. The EMR API has actions to terminate a cluster but none to stop and later restart one, the way you can with an EC2 instance. Terminate it and launch a new one when you need it, keeping your data in S3.

Related guides

Ask your AWS account in plain English

Your first 15 runs are free, with no OpenAI key needed.

npx chatwithcloud