To find idle EMR clusters, list clusters in the RUNNING and WAITING states with ListClusters, then read the IsIdle metric in the AWS/ElasticMapReduce CloudWatch namespace for the last 24 hours. A cluster that reports 1 for almost every 5-minute check and has no pending or running steps is idle. Price its instances, then attach an auto-termination policy or terminate it.
An Amazon EMR cluster on EC2 bills for every instance, every second it runs, whether Spark is busy or the cluster has sat in WAITING since last Tuesday. Long-running clusters created for an experiment, a notebook session or a one-off backfill are easy to forget, because the console shows them as healthy.
This example is for data platform and FinOps engineers who want a list with numbers. The script finds idle EMR clusters in every Region you pass, shows how much of the window each one was idle, its most recent step, its auto-termination setting and an hourly cost. It never terminates anything. With --apply it attaches an auto-termination policy to idle clusters that have none, so the cluster shuts itself down the next time it goes quiet.
What does an idle EMR cluster cost?
You pay two rates for each node: the Amazon EMR fee and the Amazon EC2 instance price, per second with a one-minute minimum, plus any EBS volumes attached to the nodes. As of September 2026, the AWS Price List shows these on-demand rates in US East (N. Virginia):
| Instance type | EMR fee per hour | EC2 per hour | Total per hour | Per 730-hour month |
|---|---|---|---|---|
| m5.xlarge | $0.048 | $0.192 | $0.240 | $175.20 |
| m5.2xlarge | $0.096 | $0.384 | $0.480 | $350.40 |
| r5.2xlarge | $0.126 | $0.504 | $0.630 | $459.90 |
| m6g.xlarge | $0.039 | $0.154 | $0.193 | $140.89 |
| r6g.xlarge | $0.0504 | $0.2016 | $0.252 | $183.96 |
Worked example: a proof-of-concept cluster with one m5.2xlarge primary node, three on-demand r5.2xlarge core nodes and four Spot r5.2xlarge task nodes costs $0.48 + 3 × $0.63 + 4 × $0.126 = $2.874 an hour before the Spot EC2 price. Left in WAITING for a month, that’s $2.874 × 730 = $2,098 plus the Spot instances. Spot prices change, so the script counts only the EMR fee for Spot nodes; the script to compare EC2 Spot price history across Availability Zones gives you the rest.
How does EMR decide a cluster is idle?
There are two signals, and they measure different things:
IsIdleis published by every cluster every five minutes. It’s 1 when no tasks or jobs are running and 0 otherwise. The EMR Management Guide warns that a 1 means the cluster was idle at that check, not for all five minutes, and suggests alarming only after it has been 1 for 30 minutes or more. The script uses the share of idle checks over the window and the unbroken idle time at the end of it.- Auto-termination uses a wider test on Amazon EMR 5.34.0 and later and 6.4.0 and later: no active YARN applications, HDFS utilization below 10%, no EMR notebook or EMR Studio connections, no on-cluster application UIs in use and no pending steps. Older supported releases only check YARN applications and Spark jobs.
Neither signal sees an engineer who is SSH’d in and running shell scripts, or non-YARN engines such as Presto, Trino and HBase. On 6.4.0 and later, a job can touch /emr/metricscollector/isbusy on the primary node to tell EMR it’s still busy.
What does the script do?
- Lists live clusters
paginateListClusterswithClusterStates: ["RUNNING", "WAITING"], thenDescribeClusterfor the release label and name. - Counts running nodes
paginateListInstanceswithInstanceStates: ["RUNNING"]works for instance groups and instance fleets alike and reports each node’s type and whether it’sON_DEMANDorSPOT. - Reads the newest step
ListStepsreturns steps newest first, so the first page shows the last step and any that are pending or running. - Measures idleness
paginateGetMetricDatafetchesIsIdleat 5-minute resolution for--hours(default 24), newest first. - Checks auto-termination
GetAutoTerminationPolicyon releases that support it (5.30.0+ and 6.1.0+). - Acts only with
--applyFor IDLE and MOSTLY IDLE clusters without a policy,PutAutoTerminationPolicywith--idle-timeoutseconds (default 3600). It never callsTerminateJobFlows.
Prerequisites
- Node.js 18 or later with
tsx, plus@aws-sdk/client-emrand@aws-sdk/client-cloudwatch. - An AWS profile the SDK can resolve, as covered in AWS SDK v3 credential providers: fromIni, fromSSO and assume role. Use a read-only one for the report.
- The Regions where you run EMR. If you’re not sure, the Cost Explorer script to get last month’s AWS cost broken down by service shows whether EMR is on your bill at all.
Which IAM permissions does it need?
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ReadEmrClusters",
"Effect": "Allow",
"Action": [
"elasticmapreduce:ListClusters",
"elasticmapreduce:DescribeCluster",
"elasticmapreduce:ListInstances",
"elasticmapreduce:ListSteps",
"elasticmapreduce:GetAutoTerminationPolicy",
"cloudwatch:GetMetricData"
],
"Resource": "*"
},
{
"Sid": "AttachAutoTerminationWithApply",
"Effect": "Allow",
"Action": "elasticmapreduce:PutAutoTerminationPolicy",
"Resource": "arn:aws:elasticmapreduce:*:123456789012:cluster/*"
}
]
}
Replace the account ID, and drop the second statement for a report-only role. If you extend the script, the IAM policy generator for TypeScript code lists the actions the new version calls.
The script to find idle EMR clusters
// find-idle-emr-clusters.ts
// Lists running and waiting Amazon EMR clusters with how much of the last N hours they were idle (IsIdle metric),
// their last step, their auto-termination policy and an on-demand hourly cost estimate.
// Report only by default. --apply attaches an auto-termination policy to idle clusters that have none.
// It never terminates a cluster.
// Usage: npx tsx find-idle-emr-clusters.ts [--regions us-east-1,eu-west-1] [--hours 24] [--csv emr.csv] [--apply] [--idle-timeout 3600]
import { writeFileSync } from "node:fs";
import {
DescribeClusterCommand,
EMRClient,
GetAutoTerminationPolicyCommand,
ListStepsCommand,
PutAutoTerminationPolicyCommand,
paginateListClusters,
paginateListInstances,
type Cluster,
} from "@aws-sdk/client-emr";
import { CloudWatchClient, paginateGetMetricData } from "@aws-sdk/client-cloudwatch";
const args = process.argv.slice(2);
const flag = (name: string): string | undefined => {
const i = args.indexOf(name);
return i >= 0 ? args[i + 1] : undefined;
};
const regions = (flag("--regions") ?? process.env.AWS_REGION ?? "us-east-1").split(",").map((r) => r.trim()).filter(Boolean);
const hours = Number(flag("--hours") ?? 24);
const csvPath = flag("--csv");
const apply = args.includes("--apply");
const idleTimeout = Number(flag("--idle-timeout") ?? 3600); // seconds; EMR accepts 60 to 604800
// On-demand USD per instance-hour in us-east-1: [Amazon EMR fee, Amazon EC2 price]. AWS Price List, September 2026.
const PRICE: Record<string, [number, number]> = {
"m5.xlarge": [0.048, 0.192],
"m5.2xlarge": [0.096, 0.384],
"m5.4xlarge": [0.192, 0.768],
"m6g.xlarge": [0.039, 0.154],
"m7g.xlarge": [0.0408, 0.1632],
"r5.xlarge": [0.063, 0.252],
"r5.2xlarge": [0.126, 0.504],
"r6g.xlarge": [0.0504, 0.2016],
"c5.xlarge": [0.043, 0.17],
};
const HOURS_PER_MONTH = 730;
interface Row {
Region: string;
ClusterId: string;
Name: string;
State: string;
Release: string;
Nodes: string;
IdlePct: number | string;
IdleForHours: number | string;
LastStep: string;
AutoTermination: string;
PerHour: string;
Verdict: string;
Action: string;
}
/** Auto-termination needs Amazon EMR 5.30.0+ or 6.1.0+ (6.14.0+ in some newer Regions). */
function supportsAutoTermination(release: string): boolean {
const m = /^emr-(\d+)\.(\d+)\./.exec(release);
if (!m) return false;
const [major, minor] = [Number(m[1]), Number(m[2])];
return major >= 7 || (major === 6 && minor >= 1) || (major === 5 && minor >= 30);
}
async function nodesAndCost(emr: EMRClient, clusterId: string): Promise<{ nodes: string; perHour: number; unpriced: boolean }> {
const byType = new Map<string, number>();
let perHour = 0;
let unpriced = false;
for await (const page of paginateListInstances({ client: emr }, { ClusterId: clusterId, InstanceStates: ["RUNNING"] })) {
for (const inst of page.Instances ?? []) {
const type = inst.InstanceType ?? "unknown";
byType.set(type, (byType.get(type) ?? 0) + 1);
const price = PRICE[type];
if (!price) unpriced = true;
else perHour += inst.Market === "SPOT" ? price[0] : price[0] + price[1]; // Spot: EMR fee only, EC2 Spot price varies
}
}
const nodes = [...byType].map(([t, n]) => `${n}x ${t}`).join(", ") || "none running";
return { nodes, perHour, unpriced };
}
async function lastStep(emr: EMRClient, clusterId: string): Promise<{ label: string; active: boolean }> {
const res = await emr.send(new ListStepsCommand({ ClusterId: clusterId })); // newest steps first
const steps = res.Steps ?? [];
const active = steps.some((s) => ["PENDING", "RUNNING", "CANCEL_PENDING"].includes(s.Status?.State ?? ""));
const newest = steps[0];
if (!newest) return { label: "no steps", active };
const t = newest.Status?.Timeline;
const when = t?.EndDateTime ?? t?.StartDateTime ?? t?.CreationDateTime;
return { label: `${newest.Status?.State ?? "?"} ${when ? when.toISOString().slice(0, 16).replace("T", " ") : ""}`.trim(), active };
}
async function idleStats(cw: CloudWatchClient, clusterId: string): Promise<{ pct: number; trailingHours: number } | undefined> {
const end = new Date();
const start = new Date(end.getTime() - hours * 3_600_000);
const values: number[] = [];
for await (const page of paginateGetMetricData({ client: cw }, {
StartTime: start,
EndTime: end,
ScanBy: "TimestampDescending", // newest datapoint first
MetricDataQueries: [{
Id: "idle",
MetricStat: {
Metric: { Namespace: "AWS/ElasticMapReduce", MetricName: "IsIdle", Dimensions: [{ Name: "JobFlowId", Value: clusterId }] },
Period: 300,
Stat: "Maximum", // EMR checks every 5 minutes: 1 = idle at that check, 0 = work was running
},
}],
})) {
for (const r of page.MetricDataResults ?? []) values.push(...(r.Values ?? []));
}
if (values.length === 0) return undefined;
const idle = values.filter((v) => v >= 1).length;
const trailing = values.findIndex((v) => v < 1);
const trailingPoints = trailing === -1 ? values.length : trailing;
return { pct: Math.round((idle / values.length) * 100), trailingHours: Math.round((trailingPoints * 5) / 60 * 10) / 10 };
}
async function autoTermination(emr: EMRClient, c: Cluster): Promise<string> {
if (!supportsAutoTermination(c.ReleaseLabel ?? "")) return "not supported";
try {
const res = await emr.send(new GetAutoTerminationPolicyCommand({ ClusterId: c.Id }));
const seconds = res.AutoTerminationPolicy?.IdleTimeout;
return seconds === undefined ? "none" : `${Math.round(seconds / 60)} min`;
} catch (err) {
return `unknown (${err instanceof Error ? err.name : "error"})`;
}
}
async function scanRegion(region: string): Promise<Row[]> {
const emr = new EMRClient({ region });
const cw = new CloudWatchClient({ region });
const rows: Row[] = [];
for await (const page of paginateListClusters({ client: emr }, { ClusterStates: ["RUNNING", "WAITING"] })) {
for (const summary of page.Clusters ?? []) {
const id = summary.Id ?? "";
const { Cluster: c } = await emr.send(new DescribeClusterCommand({ ClusterId: id }));
if (!c) continue;
const [cost, step, idle, policy] = await Promise.all([
nodesAndCost(emr, id), lastStep(emr, id), idleStats(cw, id), autoTermination(emr, c),
]);
let verdict = "in use";
if (!idle) verdict = "no IsIdle data: check the cluster age";
else if (step.active) verdict = "in use: steps pending or running";
else if (idle.pct >= 95) verdict = "IDLE";
else if (idle.pct >= 80) verdict = "MOSTLY IDLE";
const flagged = verdict === "IDLE" || verdict === "MOSTLY IDLE";
let action = flagged && policy === "none" ? "would add auto-termination (--apply)" : "-";
if (apply && flagged && policy === "none") {
try {
await emr.send(new PutAutoTerminationPolicyCommand({ ClusterId: id, AutoTerminationPolicy: { IdleTimeout: idleTimeout } }));
action = `auto-termination set: ${Math.round(idleTimeout / 60)} min`;
} catch (err) {
action = `failed: ${err instanceof Error ? err.name : String(err)}`;
}
}
rows.push({
Region: region,
ClusterId: id,
Name: c.Name ?? "",
State: c.Status?.State ?? "",
Release: c.ReleaseLabel ?? "",
Nodes: cost.nodes,
IdlePct: idle ? idle.pct : "-",
IdleForHours: idle ? idle.trailingHours : "-",
LastStep: step.label,
AutoTermination: policy,
PerHour: cost.unpriced ? `>= $${cost.perHour.toFixed(2)} (unpriced type)` : `$${cost.perHour.toFixed(2)}`,
Verdict: verdict,
Action: action,
});
}
}
return rows;
}
function toCsv(rows: Row[]): string {
const cols = Object.keys(rows[0] ?? {}) as (keyof Row)[];
const cell = (v: string | number) => `"${String(v).replace(/"/g, '""')}"`;
return [cols.join(","), ...rows.map((r) => cols.map((c) => cell(r[c])).join(","))].join("\n") + "\n";
}
async function main(): Promise<void> {
if (!Number.isInteger(idleTimeout) || idleTimeout < 60 || idleTimeout > 604_800) {
throw new Error("--idle-timeout must be a whole number of seconds from 60 to 604800");
}
const rows: Row[] = [];
for (const region of regions) {
try {
rows.push(...(await scanRegion(region)));
} catch (err) {
console.error(`${region}: ${err instanceof Error ? `${err.name}: ${err.message}` : String(err)}`);
}
}
if (rows.length === 0) {
console.log(`No running or waiting EMR clusters in ${regions.join(", ")}.`);
return;
}
console.table(rows);
const idle = rows.filter((r) => r.Verdict === "IDLE" || r.Verdict === "MOSTLY IDLE");
const perHour = idle.reduce((sum, r) => sum + Number(/\$([\d.]+)/.exec(r.PerHour)?.[1] ?? 0), 0);
console.log(`${idle.length} of ${rows.length} clusters idle over the last ${hours} hours: about $${perHour.toFixed(2)} an hour, ` +
`$${(perHour * HOURS_PER_MONTH).toFixed(0)} a month if left running (us-east-1 on-demand prices).`);
if (csvPath) {
writeFileSync(csvPath, toCsv(rows));
console.log(`Wrote ${rows.length} rows to ${csvPath}`);
}
}
main().catch((err) => {
console.error(err);
process.exit(1);
});
How do you run it?
npm install @aws-sdk/client-emr @aws-sdk/client-cloudwatch
npm install --save-dev tsx typescript @types/node
# Report only: last 24 hours, two Regions, CSV for the owners
AWS_PROFILE=readonly npx tsx find-idle-emr-clusters.ts --regions us-east-1,eu-west-1 --csv idle-emr.csv
# Attach a 30-minute auto-termination policy to idle clusters that have none
AWS_PROFILE=emr-admin npx tsx find-idle-emr-clusters.ts --regions us-east-1 --hours 48 --apply --idle-timeout 1800
Sample output
┌─────────┬─────────────┬───────────────────┬────────────────────┬───────────┬──────────────┬────────────────────────────────┬─────────┬──────────────┬──────────────────────────────┬─────────────────┬─────────┬────────────────────────────────────┬────────────────────────────────────────┐
│ (index) │ Region │ ClusterId │ Name │ State │ Release │ Nodes │ IdlePct │ IdleForHours │ LastStep │ AutoTermination │ PerHour │ Verdict │ Action │
├─────────┼─────────────┼───────────────────┼────────────────────┼───────────┼──────────────┼────────────────────────────────┼─────────┼──────────────┼──────────────────────────────┼─────────────────┼─────────┼────────────────────────────────────┼────────────────────────────────────────┤
│ 0 │ 'us-east-1' │ 'j-2AXXXXXXGAPLF' │ 'nightly-etl' │ 'WAITING' │ 'emr-7.5.0' │ '1x m5.xlarge, 4x r5.2xlarge' │ 86 │ 3.3 │ 'COMPLETED 2026-09-28 04:41' │ 'none' │ '$2.76' │ 'MOSTLY IDLE' │ 'would add auto-termination (--apply)' │
│ 1 │ 'us-east-1' │ 'j-3BXXXXXXKQ7Z1' │ 'spark-poc-mlteam' │ 'WAITING' │ 'emr-6.15.0' │ '1x m5.2xlarge, 7x r5.2xlarge' │ 100 │ 24 │ 'COMPLETED 2026-09-02 16:20' │ 'none' │ '$2.87' │ 'IDLE' │ 'would add auto-termination (--apply)' │
│ 2 │ 'us-east-1' │ 'j-1CXXXXXXW3M8P' │ 'adhoc-hive-2023' │ 'WAITING' │ 'emr-5.29.0' │ '3x m5.xlarge' │ 100 │ 23.7 │ 'no steps' │ 'not supported' │ '$0.72' │ 'IDLE' │ '-' │
│ 3 │ 'us-east-1' │ 'j-4DXXXXXXR9TT2' │ 'streaming-ingest' │ 'RUNNING' │ 'emr-7.5.0' │ '1x m5.xlarge, 2x m5.2xlarge' │ 0 │ 0 │ 'RUNNING 2026-09-14 08:00' │ 'none' │ '$1.20' │ 'in use: steps pending or running' │ '-' │
└─────────┴─────────────┴───────────────────┴────────────────────┴───────────┴──────────────┴────────────────────────────────┴─────────┴──────────────┴──────────────────────────────┴─────────────────┴─────────┴────────────────────────────────────┴────────────────────────────────────────┘
3 of 4 clusters idle over the last 24 hours: about $6.35 an hour, $4636 a month if left running (us-east-1 on-demand prices).
IDs, names and numbers are illustrative. spark-poc-mlteam has been idle for the whole window and its last step finished weeks ago; it’s the one to act on. adhoc-hive-2023 runs Amazon EMR 5.29.0, which predates auto-termination, so the script can only report it: ask the owner and terminate it by hand. nightly-etl is mostly idle because it works a few hours a night and waits the rest of the day. That pattern is cheaper as a transient cluster that terminates after its steps, or with a short idle timeout.
Auto-terminate, resize or terminate?
| Option | What happens | Use it when |
|---|---|---|
| Auto-termination policy | EMR terminates the cluster after it has been idle for IdleTimeout (60 seconds to 7 days; 1 hour by default) |
The cluster is still used, but in bursts. You keep it for the next job and stop paying for the gaps after it |
| Terminate now | All instances stop billing; HDFS data on the cluster is gone | Nobody claims it. Check that outputs are in S3, not only in HDFS |
| EMR Serverless | Applications auto-start on job submission and, by default, auto-stop after 15 idle minutes | Spark or Hive jobs are occasional and you’d rather not size clusters at all |
The Amazon EMR auto-termination policy documentation lists what to check before relying on it. It needs release 5.30.0 or 6.1.0 and later (6.14.0 in some newer Regions), it isn’t supported for non-YARN applications such as Presto, Trino or HBase, and the metrics collector must reach the public auto-termination endpoint. An API Gateway interface endpoint with private DNS in the cluster’s VPC breaks it. For SQL on data that’s already in S3, the guide to run an Athena query with AWS SDK v3 is another way to skip a standing cluster.
What should you check before acting?
- Interactive users. SSH sessions and notebook kernels that aren’t running a job don’t count as activity for
IsIdle. Ask in the team channel before a short timeout ends someone’s afternoon. - Data only in HDFS. Terminating deletes it. HDFS utilization in the console or the
HDFSUtilizationmetric tells you whether there’s anything there. - Longer cycles. A weekly job looks idle in a 24-hour window. Re-run with
--hours 168before terminating anything with a schedule-like name. - Owners. Cluster tags tell you whom to ask; the script to find untagged AWS resources lists clusters nobody has claimed.
Idle analytics clusters rarely come alone. The same review usually turns up a warehouse or search domain nobody queries; the scripts to find idle Amazon Redshift clusters and find idle Amazon OpenSearch Service domains cover those, and an AWS budget alert created with SDK v3 warns you when the next forgotten cluster starts adding up. Data migrations can leave compute behind as well; the script to find idle AWS DMS replication instances left after a migration lists replication instances with no running tasks. Streaming ingestion leaves brokers behind too; the script to find idle Amazon MSK clusters flags Kafka clusters with no producers or consumers.
Troubleshooting
- “no IsIdle data: check the cluster age”. The cluster started after the window began or is unreachable. EMR pulls metrics from the cluster, and an unreachable cluster reports nothing until it’s reachable again.
- AutoTermination shows “unknown (…)”.
GetAutoTerminationPolicyfailed; the error name is in brackets. A missingelasticmapreduce:GetAutoTerminationPolicypermission is the usual cause. - PerHour shows “unpriced type”. Add the EMR fee and EC2 price for that type to
PRICE. Prices differ by Region, so adjust them outside us-east-1. AccessDeniedExceptionin one Region only. An SCP may restrict that Region. The steps to troubleshoot AWS IAM access denied errors show how to read the message.
Ask ChatWithCloud instead
For a one-off check, ask ChatWithCloud “Which EMR clusters are in WAITING state, and when did their last step finish?” It writes AWS SDK for JavaScript v2 code, runs it locally with your profile and explains the answer; how ChatWithCloud runs AWS SDK code locally shows each step. It uses one profile and Region per session and runs changes without a confirmation step, so connect ChatWithCloud to a read-only AWS profile and leave termination to a person. More scripts like this are on the AWS practical examples hub.
Frequently asked questions
How do I know if an EMR cluster is idle?
Check the IsIdle CloudWatch metric for the cluster (dimension JobFlowId). A value of 1 for 30 minutes or more, with no pending steps, means no jobs or tasks ran in that time.
Do you pay for an EMR cluster in WAITING state?
Yes. WAITING means the cluster is up and waiting for work, and every running instance bills the EMR fee and the EC2 price until the cluster terminates.
Can I add auto-termination to a running EMR cluster?
Yes, on release 5.30.0 or 6.1.0 and later. Call PutAutoTerminationPolicy or run aws emr put-auto-termination-policy --cluster-id j-XXXX --auto-termination-policy IdleTimeout=3600.
Can I stop an EMR cluster instead of terminating it?
No. The EMR API has actions to terminate a cluster but none to stop and later restart one, the way you can with an EC2 instance. Terminate it and launch a new one when you need it, keeping your data in S3.
Related guides
Ask your AWS account in plain English
Your first 15 runs are free, with no OpenAI key needed.
npx chatwithcloud