Detect and Stop Underutilized EC2 Instances by CPU

Close view of network cables plugged into a server rack with green lights

Photo by Ryutaro Uozumi on Unsplash

To stop underutilized EC2 instances safely, list running instances with DescribeInstances, read each one’s hourly CPUUtilization for the last 14 days with CloudWatch GetMetricData, and flag those whose average and peak stay under your thresholds. Review the list first; then call StopInstances. The script below reports by default and stops only when you pass --stop.

Instances that sit at 1 or 2% CPU all month are some of the easiest AWS savings to find, and some of the riskiest to act on blindly: a quiet instance can still be a license server, a bastion host or a database replica. This example is for engineers who want a repeatable way to find and stop underutilized EC2 instances without surprising anyone. You’ll get a TypeScript script for the AWS SDK for JavaScript v3 that finds idle candidates in a region, explains why it skipped the others, and stops nothing unless you ask it to.

It’s part of our AWS SDK v3 examples for cost and operations. To confirm EC2 is where your money goes before you start, run the script to find your most expensive AWS service with Cost Explorer.

Report-only by default: without --stop the script only reads data. With --stop it calls StopInstances on every instance in the report, with no further prompt. Stopping an instance loses data on any instance store volumes and releases its auto-assigned public IPv4 address.

How does the script decide an instance is underutilized?

  1. Lists running instancespaginateDescribeInstances with the filter instance-state-name = running, in the region from AWS_REGION or your profile.
  2. Skips instances that shouldn’t be stopped this waySpot Instances, instances with an instance-store root volume, members of an Auto Scaling group (the group would replace a stopped instance), instances tagged cwc:keep-running, and anything launched inside the lookback window. Each skip is printed with its reason.
  3. Reads hourly CPUOne GetMetricData query per instance for AWS/EC2 CPUUtilization, with Period: 3600 and the Average statistic, up to 500 queries per request, following NextToken.
  4. Applies two thresholdsAn instance is flagged when the mean of its hourly averages is under --avg (default 5%) and its busiest hour is under --max (default 20%). The peak check keeps nightly batch servers off the list. Instances with fewer than 24 hourly datapoints are skipped.
  5. Stops only on requestWith --stop, it calls StopInstances in batches of 50 and prints each state change.

Prerequisites

  • Node.js 18 or later, npm and tsx.
  • @aws-sdk/client-ec2 and @aws-sdk/client-cloudwatch.
  • A profile with a default region, or AWS_REGION set. The script covers one region per run.
  • Agreement from the instance owners. Tag anything that must keep running with cwc:keep-running before you use --stop.

Which IAM permissions does it need?

The first statement is all report-only mode needs. Add the second only to the profile you’ll use with --stop, and scope it to your account and region (replace 123456789012 and us-east-1).

stop-idle-ec2-policy.json

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "ReportIdleInstances",
      "Effect": "Allow",
      "Action": [
        "ec2:DescribeInstances",
        "cloudwatch:GetMetricData"
      ],
      "Resource": "*"
    },
    {
      "Sid": "StopInstancesInOneRegion",
      "Effect": "Allow",
      "Action": "ec2:StopInstances",
      "Resource": "arn:aws:ec2:us-east-1:123456789012:instance/*"
    }
  ]
}

Keeping the read and write statements separate lets you run the report from a read-only role. The IAM policy generator for TypeScript SDK code is useful if you add calls, and the guide to review an IAM policy for least privilege covers how to tighten the resource ARNs further, for example with tag conditions.

The full script to detect and stop underutilized EC2 instances

stop-idle-ec2.ts

// stop-idle-ec2.ts
// Finds running EC2 instances whose CPUUtilization stayed low for N days.
// Report-only by default. Pass --stop to actually stop the instances it lists.
// Usage: npx tsx stop-idle-ec2.ts [--days 14] [--avg 5] [--max 20] [--stop]
import {
  EC2Client,
  paginateDescribeInstances,
  StopInstancesCommand,
  type Instance,
} from "@aws-sdk/client-ec2";
import {
  CloudWatchClient,
  GetMetricDataCommand,
  type MetricDataQuery,
} from "@aws-sdk/client-cloudwatch";

const SKIP_TAG = "cwc:keep-running"; // tag an instance with this key to exclude it

function numberArg(flag: string, fallback: number): number {
  const i = process.argv.indexOf(flag);
  const value = i === -1 ? fallback : Number(process.argv[i + 1]);
  if (!Number.isFinite(value) || value <= 0) throw new Error(`${flag} must be a positive number`);
  return value;
}

const tag = (i: Instance, key: string) => i.Tags?.find((t) => t.Key === key)?.Value;

// Why an instance must not be stopped by this script, or undefined if it's a candidate.
function skipReason(i: Instance, since: Date): string | undefined {
  if (i.InstanceLifecycle === "spot") return "Spot Instance";
  if (i.RootDeviceType === "instance-store") return "instance-store root volume";
  if (tag(i, "aws:autoscaling:groupName")) return "in an Auto Scaling group";
  if (tag(i, SKIP_TAG) !== undefined) return `tagged ${SKIP_TAG}`;
  if (i.LaunchTime && i.LaunchTime > since) return "launched inside the lookback window";
  return undefined;
}

async function main(): Promise<void> {
  const days = numberArg("--days", 14);
  const avgLimit = numberArg("--avg", 5); // mean of hourly averages, percent
  const maxLimit = numberArg("--max", 20); // highest hourly average, percent
  const stop = process.argv.includes("--stop");

  // No region given: the SDK uses AWS_REGION or the region in your AWS profile.
  const ec2 = new EC2Client({});
  const cloudwatch = new CloudWatchClient({});
  const end = new Date();
  const start = new Date(end.getTime() - days * 86_400_000);

  // 1. Running instances, split into candidates and skipped ones.
  const candidates: Instance[] = [];
  for await (const page of paginateDescribeInstances(
    { client: ec2 },
    { Filters: [{ Name: "instance-state-name", Values: ["running"] }] },
  )) {
    for (const reservation of page.Reservations ?? []) {
      for (const instance of reservation.Instances ?? []) {
        const reason = skipReason(instance, start);
        if (reason) console.log(`skip ${instance.InstanceId}: ${reason}`);
        else candidates.push(instance);
      }
    }
  }
  if (candidates.length === 0) {
    console.log("No running instances to evaluate.");
    return;
  }

  // 2. Hourly average CPUUtilization for each candidate (up to 500 queries per request).
  const values = new Map<string, number[]>();
  for (let i = 0; i < candidates.length; i += 500) {
    const batch = candidates.slice(i, i + 500);
    const queries: MetricDataQuery[] = batch.map((instance, j) => ({
      Id: `q${i + j}`,
      MetricStat: {
        Metric: {
          Namespace: "AWS/EC2",
          MetricName: "CPUUtilization",
          Dimensions: [{ Name: "InstanceId", Value: instance.InstanceId }],
        },
        Period: 3600,
        Stat: "Average",
      },
    }));
    let nextToken: string | undefined;
    do {
      const res = await cloudwatch.send(
        new GetMetricDataCommand({
          StartTime: start,
          EndTime: end,
          MetricDataQueries: queries,
          NextToken: nextToken,
        }),
      );
      for (const result of res.MetricDataResults ?? []) {
        const id = result.Id ?? "";
        values.set(id, [...(values.get(id) ?? []), ...(result.Values ?? [])]);
      }
      nextToken = res.NextToken;
    } while (nextToken);
  }

  // 3. Flag instances whose average and peak hourly CPU are both under the limits.
  const idle: { id: string; name: string; type: string; avg: number; max: number }[] = [];
  candidates.forEach((instance, index) => {
    const points = values.get(`q${index}`) ?? [];
    if (points.length < 24) {
      console.log(`skip ${instance.InstanceId}: only ${points.length} hourly datapoints`);
      return;
    }
    const avg = points.reduce((s, v) => s + v, 0) / points.length;
    const max = Math.max(...points);
    if (avg < avgLimit && max < maxLimit) {
      idle.push({
        id: instance.InstanceId ?? "",
        name: tag(instance, "Name") ?? "",
        type: instance.InstanceType ?? "",
        avg,
        max,
      });
    }
  });

  console.log(`\nUnderutilized over the last ${days} days (avg < ${avgLimit}%, peak hour < ${maxLimit}%):`);
  if (idle.length === 0) {
    console.log("none");
    return;
  }
  console.table(
    idle.map((r) => ({
      Instance: r.id,
      Name: r.name,
      Type: r.type,
      "Avg CPU %": r.avg.toFixed(1),
      "Peak hour %": r.max.toFixed(1),
    })),
  );

  if (!stop) {
    console.log("Report only. Nothing was stopped. Re-run with --stop to stop these instances.");
    return;
  }

  // 4. Stop in batches of 50 and show the state change EC2 reports.
  const ids = idle.map((r) => r.id);
  for (let i = 0; i < ids.length; i += 50) {
    const res = await ec2.send(new StopInstancesCommand({ InstanceIds: ids.slice(i, i + 50) }));
    for (const change of res.StoppingInstances ?? []) {
      console.log(`${change.InstanceId}: ${change.PreviousState?.Name} -> ${change.CurrentState?.Name}`);
    }
  }
}

main().catch((err) => {
  console.error(err);
  process.exit(1);
});

GetMetricData results for a query can be split across pages, so the script appends values per query ID rather than overwriting them. The query IDs q0, q1 and so on map back to positions in the candidates array.

How do you run it?

Terminal

npm install @aws-sdk/client-ec2 @aws-sdk/client-cloudwatch
npm install --save-dev tsx typescript

# 1. Report only (default): nothing is changed
AWS_PROFILE=readonly AWS_REGION=us-east-1 npx tsx stop-idle-ec2.ts

# Stricter thresholds over 30 days
AWS_PROFILE=readonly AWS_REGION=us-east-1 npx tsx stop-idle-ec2.ts --days 30 --avg 2 --max 10

# 2. After reviewing the report: actually stop the listed instances
AWS_PROFILE=ops AWS_REGION=us-east-1 npx tsx stop-idle-ec2.ts --stop

Sample output

Output (report-only run)

skip i-0a1b2c3d4e5f60718: in an Auto Scaling group
skip i-0f9e8d7c6b5a40312: tagged cwc:keep-running
skip i-03c4d5e6f7a8b9012: launched inside the lookback window

Underutilized over the last 14 days (avg < 5%, peak hour < 20%):
┌─────────┬───────────────────────┬──────────────────┬─────────────┬───────────┬─────────────┐
│ (index) │ Instance              │ Name             │ Type        │ Avg CPU % │ Peak hour % │
├─────────┼───────────────────────┼──────────────────┼─────────────┼───────────┼─────────────┤
│ 0       │ 'i-0123456789abcdef0' │ 'staging-api-2'  │ 'm5.large'  │ '1.8'     │ '6.4'       │
│ 1       │ 'i-0fedcba9876543210' │ 'old-jenkins'    │ 't3.medium' │ '0.6'     │ '3.1'       │
└─────────┴───────────────────────┴──────────────────┴─────────────┴───────────┴─────────────┘
Report only. Nothing was stopped. Re-run with --stop to stop these instances.

Instance IDs and figures are illustrative. With --stop, the same run ends with lines such as i-0123456789abcdef0: running -> stopping.

How much can you save by stopping idle EC2 instances?

A stopped instance isn’t charged for instance hours, but its EBS volumes, snapshots and any Elastic IP addresses are still billed. Using On-Demand Linux prices in US East (N. Virginia) as of September 2026, and AWS’s convention of 730 hours in a month:

Instance type On-Demand price per hour Arithmetic Per month if left running
t3.medium $0.0416 0.0416 × 730 $30.37
m5.large $0.096 0.096 × 730 $70.08

Stopping both instances from the sample report saves about $100.45 a month in compute, before storage. Prices differ by region, operating system and purchase option; check the Amazon EC2 On-Demand pricing page for your own instance types. Instances that must keep running but can survive an interruption may be cheaper on Spot; the script to compare EC2 Spot price history across zones shows the discount. If the instances are covered by Reserved Instances or a Savings Plan, stopping them doesn’t reduce that commitment. Once instances are stopped, their volumes are the next thing to review: the sibling script to find unattached EBS volumes and tag them for cleanup handles volumes left behind after termination. If the instances sat behind a load balancer, also find unused load balancers with no healthy targets: an idle load balancer still bills by the hour. Instances that stay stopped for weeks still pay for storage, and the script to find EC2 instances stopped for weeks and still costing you puts a monthly figure on each.

Troubleshooting and limits

  • CPU isn’t the whole picture. CloudWatch reports CPU for every instance, but memory and disk usage need the CloudWatch agent. A memory-bound cache can look idle by CPU.
  • Burstable instances. T-family instances are designed to idle most of the time, so a low average may be normal; the peak threshold matters more for them.
  • UnauthorizedOperation on StopInstances. The profile lacks ec2:StopInstances for that instance ARN. The steps to debug AWS IAM access denied and unauthorized errors apply here too.
  • UnsupportedOperation. Some instances can’t be stopped, such as those with an instance-store root volume; the script skips those it can detect.
  • No instances appear. Check the region. The script covers one region per run; loop over regions with AWS_REGION if you need more.

Ask ChatWithCloud instead

You can ask ChatWithCloud “Which EC2 instances averaged under 5% CPU over the last two weeks?” and get the list without writing the script. It generates AWS SDK for JavaScript v2 code, runs it on your machine with your profile and explains the result; the guide to list AWS resources with natural language from the terminal shows similar inventory questions. Be careful with the stopping step: ChatWithCloud runs changes without a confirmation step, so asking it to stop instances will stop them. Ask for the report with a read-only AWS profile connected to ChatWithCloud, then stop instances yourself. The ChatWithCloud security page explains what runs locally and what’s sent to the model.

Frequently asked questions

What CPU percentage counts as an underutilized EC2 instance?

There’s no single rule. Averages under 5% with no hour above 20% over two weeks are a conservative starting point; tighten or loosen them with --avg and --max for your workloads.

Do I still pay for a stopped EC2 instance?

Not for instance hours, but EBS volumes, snapshots and Elastic IP addresses keep billing while the instance is stopped.

Is it safe to stop an EC2 instance?

For EBS-backed instances the root volume and data volumes are kept, but instance store data is lost and the auto-assigned public IP changes on restart. Check with the owner first, and never stop Auto Scaling group members directly. If you can’t reach it over SSH after it starts again, work through how to troubleshoot why you can’t SSH into an EC2 instance.

Should I stop or downsize an underutilized instance?

Stop it if nobody needs it running. If it serves real traffic at low load, a smaller instance type is usually the better fix. If it runs on an old family such as M4 or C4, the script to find previous-generation EC2 instances to upgrade suggests a current type at the same time.

Related guides

Ask your AWS account in plain English

Your first 15 runs are free, with no OpenAI key needed.

npx chatwithcloud