Photo by Rafael Pol on Unsplash
To find idle SageMaker endpoints, list in-service endpoints with ListEndpoints, read each variant’s instance type and count with DescribeEndpoint and DescribeEndpointConfig, then sum the CloudWatch Invocations metric per variant over 14 days. Instance-backed endpoints with zero invocations bill every hour. Deleting one keeps its model and endpoint configuration.
A real-time SageMaker endpoint bills per instance-hour from the moment it’s in service, whether it serves a million requests or none. Endpoints get created for a demo, a load test or a model comparison, and then outlive the project. A single GPU instance left running costs more per month than many small production workloads. Notebook instances started for experiments follow the same pattern, and the script to find and stop idle SageMaker notebook instances catches those.
This example is for teams who run SageMaker AI inference and want to find idle SageMaker endpoints before the bill does. The script reports every in-service endpoint, what it runs on, how many invocations it served and an on-demand monthly estimate. It deletes only when you pass --apply, like the other scripts in the AWS SDK v3 practical examples.
Which SageMaker endpoints cost money while idle?
| Endpoint type | Billed while idle? | Activity metric the script reads |
|---|---|---|
| Real-time, instance-backed | Yes, every instance-hour | Invocations (Sum) per EndpointName and VariantName |
| Asynchronous | Yes, unless it scales to zero instances | InvocationsProcesssed, as the SageMaker docs spell it; async endpoints don’t publish Invocations |
| Serverless | No compute charge without requests, except provisioned concurrency | Invocations; reported but never deleted by the script |
| Inference components | Yes, unless the endpoint scales to zero instances | Invocations are per component; the script marks these for a manual check |
The script lists only endpoints in the InService status and uses the live CurrentInstanceCount, so autoscaled endpoints are priced at their current size, not their initial one.
What does an idle endpoint cost?
On-demand real-time inference prices in US East (N. Virginia) as of September 2026, from the AWS Price List API and the Amazon SageMaker AI pricing page. Prices vary by Region and change, so check the page before you quote them.
| Instance type | Per hour | Per month (730 hours) |
|---|---|---|
| ml.t2.medium | $0.056 | $40.88 |
| ml.m5.large | $0.115 | $83.95 |
| ml.c5.xlarge | $0.204 | $148.92 |
| ml.m5.xlarge | $0.23 | $167.90 |
| ml.g6.xlarge | $1.1267 | $822.49 |
| ml.g5.xlarge | $1.408 | $1,027.84 |
Worked example: a forgotten A/B test endpoint with two variants, each on one ml.m5.xlarge, costs 2 × $0.23 × 730 = $335.80 a month. One idle ml.g5.xlarge endpoint for a demo costs 1 × $1.408 × 730 = $1,027.84 a month. The script does the same arithmetic with live prices from pricing:GetProducts; pass --no-price to skip that call.
What does the script do?
- Lists endpoints
paginateListEndpointswithStatusEquals: InServicein each Region you pass. - Reads the configuration
DescribeEndpointfor live instance counts and serverless settings,DescribeEndpointConfigfor instance types, async settings and whether variants host inference components. - Sums invocationsOne
GetMetricDatacall per endpoint, dailySumover--days(default 14) for each variant. - Prices the idle capacityHourly on-demand price × instance count × 730, cached per instance type.
- Deletes only on requestWith
--apply,DeleteEndpointfor endpoints markedIDLE. Serverless and inference-component endpoints are never deleted.
Prerequisites
- Node.js 18 or later, npm,
tsx, and@aws-sdk/client-sagemaker,@aws-sdk/client-cloudwatchand@aws-sdk/client-pricing. - A read-only profile for the report; the guide to AWS SDK v3 credential providers such as fromIni and fromSSO covers profile setup.
- A list of endpoints that are idle on purpose, such as a disaster-recovery copy or a model that serves one monthly batch.
Which IAM permissions does it need?
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ReadEndpointsMetricsPrices",
"Effect": "Allow",
"Action": [
"sagemaker:ListEndpoints",
"cloudwatch:GetMetricData",
"pricing:GetProducts"
],
"Resource": "*"
},
{
"Sid": "DescribeEndpoints",
"Effect": "Allow",
"Action": [
"sagemaker:DescribeEndpoint",
"sagemaker:DescribeEndpointConfig"
],
"Resource": [
"arn:aws:sagemaker:*:123456789012:endpoint/*",
"arn:aws:sagemaker:*:123456789012:endpoint-config/*"
]
},
{
"Sid": "DeleteWithApply",
"Effect": "Allow",
"Action": "sagemaker:DeleteEndpoint",
"Resource": "arn:aws:sagemaker:*:123456789012:endpoint/*"
}
]
}
Remove the last statement for a report-only role. The free IAM policy generator for TypeScript code can double-check the list against the script.
The script to find idle SageMaker endpoints
// find-idle-sagemaker-endpoints.ts
// Finds SageMaker AI endpoints with no invocations in the last --days days, shows what each variant
// runs on (instances, serverless or async) and estimates the monthly instance cost from the AWS Price
// List API. Report only unless you pass --apply, which deletes idle instance-backed endpoints
// (the endpoint configuration and model are kept, so you can recreate the endpoint later).
// Usage:
// npx tsx find-idle-sagemaker-endpoints.ts [--regions us-east-1,eu-west-1] [--days 14] [--no-price] [--apply]
import { CloudWatchClient, GetMetricDataCommand, type MetricDataQuery } from "@aws-sdk/client-cloudwatch";
import { GetProductsCommand, PricingClient } from "@aws-sdk/client-pricing";
import { DeleteEndpointCommand, DescribeEndpointCommand, DescribeEndpointConfigCommand, SageMakerClient, paginateListEndpoints } from "@aws-sdk/client-sagemaker";
const args = process.argv.slice(2);
const flag = (name: string): string | undefined => {
const i = args.indexOf(name);
return i >= 0 ? args[i + 1] : undefined;
};
const regions = (flag("--regions") ?? process.env.AWS_REGION ?? "us-east-1").split(",").map((r) => r.trim()).filter(Boolean);
const days = Number(flag("--days") ?? "14");
const withPrice = !args.includes("--no-price");
const apply = args.includes("--apply");
const HOURS_PER_MONTH = 730;
interface Row {
Region: string;
Endpoint: string;
Mode: string;
Capacity: string;
Invocations: number;
UsdPerMonth: string;
Status: string;
}
// On-demand hosting price per instance-hour from the Price List API (its endpoint lives in us-east-1).
const pricing = new PricingClient({ region: "us-east-1" });
const priceCache = new Map<string, number | undefined>();
async function hourlyPrice(region: string, instanceType: string): Promise<number | undefined> {
const key = `${region}/${instanceType}`;
if (priceCache.has(key)) return priceCache.get(key);
let price: number | undefined;
try {
const { PriceList } = await pricing.send(
new GetProductsCommand({
ServiceCode: "AmazonSageMaker",
Filters: [
{ Type: "TERM_MATCH", Field: "regionCode", Value: region },
{ Type: "TERM_MATCH", Field: "instanceType", Value: `${instanceType}-Hosting` },
],
MaxResults: 10,
}),
);
for (const item of Array.isArray(PriceList) ? PriceList : []) {
const product = JSON.parse(item) as {
terms?: { OnDemand?: Record<string, { priceDimensions?: Record<string, { unit?: string; pricePerUnit?: { USD?: string } }> }> };
};
for (const term of Object.values(product.terms?.OnDemand ?? {})) {
for (const dim of Object.values(term.priceDimensions ?? {})) {
const usd = Number(dim.pricePerUnit?.USD);
if (dim.unit === "Hrs" && usd > 0) price = usd;
}
}
if (price !== undefined) break;
}
} catch {
price = undefined; // no pricing:GetProducts permission or no match: leave the cost column empty
}
priceCache.set(key, price);
return price;
}
async function scanRegion(region: string): Promise<Row[]> {
const sm = new SageMakerClient({ region });
const cw = new CloudWatchClient({ region });
const end = new Date();
const start = new Date(end.getTime() - days * 86_400_000);
const rows: Row[] = [];
for await (const page of paginateListEndpoints({ client: sm }, { StatusEquals: "InService" })) {
for (const e of page.Endpoints ?? []) {
const name = e.EndpointName ?? "";
const ep = await sm.send(new DescribeEndpointCommand({ EndpointName: name }));
const cfg = await sm.send(new DescribeEndpointConfigCommand({ EndpointConfigName: ep.EndpointConfigName }));
const isAsync = cfg.AsyncInferenceConfig !== undefined;
const usesComponents = (cfg.ProductionVariants ?? []).some((v) => !v.ModelName && !v.ServerlessConfig);
// Async endpoints don't publish Invocations; they publish InvocationsProcesssed (spelled with three s
// in the SageMaker docs), so both spellings are summed.
const metricNames = isAsync ? ["InvocationsProcesssed", "InvocationsProcessed"] : ["Invocations"];
const queries: MetricDataQuery[] = [];
for (const [vi, v] of (cfg.ProductionVariants ?? []).entries()) {
for (const [mi, metric] of metricNames.entries()) {
queries.push({
Id: `v${vi}m${mi}`,
MetricStat: {
Metric: {
Namespace: "AWS/SageMaker",
MetricName: metric,
Dimensions: [
{ Name: "EndpointName", Value: name },
{ Name: "VariantName", Value: v.VariantName },
],
},
Period: 86_400,
Stat: "Sum",
},
});
}
}
let invocations = 0;
if (queries.length) {
const data = await cw.send(new GetMetricDataCommand({ StartTime: start, EndTime: end, MetricDataQueries: queries }));
for (const r of data.MetricDataResults ?? []) invocations += (r.Values ?? []).reduce((a, b) => a + b, 0);
}
// What is billed while idle: instance-hours for real-time and async variants, provisioned
// concurrency for serverless, nothing else for on-demand serverless.
let hourly = 0;
let priced = true;
const capacity: string[] = [];
for (const v of cfg.ProductionVariants ?? []) {
const live = ep.ProductionVariants?.find((s) => s.VariantName === v.VariantName);
if (v.ServerlessConfig) {
const pc = live?.CurrentServerlessConfig?.ProvisionedConcurrency ?? v.ServerlessConfig.ProvisionedConcurrency ?? 0;
capacity.push(`${v.VariantName}: serverless ${v.ServerlessConfig.MemorySizeInMB} MB${pc ? `, ${pc} provisioned` : ""}`);
if (pc) priced = false;
continue;
}
const count = live?.CurrentInstanceCount ?? v.InitialInstanceCount ?? 0;
const type = v.InstanceType ?? "?";
capacity.push(`${v.VariantName}: ${count} x ${type}`);
const price = withPrice && count > 0 ? await hourlyPrice(region, type) : undefined;
if (price === undefined) priced = priced && count === 0;
else hourly += price * count;
}
const mode = isAsync ? "async" : usesComponents ? "inference components" : capacity.some((c) => c.includes("serverless")) ? "serverless" : "real-time";
let status = invocations > 0 ? "in use" : "IDLE";
if (status === "IDLE" && usesComponents) status = "IDLE? check per-component metrics";
if (status === "IDLE" && mode === "serverless" && priced && hourly === 0) status = "idle, no compute charge";
rows.push({
Region: region,
Endpoint: name,
Mode: mode,
Capacity: capacity.join("; "),
Invocations: invocations,
UsdPerMonth: hourly > 0 ? (hourly * HOURS_PER_MONTH).toFixed(2) + (priced ? "" : "+") : priced ? "0.00" : "?",
Status: status,
});
}
}
if (apply) {
for (const r of rows) {
if (r.Status !== "IDLE" || r.Mode === "serverless") continue;
try {
await sm.send(new DeleteEndpointCommand({ EndpointName: r.Endpoint }));
r.Status = "DELETED (config and model kept)";
} catch (err) {
r.Status = `error: ${err instanceof Error ? err.name : String(err)}`;
}
}
}
return rows;
}
async function main(): Promise<void> {
const rows: Row[] = [];
for (const region of regions) {
try {
rows.push(...(await scanRegion(region)));
} catch (err) {
console.error(`${region}: ${err instanceof Error ? `${err.name}: ${err.message}` : String(err)}`);
}
}
if (rows.length) console.table(rows);
const idle = rows.filter((r) => r.Status.startsWith("IDLE") || r.Status.startsWith("DELETED"));
const monthly = idle.reduce((sum, r) => sum + (Number.parseFloat(r.UsdPerMonth) || 0), 0);
console.log(`${rows.length} in-service endpoints, ${idle.length} with no invocations in ${days} days (about $${monthly.toFixed(2)}/month on demand)`);
if (!apply) console.log("Report only. Re-run with --apply to delete idle instance-backed endpoints.");
}
main().catch((err) => {
console.error(err);
process.exit(1);
});
How do you run it?
npm install @aws-sdk/client-sagemaker @aws-sdk/client-cloudwatch @aws-sdk/client-pricing
npm install --save-dev tsx typescript @types/node
# Report on two Regions, 30-day window
AWS_PROFILE=readonly npx tsx find-idle-sagemaker-endpoints.ts --regions us-east-1,us-west-2 --days 30
# Delete what is idle after review
AWS_PROFILE=ml-admin npx tsx find-idle-sagemaker-endpoints.ts --regions us-east-1 --days 30 --apply
Sample output
┌─────────┬─────────────┬──────────────────┬──────────────┬───────────────────────────────────────────────────────────┬─────────────┬─────────────┬───────────────────────────┐
│ (index) │ Region │ Endpoint │ Mode │ Capacity │ Invocations │ UsdPerMonth │ Status │
├─────────┼─────────────┼──────────────────┼──────────────┼───────────────────────────────────────────────────────────┼─────────────┼─────────────┼───────────────────────────┤
│ 0 │ 'us-east-1' │ 'churn-ab-test' │ 'real-time' │ 'control: 1 x ml.m5.xlarge; challenger: 1 x ml.m5.xlarge' │ 0 │ '335.80' │ 'IDLE' │
│ 1 │ 'us-east-1' │ 'llm-demo' │ 'real-time' │ 'AllTraffic: 1 x ml.g5.xlarge' │ 0 │ '1027.84' │ 'IDLE' │
│ 2 │ 'us-east-1' │ 'fraud-scoring' │ 'real-time' │ 'AllTraffic: 2 x ml.c5.xlarge' │ 184223 │ '297.84' │ 'in use' │
│ 3 │ 'us-west-2' │ 'doc-classifier' │ 'serverless' │ 'AllTraffic: serverless 2048 MB' │ 0 │ '0.00' │ 'idle, no compute charge' │
└─────────┴─────────────┴──────────────────┴──────────────┴───────────────────────────────────────────────────────────┴─────────────┴─────────────┴───────────────────────────┘
4 in-service endpoints, 2 with no invocations in 30 days (about $1363.64/month on demand)
Report only. Re-run with --apply to delete idle instance-backed endpoints.
Names are illustrative. churn-ab-test and llm-demo haven’t served a request in 30 days and cost about $1,364 a month together. doc-classifier is idle too, but a serverless endpoint without provisioned concurrency costs nothing for compute, so it isn’t counted or deleted.
What should you check before deleting an endpoint?
- Who calls it. Search your Lambda functions, API Gateway integrations and application config for the endpoint name. Batch jobs that run monthly look idle over 14 days; use
--days 45for those. - That you can recreate it.
DeleteEndpointkeeps the endpoint configuration and the SageMaker model, and deleting a model never removes its artifacts in S3. Recreate withaws sagemaker create-endpoint --endpoint-name churn-ab-test --endpoint-config-name churn-ab-test-config. Don’t delete the endpoint configuration while the endpoint still exists. - Whether it should scale instead. Endpoints built on inference components can scale to zero instances with managed instance scaling, and async endpoints can scale in to zero. Both keep the endpoint but stop the hourly instance charge; after scaling in to zero, requests fail for several minutes while capacity is provisioned again.
- The execution role. SageMaker needs the model’s execution role to clean up after
DeleteEndpoint, so don’t remove that role first.
Idle endpoints are rarely the only forgotten resource. The same pattern works for databases, data processing clusters and networks: find idle RDS instances with no connections, find idle ElastiCache clusters, find idle Amazon EMR clusters with the IsIdle metric and find idle NAT gateways costing you money. To see how much SageMaker costs in total, get last month’s AWS cost broken down by service, and to catch the next forgotten endpoint early, create an AWS budget alert with AWS SDK v3.
Troubleshooting
- An endpoint in use shows zero invocations. Clients may call an endpoint of the same name in another Region, or traffic is lower than one request per
--dayswindow. For async endpoints the script sums both spellings of the processed-invocations metric. - The cost column shows
?. The Price List API returned no match orpricing:GetProductswas denied. The instance count and type are still correct. ThrottlingExceptionon large accounts. The SDK retries with backoff; the guide to configure retry and timeout settings in AWS SDK for JavaScript v3 shows how to raisemaxAttempts.- Delete fails with a validation error. The endpoint is being created or updated. Wait until it’s
InServiceand run again.
Ask ChatWithCloud instead
For a quick look, ask ChatWithCloud “Which SageMaker endpoints in us-east-1 had no invocations in the last 14 days?” It writes AWS SDK for JavaScript v2 code, runs it locally with your AWS profile and explains the answer; how ChatWithCloud works with your AWS profile covers what runs where. It can be wrong, and it runs changes without a confirmation step, so connect ChatWithCloud through a read-only profile and delete with the script. The guide to ask AI why your AWS bill increased helps when the question starts from the bill.
Frequently asked questions
Do SageMaker endpoints charge when not in use?
Real-time instance-backed endpoints charge for every instance-hour while in service, with or without traffic. On-demand serverless endpoints charge for compute only while processing requests.
Does deleting a SageMaker endpoint delete the model?
No. DeleteEndpoint keeps the endpoint configuration and the model, and the model artifacts in S3 stay too. You can recreate the endpoint from the same configuration.
Can you stop a SageMaker endpoint instead of deleting it?
There is no stop action. You delete it and recreate it later, or use an endpoint type that scales to zero instances, such as inference components or async inference.
How do I see SageMaker endpoint invocations?
In CloudWatch, namespace AWS/SageMaker, metric Invocations with the EndpointName and VariantName dimensions, using the Sum statistic.
Related guides
Ask your AWS account in plain English
Your first 15 runs are free, with no OpenAI key needed.
npx chatwithcloud