Photo by Sebastian Herrmann on Unsplash
ECS RunTask with AWS SDK v3 is RunTaskCommand from @aws-sdk/client-ecs. Pass the cluster, taskDefinition, launchType: "FARGATE" (or a capacity provider strategy) and an awsvpcConfiguration with subnets and security groups. Check the failures array, wait with waitUntilTasksStopped, then read the container’s exitCode from DescribeTasks.
Database migrations, a nightly report, a one-time backfill: jobs that should run once, in the same image and network as your service, and then stop. RunTask starts a standalone task outside any service, so nothing restarts it when it exits. This guide is for Node.js and TypeScript developers using ECS RunTask with AWS SDK v3 from CI or a script, who need more than “the task started”: whether it succeeded, what it printed, and the least IAM that allows it. For commands that must run on existing EC2 instances rather than in a new container, SSM SendCommand with AWS SDK v3 is the equivalent.
Every sample was type-checked with strict tsc against @aws-sdk/client-ecs and @aws-sdk/client-cloudwatch-logs (SDK 3.1142.0) in September 2026, and run against aws-sdk-client-mock: a successful task, a failed exit code on Fargate Spot, and a placement failure.
Which RunTask parameters matter for a one-off task?
| Parameter | What to pass | Why it matters |
|---|---|---|
taskDefinition |
family:revision or ARN |
Without a revision, the latest ACTIVE revision runs. Pin it in CI. |
launchType / capacityProviderStrategy |
FARGATE, or [{ capacityProvider: "FARGATE_SPOT", weight: 1 }] |
Mutually exclusive. Fargate Spot requires a strategy. |
networkConfiguration.awsvpcConfiguration |
subnets, securityGroups, assignPublicIp |
Required for awsvpc task definitions, which Fargate uses. |
overrides.containerOverrides |
name, command, environment |
Runs a different command in the same image. All overrides together are limited to 8,192 characters of JSON. |
count |
1 to 10 | Up to 10 tasks per call. |
clientToken |
A unique string, up to 64 characters | Makes the request idempotent, so a retry doesn’t start a second task. |
startedBy |
Up to 128 characters | Lets ListTasks find every task from one job. |
All of these come from the RunTask API reference. If the subnets are private, the task needs a NAT gateway or VPC endpoints to pull its image; assignPublicIp: "ENABLED" is only for public subnets.
Prerequisites
- A cluster and a Fargate task definition whose container uses the
awslogslog driver withawslogs-groupandawslogs-stream-prefixset. - Node.js 18 or later with
@aws-sdk/client-ecsand@aws-sdk/client-cloudwatch-logs. - Credentials for a role with the policy below. The guide to AWS SDK v3 credential providers covers CI roles and profiles.
How to call ECS RunTask with AWS SDK v3, step by step
- Build the inputTask definition, cluster, network settings and a container override with the command to run.
- Send
RunTaskCommandWith a freshclientTokenfromcrypto.randomUUID(). - Check
failuresA placement problem comes back in the response with HTTP 200, not as an exception. - Wait for
STOPPEDwaitUntilTasksStoppedpollsDescribeTasksuntil every task’slastStatusisSTOPPED. - Read the result
DescribeTasksonce more forstopCode,stoppedReasonand the container’sexitCode. - Fetch the logs
FilterLogEventson the streamprefix/container-name/task-id.
Example: a reusable runOneOffTask module
// run-task.ts: run a one-off ECS task on Fargate, wait for it to stop, return its exit code and logs.
import { randomUUID } from "node:crypto";
import {
ECSClient,
RunTaskCommand,
DescribeTasksCommand,
waitUntilTasksStopped,
type Failure,
type RunTaskCommandInput,
} from "@aws-sdk/client-ecs";
import { CloudWatchLogsClient, paginateFilterLogEvents } from "@aws-sdk/client-cloudwatch-logs";
export interface OneOffTask {
cluster: string;
taskDefinition: string; // "family:revision", or "family" for the latest ACTIVE revision
container: string; // the container whose command and exit code you care about
command?: string[];
environment?: Record<string, string>;
subnets: string[];
securityGroups: string[];
assignPublicIp?: boolean; // only for public subnets without a NAT gateway
spot?: boolean; // FARGATE_SPOT through a capacity provider strategy
startedBy?: string;
maxWaitSeconds?: number;
logGroup?: string; // awslogs-group of the container
logStreamPrefix?: string; // awslogs-stream-prefix of the container
}
export interface TaskResult {
taskArn: string;
exitCode?: number;
stopCode?: string;
stoppedReason?: string;
containerReason?: string;
logs: string[];
}
export class RunTaskFailed extends Error {
readonly failures: Failure[];
constructor(failures: Failure[]) {
super(`RunTask placed no task: ${failures.map((f) => `${f.reason ?? "unknown"} ${f.detail ?? ""} ${f.arn ?? ""}`.trim()).join("; ") || "no tasks returned"}`);
this.name = "RunTaskFailed";
this.failures = failures;
}
}
const ecs = new ECSClient({});
const logs = new CloudWatchLogsClient({});
export async function runOneOffTask(t: OneOffTask, waiterMinDelay = 6): Promise<TaskResult> {
const input: RunTaskCommandInput = {
cluster: t.cluster,
taskDefinition: t.taskDefinition,
count: 1,
clientToken: randomUUID(), // makes a retried request return the same task instead of starting a second one
startedBy: t.startedBy ?? "one-off",
networkConfiguration: {
awsvpcConfiguration: { subnets: t.subnets, securityGroups: t.securityGroups, assignPublicIp: t.assignPublicIp ? "ENABLED" : "DISABLED" },
},
overrides: {
containerOverrides: [
{
name: t.container,
command: t.command,
environment: Object.entries(t.environment ?? {}).map(([name, value]) => ({ name, value })),
},
],
},
};
// A task uses either a launch type or a capacity provider strategy, never both.
if (t.spot) input.capacityProviderStrategy = [{ capacityProvider: "FARGATE_SPOT", weight: 1 }];
else input.launchType = "FARGATE";
const run = await ecs.send(new RunTaskCommand(input));
const taskArn = run.tasks?.[0]?.taskArn;
if (!taskArn || run.failures?.length) throw new RunTaskFailed(run.failures ?? []);
await waitUntilTasksStopped(
{ client: ecs, maxWaitTime: t.maxWaitSeconds ?? 3600, minDelay: waiterMinDelay },
{ cluster: t.cluster, tasks: [taskArn] },
);
const described = await ecs.send(new DescribeTasksCommand({ cluster: t.cluster, tasks: [taskArn] }));
const task = described.tasks?.[0];
const container = task?.containers?.find((c) => c.name === t.container);
const result: TaskResult = {
taskArn,
exitCode: container?.exitCode,
stopCode: task?.stopCode,
stoppedReason: task?.stoppedReason,
containerReason: container?.reason,
logs: [],
};
if (t.logGroup && t.logStreamPrefix) {
// awslogs names the stream prefix/container-name/task-id
const taskId = taskArn.split("/").pop() ?? "";
const stream = `${t.logStreamPrefix}/${t.container}/${taskId}`;
for await (const page of paginateFilterLogEvents({ client: logs }, { logGroupName: t.logGroup, logStreamNames: [stream] })) {
for (const e of page.events ?? []) result.logs.push(e.message ?? "");
}
}
return result;
}
The waiter’s defaults for this operation are a 6-second minimum and 600-second maximum delay between polls; maxWaitTime is the overall limit. The stream name follows the awslogs-stream-prefix format that the ECS LogConfiguration reference documents: prefix, container name, task ID. The SDK retries throttled or failed calls on its own, which is exactly why the clientToken is there; the AWS SDK v3 retries and timeouts guide shows how to tune that.
Example: run database migrations from CI
// run-migration.ts: run database migrations as a one-off ECS task from CI and fail the job if they fail.
// Usage: CLUSTER=prod SUBNETS=subnet-0a,subnet-0b SECURITY_GROUP=sg-0abc npx tsx run-migration.ts
import { runOneOffTask, RunTaskFailed } from "./run-task.ts";
const env = (name: string): string => {
const v = process.env[name];
if (!v) throw new Error(`${name} is not set`);
return v;
};
try {
const result = await runOneOffTask({
cluster: env("CLUSTER"),
taskDefinition: process.env.TASK_DEFINITION ?? "api",
container: "app",
command: ["npm", "run", "migrate"],
environment: { LOG_LEVEL: "info" },
subnets: env("SUBNETS").split(","),
securityGroups: [env("SECURITY_GROUP")],
startedBy: `ci-${process.env.GITHUB_RUN_ID ?? "local"}`,
maxWaitSeconds: 1800,
logGroup: "/ecs/api",
logStreamPrefix: "api",
});
for (const line of result.logs) console.log(line);
console.log(`task ${result.taskArn} stopped (${result.stopCode}), exit code ${result.exitCode}`);
if (result.exitCode !== 0) console.error("stopped reason:", result.stoppedReason, "| container reason:", result.containerReason);
process.exit(result.exitCode === 0 ? 0 : 1);
} catch (err) {
if (err instanceof RunTaskFailed) console.error(err.message, err.failures);
else console.error(err instanceof Error ? `${err.name}: ${err.message}` : err);
process.exit(1);
}
A run against mocked ECS and CloudWatch Logs clients prints:
> [email protected] migrate
Applied 20260928_add_invoice_index
Applied 20260929_backfill_currency
2 migrations applied
task arn:aws:ecs:us-east-1:123456789012:task/prod/9f3c1e7a52b84d0f stopped (EssentialContainerExited), exit code 0
The process exit code mirrors the container’s, so the CI step fails when the migration fails. If the job needs several steps with retries between them, a state machine fits better; starting a Step Functions execution with SDK v3 covers that. For jobs on a timetable, creating an EventBridge Scheduler schedule with SDK v3 can start the task without any code of yours running.
What do failures, stopCode and exitCode tell you?
failuresin the RunTask response. Each entry has areason,detailandarn. The ECS guide’s list of API failure reasons covers values such asRESOURCE:*,ATTRIBUTEandLOCATION, which mostly apply to EC2 capacity. Treat any entry, or an emptytasksarray, as “nothing started”.stopCode. SDK v3 types it asTaskFailedToStart,EssentialContainerExited,UserInitiated,ServiceSchedulerInitiated,SpotInterruption,TerminationNoticeorInfrastructureHealth.EssentialContainerExitedis the normal end of a one-off job.exitCode. The container process’s exit status. It’s missing if the container never started, for example when the image pull failed;stoppedReasonand the container’sreasonsay why.SpotInterruption. Fargate Spot capacity was reclaimed. Rerun the job, or use On-Demand for work that can’t be restarted. The report to find ECS services that could use Fargate Spot discusses which workloads suit it.
Stopped tasks stay visible to DescribeTasks for a limited time (the ECS guide says at least an hour), so read the result right away and keep the logs.
Permissions needed
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "RunApiTaskInProdCluster",
"Effect": "Allow",
"Action": "ecs:RunTask",
"Resource": "arn:aws:ecs:us-east-1:123456789012:task-definition/api:*",
"Condition": { "ArnEquals": { "ecs:cluster": "arn:aws:ecs:us-east-1:123456789012:cluster/prod" } }
},
{
"Sid": "WatchTasks",
"Effect": "Allow",
"Action": "ecs:DescribeTasks",
"Resource": "arn:aws:ecs:us-east-1:123456789012:task/prod/*"
},
{
"Sid": "PassTaskRoles",
"Effect": "Allow",
"Action": "iam:PassRole",
"Resource": [
"arn:aws:iam::123456789012:role/api-task-execution",
"arn:aws:iam::123456789012:role/api-task"
],
"Condition": { "StringEquals": { "iam:PassedToService": "ecs-tasks.amazonaws.com" } }
},
{
"Sid": "ReadTaskLogs",
"Effect": "Allow",
"Action": "logs:FilterLogEvents",
"Resource": "arn:aws:logs:us-east-1:123456789012:log-group:/ecs/api"
}
]
}
iam:PassRole is the one people forget: running a task hands its execution role and task role to ECS, so the caller must be allowed to pass exactly those roles. If you add tags or propagateTags, expect to grant ecs:TagResource as well, because ECS checks tag-on-create permissions. To derive actions from your own code, see finding the IAM actions in AWS SDK for JavaScript code or the IAM policy generator for TypeScript code. Before granting RunTask on a task definition, check what it can do; the audits to find privileged ECS task definitions and find secrets in ECS task definitions cover the usual risks.
Troubleshooting and common mistakes
- RunTask rejects the network configuration. A Fargate task definition uses
awsvpc, which needsnetworkConfiguration. Passing it for abridgeorhosttask definition is also an error. - Both
launchTypeandcapacityProviderStrategyset. The API accepts one or the other. With neither, the cluster’s default capacity provider strategy is used. - The task stops with no exit code. It never started: check
stoppedReasonfor image pull or secret retrieval problems, which usually point at the execution role or at private subnets without a route to ECR. - The waiter times out. The job runs longer than
maxWaitSeconds, or it hangs. Stop it withStopTaskCommand(andecs:StopTask) so it doesn’t run forever. DescribeTasksreturnsMISSINGright after RunTask. The ECS API is eventually consistent. The waiter keeps polling, which handles it.- Access denied when ECS receives the task roles. Add the role ARNs to the
PassTaskRolesstatement, then follow troubleshooting AWS IAM access denied errors.
How do you test RunTask code?
Mock the clients, not AWS. This test covers a successful run with paged logs, a failing exit code on Fargate Spot, and a placement failure:
// run-task.test.ts: exercises run-task.ts with aws-sdk-client-mock (run with: npx tsx test/run-task.test.ts)
import assert from "node:assert/strict";
import { mockClient } from "aws-sdk-client-mock";
import { ECSClient, RunTaskCommand, DescribeTasksCommand } from "@aws-sdk/client-ecs";
import { CloudWatchLogsClient, FilterLogEventsCommand } from "@aws-sdk/client-cloudwatch-logs";
import { runOneOffTask, RunTaskFailed, type OneOffTask } from "../src/run-task.ts";
const ecs = mockClient(ECSClient);
const logs = mockClient(CloudWatchLogsClient);
const arn = "arn:aws:ecs:us-east-1:123456789012:task/prod/0a1b2c3d4e5f";
const job: OneOffTask = {
cluster: "prod", taskDefinition: "api:42", container: "app", command: ["npm", "run", "migrate"],
subnets: ["subnet-0a"], securityGroups: ["sg-0abc"], logGroup: "/ecs/api", logStreamPrefix: "api",
};
ecs.on(RunTaskCommand).resolves({ tasks: [{ taskArn: arn, lastStatus: "PROVISIONING" }], failures: [] });
ecs.on(DescribeTasksCommand)
.resolvesOnce({ tasks: [{ taskArn: arn, lastStatus: "RUNNING" }] })
.resolves({
tasks: [{ taskArn: arn, lastStatus: "STOPPED", stopCode: "EssentialContainerExited",
containers: [{ name: "app", exitCode: 0 }] }],
});
logs.on(FilterLogEventsCommand)
.resolvesOnce({ events: [{ message: "Running 3 migrations" }], nextToken: "t1" })
.resolves({ events: [{ message: "Done" }] });
const ok = await runOneOffTask(job, 1);
assert.equal(ok.exitCode, 0);
assert.deepEqual(ok.logs, ["Running 3 migrations", "Done"]);
const sent = ecs.commandCalls(RunTaskCommand)[0].args[0].input;
assert.equal(sent.launchType, "FARGATE");
assert.equal(sent.capacityProviderStrategy, undefined);
assert.equal(sent.networkConfiguration?.awsvpcConfiguration?.assignPublicIp, "DISABLED");
assert.deepEqual(sent.overrides?.containerOverrides?.[0].command, ["npm", "run", "migrate"]);
assert.equal(typeof sent.clientToken, "string");
assert.deepEqual(logs.commandCalls(FilterLogEventsCommand)[0].args[0].input.logStreamNames, ["api/app/0a1b2c3d4e5f"]);
// Spot uses a capacity provider strategy instead of a launch type.
ecs.on(DescribeTasksCommand).resolves({ tasks: [{ taskArn: arn, lastStatus: "STOPPED", containers: [{ name: "app", exitCode: 3 }] }] });
const failed = await runOneOffTask({ ...job, spot: true, logGroup: undefined }, 1);
assert.equal(failed.exitCode, 3);
const spot = ecs.commandCalls(RunTaskCommand)[1].args[0].input;
assert.equal(spot.launchType, undefined);
assert.deepEqual(spot.capacityProviderStrategy, [{ capacityProvider: "FARGATE_SPOT", weight: 1 }]);
// Placement failures come back in the response, not as an exception.
ecs.on(RunTaskCommand).resolves({ tasks: [], failures: [{ reason: "ATTRIBUTE", arn: "arn:aws:ecs:us-east-1:123456789012:cluster/prod" }] });
await assert.rejects(runOneOffTask(job, 1), (err: unknown) => err instanceof RunTaskFailed && err.failures[0].reason === "ATTRIBUTE");
console.log("run-task tests passed");
Passing a 1-second minDelay keeps the waiter fast in tests. The guide to mocking AWS SDK v3 clients in unit tests shows the same setup in Jest and Vitest, and the client-ecs package in the AWS SDK repository lists every command and waiter.
Limits of this approach
- It waits in the calling process. For jobs that run for hours, let the caller exit and react to the ECS task state change event instead.
- Only one container’s exit code is checked. If a sidecar is marked essential and exits first, it stops the task; check
stopCodeand every container. FilterLogEventsreads what the container wrote. For searching logs across many runs, running CloudWatch Logs Insights queries with SDK v3 is faster.- Container Insights is not needed for any of this, but it helps when a job is slow; checking whether ECS Container Insights is enabled shows the settings and cost.
Coming from v2? The v2 call was new AWS.ECS().runTask(params).promise() with the same parameter names. The free AWS SDK v2 to v3 converter drafts the change, and migrating a Node.js app from AWS SDK v2 to v3 covers the rest of the codebase.
Frequently asked questions
How do I run an ECS task with a different command?
Pass overrides.containerOverrides with the container’s name and a command array. It replaces the command from the task definition or image for that run only.
How do I get the exit code of an ECS task?
Wait until the task is STOPPED, then call DescribeTasks and read containers[].exitCode for the container you care about.
Why does RunTask succeed but no task starts?
Placement problems are returned in the failures array of a successful response. Check it on every call.
Can I run a task on Fargate Spot with RunTask?
Yes, with capacityProviderStrategy: [{ capacityProvider: "FARGATE_SPOT", weight: 1 }] and no launchType. The cluster must have the Fargate Spot capacity provider associated.
Related guides
Ask your AWS account in plain English
Your first 15 runs are free, with no OpenAI key needed.
npx chatwithcloud