Run Commands on EC2 With SSM SendCommand and AWS SDK v3

Dark computer screen showing lines of green command line text

Photo by Jake Walker on Unsplash

SSM SendCommand with AWS SDK v3 is SendCommandCommand from @aws-sdk/client-ssm. Send DocumentName: "AWS-RunShellScript", your script lines in Parameters.commands, and either InstanceIds (up to 50) or tag Targets. Wait with waitUntilCommandExecuted, then read Status, ResponseCode and StandardOutputContent from GetCommandInvocation.

Run Command, part of AWS Systems Manager, runs scripts on managed instances without SSH, open ports or key pairs. The SSM Agent on each instance picks up the command, runs it and reports back. This guide is for Node.js and TypeScript developers who want to use SSM send command from AWS SDK v3 in a deploy script, an ops tool or a Lambda function: one instance or a whole tagged fleet, with results you can act on and IAM that can’t reach the wrong machines.

Every sample was type-checked with strict tsc against @aws-sdk/client-ssm (SDK 3.1142.0) in September 2026 and run against aws-sdk-client-mock: a success, a failing script, and a tagged fleet with paged invocations.

Which Run Command API calls do you need?

SDK v3 command Use it to Output returned
SendCommandCommand Start the command on instances or tag targets Command.CommandId
GetCommandInvocationCommand Read one instance’s result First 24,000 characters of stdout, 8,000 of stderr
ListCommandsCommand Read the overall status across all targets Counts: targets, completed, errors, delivery timeouts
ListCommandInvocationsCommand List every instance’s result (with Details: true) Plugin Output, up to 2,500 characters
waitUntilCommandExecuted Poll GetCommandInvocation until one invocation finishes Throws on Failed, TimedOut, Cancelled

The limits come from the Systems Manager API reference. The waiter behavior comes from the SDK source: it retries on Pending, InProgress and Delayed, and on InvocationDoesNotExist, which you can get right after SendCommand because the Run Command API is eventually consistent.

Prerequisites

  • Instances that are managed nodes: SSM Agent running, an instance profile with the Systems Manager core permissions, and a route to the SSM endpoints. The report to find EC2 instances not managed by SSM lists the ones that aren’t.
  • Node.js 18 or later with @aws-sdk/client-ssm.
  • A role with the policy below; the guide to AWS SDK v3 credential providers covers profiles, SSO and CI roles.

How to use SSM SendCommand with AWS SDK v3, step by step

  1. Pick the documentAWS-RunShellScript for Linux, AWS-RunPowerShellScript for Windows. Both take a commands list and an executionTimeout.
  2. Pick the targetsInstanceIds for a few known instances, or Targets such as [{ Key: "tag:Role", Values: ["web"] }] for a fleet. Not both.
  3. Set the safety limitsMaxConcurrency (default 50) and MaxErrors (default 0) decide how fast the command spreads and when it stops.
  4. Send and keep the CommandIdEverything else is looked up by it.
  5. WaitFor one instance, waitUntilCommandExecuted. For a fleet, poll ListCommands until the command reaches a terminal status.
  6. Read resultsGetCommandInvocation per instance, or ListCommandInvocations with Details: true for all of them.

Example: a run-command module

run-command.ts

// run-command.ts: run shell commands on EC2 instances with SSM Run Command and collect the results.
import { setTimeout as sleep } from "node:timers/promises";
import {
  SSMClient,
  SendCommandCommand,
  GetCommandInvocationCommand,
  ListCommandsCommand,
  paginateListCommandInvocations,
  waitUntilCommandExecuted,
  type Target,
} from "@aws-sdk/client-ssm";

export interface ShellJob {
  commands: string[];
  instanceIds?: string[]; // up to 50 IDs
  targets?: Target[]; // or tags, e.g. [{ Key: "tag:Role", Values: ["web"] }]
  executionTimeoutSeconds?: number; // how long the script may run (document parameter, default 3600)
  deliveryTimeoutSeconds?: number; // TimeoutSeconds: how long to wait for the command to start (30 or more)
  maxConcurrency?: string; // "10" or "10%"; default 50
  maxErrors?: string; // "0", "1" or "5%"; default 0
  comment?: string; // up to 100 characters
  logGroup?: string; // send full output to CloudWatch Logs
}

export interface InvocationResult {
  instanceId: string;
  status: string;
  statusDetails: string;
  responseCode?: number;
  stdout: string;
  stderr: string;
}

const ssm = new SSMClient({});
const TERMINAL = new Set(["Success", "Failed", "TimedOut", "Cancelled"]);

export async function sendShellCommand(job: ShellJob): Promise<string> {
  if (!job.instanceIds?.length === !job.targets?.length) throw new Error("pass either instanceIds or targets");
  const res = await ssm.send(
    new SendCommandCommand({
      DocumentName: "AWS-RunShellScript",
      Parameters: { commands: job.commands, executionTimeout: [String(job.executionTimeoutSeconds ?? 3600)] },
      InstanceIds: job.instanceIds,
      Targets: job.targets,
      TimeoutSeconds: job.deliveryTimeoutSeconds ?? 600,
      MaxConcurrency: job.maxConcurrency,
      MaxErrors: job.maxErrors,
      Comment: job.comment,
      CloudWatchOutputConfig: job.logGroup ? { CloudWatchOutputEnabled: true, CloudWatchLogGroupName: job.logGroup } : undefined,
    }),
  );
  const id = res.Command?.CommandId;
  if (!id) throw new Error("SendCommand returned no CommandId");
  return id;
}

/** One instance: wait for the invocation, then read up to 24,000 characters of stdout and 8,000 of stderr. */
export async function runOnInstance(instanceId: string, commands: string[], maxWaitSeconds = 900, minDelay = 5): Promise<InvocationResult> {
  const commandId = await sendShellCommand({ commands, instanceIds: [instanceId], comment: "runOnInstance" });
  try {
    await waitUntilCommandExecuted({ client: ssm, maxWaitTime: maxWaitSeconds, minDelay }, { CommandId: commandId, InstanceId: instanceId });
  } catch {
    // The waiter throws for Failed, TimedOut and Cancelled too; the invocation below says which.
  }
  const inv = await ssm.send(new GetCommandInvocationCommand({ CommandId: commandId, InstanceId: instanceId }));
  return {
    instanceId,
    status: inv.Status ?? "Unknown",
    statusDetails: inv.StatusDetails ?? "",
    responseCode: inv.ResponseCode,
    stdout: inv.StandardOutputContent ?? "",
    stderr: inv.StandardErrorContent ?? "",
  };
}

/** Many instances: poll the parent command until it is terminal, then list every invocation with its output. */
export async function runOnFleet(job: ShellJob, pollSeconds = 10, maxWaitSeconds = 3600): Promise<{ status: string; results: InvocationResult[] }> {
  const commandId = await sendShellCommand(job);
  const deadline = Date.now() + maxWaitSeconds * 1000;
  let status = "Pending";
  let details = "Pending";
  // StatusDetails adds Incomplete: some invocations failed, but fewer than MaxErrors allows.
  while (!TERMINAL.has(status) && details !== "Incomplete") {
    if (Date.now() > deadline) throw new Error(`command ${commandId} still ${details} after ${maxWaitSeconds}s`);
    await sleep(pollSeconds * 1000);
    const cmd = (await ssm.send(new ListCommandsCommand({ CommandId: commandId }))).Commands?.[0];
    status = cmd?.Status ?? "Pending";
    details = cmd?.StatusDetails ?? status;
  }
  const results: InvocationResult[] = [];
  for await (const page of paginateListCommandInvocations({ client: ssm }, { CommandId: commandId, Details: true })) {
    for (const inv of page.CommandInvocations ?? []) {
      const plugin = inv.CommandPlugins?.[0];
      results.push({
        instanceId: inv.InstanceId ?? "",
        status: inv.Status ?? "Unknown",
        statusDetails: inv.StatusDetails ?? "",
        responseCode: plugin?.ResponseCode,
        stdout: plugin?.Output ?? "", // at most 2,500 characters here; use S3 or CloudWatch Logs for more
        stderr: "",
      });
    }
  }
  return { status: details, results };
}

runOnInstance swallows the waiter’s error on purpose: a failed script is a result you want to read, not an exception. runOnFleet pages through invocations with paginateListCommandInvocations; the guide to AWS SDK v3 paginators explains how those work.

Example: check a tagged fleet, 10% at a time

check-fleet.ts

// check-fleet.ts: check disk space and nginx on every instance tagged Role=web, 10% at a time.
// Usage: AWS_PROFILE=ops npx tsx check-fleet.ts
import { runOnFleet } from "./run-command.ts";

const { status, results } = await runOnFleet({
  commands: ["set -e", "df -h / | tail -1 | awk '{print $5}'", "systemctl is-active nginx"],
  targets: [{ Key: "tag:Role", Values: ["web"] }],
  maxConcurrency: "10%",
  maxErrors: "1",
  executionTimeoutSeconds: 120,
  comment: "disk and nginx check",
  logGroup: "/aws/ssm/check-fleet", // the instance profile writes here; custom policies often allow /aws/ssm/*
});
console.log(`command status: ${status}`);
console.table(results.map((r) => ({ instance: r.instanceId, status: r.statusDetails, exit: r.responseCode, output: r.stdout.trim().replace(/\n/g, " | ") })));
process.exitCode = results.every((r) => r.status === "Success") ? 0 : 1;

A run against a mocked SSM client, where two instances report nginx as not active:

Output

command status: Failed
┌─────────┬───────────────────────┬──────────────┬──────┬──────────────────┐
│ (index) │ instance              │ status       │ exit │ output           │
├─────────┼───────────────────────┼──────────────┼──────┼──────────────────┤
│ 0       │ 'i-0a1b2c3d4e5f60001' │ 'Success'    │ 0    │ '41% | active'   │
│ 1       │ 'i-0a1b2c3d4e5f60002' │ 'Success'    │ 0    │ '63% | active'   │
│ 2       │ 'i-0a1b2c3d4e5f60003' │ 'Failed'     │ 3    │ '88% | inactive' │
│ 3       │ 'i-0a1b2c3d4e5f60004' │ 'Failed'     │ 3    │ '35% | failed'   │
│ 4       │ 'i-0a1b2c3d4e5f60005' │ 'Terminated' │ -1   │ ''               │
└─────────┴───────────────────────┴──────────────┴──────┴──────────────────┘

With MaxErrors: "1", the second failure pushed the command over its limit, so Run Command stopped sending it and the last instance shows Terminated. That’s the point of the limit: a broken script stops after a couple of machines instead of reaching all of them.

How do timeouts, MaxConcurrency and MaxErrors work?

  • TimeoutSeconds (30 to 2,592,000) is the delivery timeout: if the command hasn’t started on an instance by then, it won’t run there. The module sets 600.
  • executionTimeout is a document parameter, a string, 3,600 seconds by default for AWS-RunShellScript. The Systems Manager guide says the total timeout is the two added together, and that a command that runs too long ends as Execution Timed Out.
  • MaxConcurrency is a number or a percentage of targets running at once. Small batches let early failures stop the rollout.
  • MaxErrors counts Failed and Execution Timed Out invocations. Delivery timeouts and undeliverable instances don’t count against it, so an offline fleet won’t stop the command; check those statuses separately.

Spreading a risky change across a small batch first and stopping on errors is the same idea as canarying releases in Google’s SRE workbook. For patching specifically, Patch Manager already does this; checking SSM patch compliance shows how to read its results.

Where does the command output go?

The API returns only part of it: 24,000 characters of stdout and 8,000 of stderr from GetCommandInvocation, and 2,500 characters per plugin from ListCommandInvocations. For anything longer, set OutputS3BucketName (and OutputS3KeyPrefix), or CloudWatchOutputConfig with a log group, as check-fleet.ts does. The instance writes that output with its own instance profile, not your role. With AmazonSSMManagedInstanceCore and CloudWatchAgentServerPolicy attached, CloudWatch output needs no extra setup; custom instance policies in the Systems Manager guide grant log access on /aws/ssm/* groups, which is why the example uses that prefix. Give the log group a retention period; setting CloudWatch Logs retention for all log groups does it in bulk.

Keep secrets out of commands: parameters are stored with the command and shown in the console and in ListCommands. Read secrets on the instance instead, for example from Parameter Store; getting an SSM parameter with AWS SDK v3 covers the API, and the audit to find plaintext SSM parameters checks they’re encrypted.

Permissions needed

send-command-policy.json

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "UseRunShellScriptOnly",
      "Effect": "Allow",
      "Action": "ssm:SendCommand",
      "Resource": "arn:aws:ssm:us-east-1:*:document/AWS-RunShellScript"
    },
    {
      "Sid": "OnlyWebInstances",
      "Effect": "Allow",
      "Action": "ssm:SendCommand",
      "Resource": "arn:aws:ec2:us-east-1:123456789012:instance/*",
      "Condition": { "StringEquals": { "ssm:resourceTag/Role": "web" } }
    },
    {
      "Sid": "ReadResults",
      "Effect": "Allow",
      "Action": ["ssm:GetCommandInvocation", "ssm:ListCommands", "ssm:ListCommandInvocations"],
      "Resource": "*"
    }
  ]
}

SendCommand is authorized against both the document and each instance, so you scope it twice: which document may run, and on which instances. The service authorization reference lists ssm:resourceTag/${TagKey} as a condition key for it. The three read actions don’t support resource-level permissions. Test the tag condition with both InstanceIds and Targets; instances the caller may not reach show the invocation status AccessDenied. SSM Agent runs with root permissions on Linux and SYSTEM on Windows, so permission to send AWS-RunShellScript is effectively root access: treat it like SSH access. To derive actions from your code, see finding the IAM actions in AWS SDK for JavaScript code or the IAM policy generator for TypeScript code. Shared documents are a risk of their own; finding public SSM documents checks yours.

Troubleshooting and common mistakes

  • InvalidInstanceId. The API reference lists the causes: no permission for that node, SSM Agent not running or not registered, or an instance that is shutting down or terminated.
  • InvalidParameters. A parameter isn’t defined in the document, or a required one is missing. executionTimeout must be a string inside an array.
  • Status Success but nothing happened. The Systems Manager guide warns that Success only means an exit code of 0, and that targeting a tag with no matching instances can also return Success. Use set -e and check TargetCount.
  • Delivery Timed Out on some instances. They were offline or not managed. They don’t count against MaxErrors, so look for them in the results.
  • Commands hang in InProgress. The agent went away mid-run; the status turns terminal when the execution timeout passes.
  • SSH still needed for debugging? Usually not; the example on troubleshooting EC2 SSH connections covers the cases where you do.

How do you test SendCommand code?

run-command.test.ts

// run-command.test.ts: exercises run-command.ts with aws-sdk-client-mock (run with: npx tsx test/run-command.test.ts)
import assert from "node:assert/strict";
import { mockClient } from "aws-sdk-client-mock";
import {
  SSMClient,
  SendCommandCommand,
  GetCommandInvocationCommand,
  ListCommandsCommand,
  ListCommandInvocationsCommand,
  InvocationDoesNotExist,
} from "@aws-sdk/client-ssm";
import { runOnInstance, runOnFleet, sendShellCommand } from "../src/run-command.ts";

const ssm = mockClient(SSMClient);
const id = "3f1c2a9e-7b4d-4e8a-9c0f-1a2b3c4d5e6f";
ssm.on(SendCommandCommand).resolves({ Command: { CommandId: id } });

// Single instance: the first read races the command (InvocationDoesNotExist), then it succeeds.
ssm.on(GetCommandInvocationCommand)
  .rejectsOnce(new InvocationDoesNotExist({ message: "not yet", $metadata: { httpStatusCode: 400 } }))
  .resolves({ Status: "Success", StatusDetails: "Success", ResponseCode: 0, StandardOutputContent: "active\n", StandardErrorContent: "" });
const one = await runOnInstance("i-0abc1234def567890", ["systemctl is-active nginx"], 60, 1);
assert.equal(one.status, "Success");
assert.equal(one.stdout, "active\n");
const sent = ssm.commandCalls(SendCommandCommand)[0].args[0].input;
assert.equal(sent.DocumentName, "AWS-RunShellScript");
assert.deepEqual(sent.Parameters, { commands: ["systemctl is-active nginx"], executionTimeout: ["3600"] });
assert.deepEqual(sent.InstanceIds, ["i-0abc1234def567890"]);

// A failing script: the waiter throws, the invocation still reports exit code and stderr.
ssm.on(GetCommandInvocationCommand).resolves({ Status: "Failed", StatusDetails: "Failed", ResponseCode: 3, StandardOutputContent: "", StandardErrorContent: "inactive" });
const bad = await runOnInstance("i-0abc1234def567890", ["systemctl is-active nginx"], 60, 1);
assert.deepEqual([bad.status, bad.responseCode, bad.stderr], ["Failed", 3, "inactive"]);

// Fleet by tag: poll ListCommands, then page through invocations.
ssm.on(ListCommandsCommand)
  .resolvesOnce({ Commands: [{ CommandId: id, Status: "InProgress", StatusDetails: "InProgress" }] })
  .resolves({ Commands: [{ CommandId: id, Status: "Success", StatusDetails: "Incomplete" }] });
ssm.on(ListCommandInvocationsCommand)
  .resolvesOnce({ CommandInvocations: [{ InstanceId: "i-01", Status: "Success", StatusDetails: "Success", CommandPlugins: [{ ResponseCode: 0, Output: "ok" }] }], NextToken: "n" })
  .resolves({ CommandInvocations: [{ InstanceId: "i-02", Status: "TimedOut", StatusDetails: "Delivery Timed Out", CommandPlugins: [{ ResponseCode: -1, Output: "" }] }] });
const fleet = await runOnFleet({ commands: ["uptime"], targets: [{ Key: "tag:Role", Values: ["web"] }], maxConcurrency: "10%", maxErrors: "1" }, 0.01, 30);
assert.equal(fleet.status, "Incomplete");
assert.deepEqual(fleet.results.map((r) => [r.instanceId, r.statusDetails, r.responseCode]), [["i-01", "Success", 0], ["i-02", "Delivery Timed Out", -1]]);
const tagged = ssm.commandCalls(SendCommandCommand)[2].args[0].input;
assert.deepEqual([tagged.Targets, tagged.InstanceIds, tagged.MaxConcurrency, tagged.MaxErrors], [[{ Key: "tag:Role", Values: ["web"] }], undefined, "10%", "1"]);

await assert.rejects(sendShellCommand({ commands: ["uptime"] }), /either instanceIds or targets/);
console.log("run-command tests passed");

The first case checks that InvocationDoesNotExist is retried by the waiter rather than surfaced. The guide to mocking AWS SDK v3 clients in unit tests shows the same pattern in Jest and Vitest.

Limits of this approach

  • Polling from the caller is fine for minutes, not hours. For long jobs, subscribe to Run Command status changes through EventBridge or SNS and let the caller exit.
  • InstanceIds takes at most 50 IDs and Targets at most 5 key-value filters per call.
  • Run Command runs a script; it doesn’t keep machines in a desired state. For repeated configuration, use State Manager associations or your configuration tool. If a job doesn’t need a particular machine, one-off ECS tasks started with RunTask in AWS SDK v3 run it in a fresh container instead.
  • If you’re coming from v2, new AWS.SSM().sendCommand(params).promise() takes the same parameters; the free AWS SDK v2 to v3 converter drafts the change.

To investigate what went wrong across a fleet in plain English, troubleshooting AWS infrastructure with an AI CLI shows another route.

Frequently asked questions

How do I get the output of an SSM SendCommand?

Call GetCommandInvocation with the CommandId and InstanceId. It returns up to 24,000 characters of stdout; send output to S3 or CloudWatch Logs for more.

How many instances can SendCommand target?

Up to 50 with InstanceIds. With tag Targets you can reach many more, controlled by MaxConcurrency.

Why does GetCommandInvocation throw InvocationDoesNotExist?

You asked too soon after SendCommand, or the IDs don’t match. The SDK’s waitUntilCommandExecuted retries on this error.

What is the difference between TimeoutSeconds and executionTimeout?

TimeoutSeconds limits how long the command waits to start on an instance. executionTimeout limits how long the script may run once started.

Related guides

Ask your AWS account in plain English

Your first 15 runs are free, with no OpenAI key needed.

npx chatwithcloud