Read an S3 Object With AWS SDK v3 (GetObject)

An open book resting on a wooden table in front of tall library shelves

Photo by Jayanth Muppaneni on Unsplash

S3 GetObject in AWS SDK v3 returns the object’s Body as a stream, not a Buffer. Send new GetObjectCommand({ Bucket, Key }) with an S3Client, then call await Body.transformToString() for text, transformToByteArray() for bytes, or pipe the Node.js stream to a file with pipeline() for large objects. Catch NoSuchKey for missing keys, and read each body exactly once.

Reading a file from S3 is one of the first things people write with the AWS SDK, and one of the first things that breaks after a v2 to v3 migration. In v2, getObject().promise() handed you a Buffer. In v3, Body is a stream with helper methods, so Body.toString() prints [object Object] and JSON.parse(Body) fails.

This guide is for Node.js and TypeScript developers on AWS SDK for JavaScript v3. You’ll get one typed module that covers S3 GetObject in AWS SDK v3 for text, JSON, bytes, byte ranges and large files, plus missing keys and conditional reads, with the IAM permissions and the mistakes that cause memory and socket problems. Every function here was type-checked with strict tsc and run against @aws-sdk/client-s3 3.1141.0 in September 2026. If you only need to know whether a key exists, the script to check if an S3 object exists in TypeScript uses HeadObject and downloads nothing. If a browser or another service should download the file directly, create a presigned S3 download URL with AWS SDK v3 instead.

What does GetObject return in AWS SDK v3?

The output has the object’s metadata as fields (ContentType, ContentLength, ETag, LastModified, Metadata) and the data in Body. In Node.js, Body is a Readable stream with three helpers mixed in:

Method Returns Use it for
transformToString(encoding?) Promise<string> Text, CSV and JSON files that fit comfortably in memory
transformToByteArray() Promise<Uint8Array> Images, gzip and other binary data you process in memory
transformToWebStream() ReadableStream Web-stream APIs such as Response or TextDecoderStream
The stream itself Readable Large files: pipe to disk or another stream without buffering

Each body can be consumed once. A second call to any helper throws The stream has already been transformed., straight from the SDK’s stream code. If you need the text twice, keep the string.

Prerequisites

How to read an S3 object with GetObject, step by step

  1. Create one client and reuse itnew S3Client({}) at module level. Creating a client per request throws away its connection pool.
  2. Send GetObjectCommandPass Bucket and Key, plus Range, VersionId or IfNoneMatch when you need them.
  3. Check that Body existsIt’s typed as optional, so narrow it before calling a helper.
  4. Consume the body oncePick one: a transformTo* helper for small objects, or pipeline() for large ones.
  5. Handle the expected errorsNoSuchKey for a missing key, a 304 for an unchanged object, AccessDenied for permissions.

Example: a typed module for reading S3 objects

s3-read.ts

// s3-read.ts: read S3 objects with AWS SDK for JavaScript v3 (GetObject)
import { createWriteStream } from "node:fs";
import { Readable } from "node:stream";
import { pipeline } from "node:stream/promises";
import { GetObjectCommand, NoSuchKey, S3Client, S3ServiceException } from "@aws-sdk/client-s3";

const s3 = new S3Client({}); // Region and credentials come from the environment or profile

/** Whole object as text. Fine for small files: the entire body is held in memory. */
export async function readText(Bucket: string, Key: string): Promise<string> {
  const { Body } = await s3.send(new GetObjectCommand({ Bucket, Key }));
  if (!Body) throw new Error(`s3://${Bucket}/${Key} returned no body`);
  return Body.transformToString("utf-8");
}

/** Parse a JSON object. The type parameter is a promise to the compiler, not a runtime check. */
export async function readJson<T>(Bucket: string, Key: string): Promise<T> {
  return JSON.parse(await readText(Bucket, Key)) as T;
}

/** Raw bytes, for images, gzip files or anything that isn't text. */
export async function readBytes(Bucket: string, Key: string): Promise<Uint8Array> {
  const { Body } = await s3.send(new GetObjectCommand({ Bucket, Key }));
  if (!Body) throw new Error(`s3://${Bucket}/${Key} returned no body`);
  return Body.transformToByteArray();
}

/** Stream a large object to disk without loading it into memory. */
export async function downloadToFile(Bucket: string, Key: string, path: string): Promise<number> {
  const { Body, ContentLength } = await s3.send(new GetObjectCommand({ Bucket, Key }));
  if (!(Body instanceof Readable)) throw new Error("Expected a Node.js Readable body (are you running in a browser?)");
  await pipeline(Body, createWriteStream(path)); // handles backpressure and closes both streams on error
  return ContentLength ?? 0;
}

/** Read bytes start..end inclusive. S3 answers 206 Partial Content with a Content-Range header. */
export async function readRange(Bucket: string, Key: string, start: number, end: number): Promise<{ bytes: Uint8Array; range?: string }> {
  const { Body, ContentRange } = await s3.send(new GetObjectCommand({ Bucket, Key, Range: `bytes=${start}-${end}` }));
  if (!Body) throw new Error(`s3://${Bucket}/${Key} returned no body`);
  return { bytes: await Body.transformToByteArray(), range: ContentRange };
}

/** Return undefined when the object is missing instead of throwing. */
export async function readTextIfExists(Bucket: string, Key: string): Promise<string | undefined> {
  try {
    return await readText(Bucket, Key);
  } catch (err) {
    if (err instanceof NoSuchKey) return undefined; // 404: needs s3:ListBucket, or S3 answers 403 instead
    throw err;
  }
}

/** Download only if the object changed since the ETag you cached. */
export async function readIfChanged(Bucket: string, Key: string, etag?: string): Promise<{ changed: boolean; etag?: string; text?: string }> {
  try {
    const res = await s3.send(new GetObjectCommand({ Bucket, Key, IfNoneMatch: etag }));
    return { changed: true, etag: res.ETag, text: await res.Body?.transformToString() };
  } catch (err) {
    if (err instanceof S3ServiceException && err.$metadata.httpStatusCode === 304) return { changed: false, etag };
    throw err;
  }
}

Use it like any other module: const cfg = await readJson<AppConfig>("my-config-bucket", "prod/app.json"). readJson trusts the file’s shape; validate it with a schema library if the file comes from outside your team. For tests, the guide to mock AWS SDK v3 clients in Jest and Vitest shows how to fake a Body with sdkStreamMixin.

How do you stream a large S3 object to a file?

transformToString and transformToByteArray hold the whole object in memory. That’s fine for a 200 KB config file and a crash for a 6 GB export. downloadToFile passes the Body stream to pipeline from node:stream/promises, which moves data in chunks, respects backpressure and destroys both streams if either fails. The Node.js stream documentation explains why pipeline is safer than .pipe(), which doesn’t forward errors or clean up.

The same pattern works for any destination stream: a gzip transform, a CSV parser, or an upload to another bucket. For large writes in the other direction, see upload large files and streams to S3 with AWS SDK v3, and to copy between buckets without downloading at all, copy and move S3 objects with AWS SDK v3. When the object is a scanned PDF or an image and you need its text, extract text from PDFs and images with Amazon Textract and SDK v3 reads it straight from the bucket.

Range reads, ETags and conditional GETs

  • Range: "bytes=0-1023" returns only those bytes, inclusive, with status 206 and a ContentRange such as bytes 0-1023/16000. Use it to read a file header, resume a download or split a big object into parallel parts. S3 doesn’t support multiple ranges in one GET.
  • IfNoneMatch: etag returns the object only if its ETag changed. If it didn’t, S3 answers 304 Not Modified, and the SDK throws. In version 3.1141.0 the thrown error’s name is Unknown, so readIfChanged checks $metadata.httpStatusCode === 304 instead of the name.
  • IfMatch: etag does the opposite: it fails with 412 Precondition Failed if the object changed since you read the ETag.
  • VersionId reads an older version in a versioned bucket and needs s3:GetObjectVersion.

Every GET is a billable request and every byte leaving the Region is data transfer. The guides to calculate S3 GET and PUT request costs and estimate AWS S3 data transfer out cost show what a hot read path adds up to.

Which IAM permissions does GetObject need?

s3-read-policy.json

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "ReadObjects",
      "Effect": "Allow",
      "Action": "s3:GetObject",
      "Resource": "arn:aws:s3:::my-config-bucket/*"
    },
    {
      "Sid": "Get404InsteadOf403ForMissingKeys",
      "Effect": "Allow",
      "Action": "s3:ListBucket",
      "Resource": "arn:aws:s3:::my-config-bucket"
    }
  ]
}

s3:GetObject goes on the object ARN (bucket/*); s3:ListBucket goes on the bucket ARN. The second statement is optional, but without it S3 answers 403 Access Denied instead of 404 for a missing key, so your NoSuchKey branch never runs. Add s3:GetObjectVersion if you pass VersionId, and kms:Decrypt on the key for objects encrypted with SSE-KMS. The free IAM policy generator for TypeScript code drafts this from your module, and the guide to find the IAM actions your AWS SDK for JavaScript code needs explains the mapping.

Troubleshooting and common mistakes

  • [object Object] or Unexpected token in JSON.parse. You passed the stream, not its contents. Await Body.transformToString() first. If you’re migrating, the AWS SDK JavaScript v2 to v3 converter drafts the new calls; check every place its output reads Body.
  • The stream has already been transformed. The body was read twice, often once for logging and once for parsing. Read it once into a variable.
  • Requests hang, then @smithy/node-http-handler:WARN - socket usage at capacity=50 and ... additional requests are enqueued. Bodies you never read keep their sockets busy, and the default pool has 50. Always consume or destroy Body, even when you only needed the headers; for headers alone, use HeadObject.
  • AccessDenied for a key that doesn’t exist. The caller lacks s3:ListBucket, as above. The steps to troubleshoot AWS IAM access denied errors help with the rest: bucket policies, KMS key policies and SCPs.
  • InvalidObjectState. The object is in S3 Glacier Flexible Retrieval, Glacier Deep Archive or an Intelligent-Tiering archive tier. Restore it with RestoreObject and read it when the restore finishes.
  • Out-of-memory crashes. A helper was used on a large object. Switch to downloadToFile or process the stream in chunks.

Limits: what S3 GetObject in AWS SDK v3 can’t do

  • It reads one object per request. There’s no batch GET; to read many objects, run several requests with a concurrency limit.
  • It returns bytes, not rows. To filter inside a CSV or Parquet file, download and parse it, or query the data with Athena.
  • Archived storage classes need a restore first, which takes minutes to hours depending on the class and tier.
  • The stream can fail mid-download on a network error. Retries in the SDK cover the request, not a half-read body, so resume with a Range from the last byte you wrote. The guide to configure retry and timeout settings in AWS SDK for JavaScript v3 covers timeouts for slow streams.

To see which buckets hold the most data before you start reading from them, ChatWithCloud can answer in plain English; the walkthrough to ask AI which S3 buckets are largest shows the questions and the calls it makes.

Frequently asked questions

How do I get the S3 object body as a string in AWS SDK v3?

Use const text = await response.Body?.transformToString();. Pass an encoding such as "utf-8" if you want to be explicit. It reads the whole body into memory.

Why is Body a stream instead of a Buffer in SDK v3?

So you can process large objects without loading them into memory. The transformTo* helpers give you the v2-style result when the object is small.

How do I handle NoSuchKey in AWS SDK v3?

Import NoSuchKey from @aws-sdk/client-s3 and check err instanceof NoSuchKey. Make sure the caller has s3:ListBucket, or S3 returns 403 instead of 404.

Can I read only part of an S3 object?

Yes. Set Range: "bytes=start-end" on GetObjectCommand. S3 returns status 206 with just those bytes, one range per request.

Related guides

Ask your AWS account in plain English

Your first 15 runs are free, with no OpenAI key needed.

npx chatwithcloud