Invoke Amazon Bedrock Models With AWS SDK v3 (Converse API)

Abstract blue and cyan light trails forming a network of connected points on a dark background

Photo by Shubham Dhage on Unsplash

To invoke a Bedrock model with AWS SDK v3, install @aws-sdk/client-bedrock-runtime and send a ConverseCommand with a modelId (or inference profile ID such as us.amazon.nova-lite-v1:0) and a messages array. Read the answer from output.message.content. Use ConverseStreamCommand to stream tokens, and InvokeModelCommand only when you need a model’s native request format.

Amazon Bedrock gives you foundation models from Amazon and other providers behind one AWS API, billed to your AWS account and authorized with IAM. The SDK part is short. What trips people up is everything around it: which model ID to pass, why the same ID works in one Region and fails in another, which IAM resources to allow, and what to do when calls start throttling.

This guide is for Node.js and TypeScript developers adding a model call to an app, a Lambda function or a script. You’ll get a typed module that covers the Bedrock Converse API, streaming and the lower-level InvokeModel call with AWS SDK v3, plus the IAM policy and the fixes for the errors you’re most likely to see.

Converse or InvokeModel: which should you use?

ConverseCommand / ConverseStreamCommand InvokeModelCommand / InvokeModelWithResponseStreamCommand
Request body One shape for every model that supports messages: messages, system, inferenceConfig The provider’s own JSON, different for each model family
Response Normalized: output, stopReason, usage, metrics Raw bytes you decode and parse per model
Switching models Change the modelId Rewrite the body and the parser
Model-specific options Through additionalModelRequestFields Directly in the body
IAM action bedrock:InvokeModel (Converse), bedrock:InvokeModelWithResponseStream (ConverseStream) The same two actions

Default to Converse. Use InvokeModel for models that don’t take messages, such as embedding models, or when a provider feature isn’t exposed through Converse.

Prerequisites

How to invoke a Bedrock model with AWS SDK v3, step by step

  1. Create one runtime clientnew BedrockRuntimeClient({ retryMode: "adaptive" }) at module level. The runtime client (bedrock-runtime) sends prompts; the separate @aws-sdk/client-bedrock client manages models and profiles.
  2. Pick a model or inference profile IDFor example amazon.nova-lite-v1:0 in a Region that serves it directly, or the US cross-Region profile us.amazon.nova-lite-v1:0.
  3. Build the messagesEach message has a role (user or assistant) and a content array of blocks such as { text }. Put instructions in system.
  4. Set limitsinferenceConfig.maxTokens caps the answer and the output-token cost; temperature and topP control randomness.
  5. Send and read the resultJoin the text blocks in output.message.content, and check stopReason: max_tokens means the answer was cut off.
  6. Record usageusage.inputTokens and usage.outputTokens are what you pay for. Log them per request.

Example: a typed Bedrock module

bedrock.ts

// bedrock.ts
// Call Amazon Bedrock models with AWS SDK for JavaScript v3: Converse for a full reply,
// ConverseStream for token-by-token output, and InvokeModel for a model's native request body.
import {
  BedrockRuntimeClient,
  ConverseCommand,
  ConverseStreamCommand,
  InvokeModelCommand,
  ThrottlingException,
  ValidationException,
  type Message,
} from "@aws-sdk/client-bedrock-runtime";

// Adaptive retry mode slows the client down when Bedrock returns ThrottlingException.
const bedrock = new BedrockRuntimeClient({ retryMode: "adaptive", maxAttempts: 6 });

// A US cross-Region inference profile for Amazon Nova Lite. The plain model ID
// "amazon.nova-lite-v1:0" also works in Regions that serve the model in-Region.
export const MODEL_ID = process.env.BEDROCK_MODEL_ID ?? "us.amazon.nova-lite-v1:0";

export interface Reply {
  text: string;
  stopReason: string;
  inputTokens: number;
  outputTokens: number;
  latencyMs: number;
}

/** One request, one complete answer. Pass the running history to keep a conversation. */
export async function ask(history: Message[], system = "You are a concise assistant for AWS engineers."): Promise<Reply> {
  try {
    const res = await bedrock.send(
      new ConverseCommand({
        modelId: MODEL_ID,
        system: [{ text: system }],
        messages: history,
        inferenceConfig: { maxTokens: 512, temperature: 0.2 },
      }),
    );
    const text = (res.output?.message?.content ?? []).map((block) => block.text ?? "").join("");
    return {
      text,
      stopReason: res.stopReason ?? "unknown",
      inputTokens: res.usage?.inputTokens ?? 0,
      outputTokens: res.usage?.outputTokens ?? 0,
      latencyMs: res.metrics?.latencyMs ?? 0,
    };
  } catch (err) {
    if (err instanceof ValidationException && /inference profile/i.test(err.message)) {
      throw new Error(`${MODEL_ID} needs an inference profile ID here, such as us.${MODEL_ID}`, { cause: err });
    }
    if (err instanceof ThrottlingException) {
      throw new Error("Bedrock is still throttling after retries: lower concurrency or request a quota increase", { cause: err });
    }
    throw err;
  }
}

/** Stream the answer, calling onText for each chunk as the model writes it. */
export async function askStream(prompt: string, onText: (chunk: string) => void): Promise<Reply> {
  const res = await bedrock.send(
    new ConverseStreamCommand({
      modelId: MODEL_ID,
      messages: [{ role: "user", content: [{ text: prompt }] }],
      inferenceConfig: { maxTokens: 512 },
    }),
  );
  const reply: Reply = { text: "", stopReason: "unknown", inputTokens: 0, outputTokens: 0, latencyMs: 0 };
  for await (const event of res.stream ?? []) {
    if (event.contentBlockDelta?.delta?.text) {
      reply.text += event.contentBlockDelta.delta.text;
      onText(event.contentBlockDelta.delta.text);
    } else if (event.messageStop) {
      reply.stopReason = event.messageStop.stopReason ?? "unknown";
    } else if (event.metadata) {
      reply.inputTokens = event.metadata.usage?.inputTokens ?? 0;
      reply.outputTokens = event.metadata.usage?.outputTokens ?? 0;
      reply.latencyMs = event.metadata.metrics?.latencyMs ?? 0;
    }
  }
  return reply;
}

/** InvokeModel sends the provider's own JSON. This body is the Amazon Nova format. */
export async function invokeNova(prompt: string): Promise<string> {
  const res = await bedrock.send(
    new InvokeModelCommand({
      modelId: MODEL_ID,
      contentType: "application/json",
      accept: "application/json",
      body: JSON.stringify({
        messages: [{ role: "user", content: [{ text: prompt }] }],
        inferenceConfig: { maxTokens: 512 },
      }),
    }),
  );
  const json = JSON.parse(new TextDecoder().decode(res.body));
  return json.output?.message?.content?.[0]?.text ?? "";
}

A caller that holds a two-turn conversation, streams a second answer and makes one native call:

chat.ts

// chat.ts
// Usage: AWS_PROFILE=dev AWS_REGION=us-east-1 npx tsx chat.ts "Explain what a NAT gateway charges for"
import type { Message } from "@aws-sdk/client-bedrock-runtime";
import { ask, askStream, invokeNova, MODEL_ID } from "./bedrock.js";

const question = process.argv[2] ?? "In two sentences, what does an AWS NAT gateway charge for?";
console.log(`Model: ${MODEL_ID}\n`);

// 1. Converse: a two-turn conversation. Converse is stateless, so you send the whole history each time.
const history: Message[] = [{ role: "user", content: [{ text: question }] }];
const first = await ask(history);
console.log(`Converse -> ${first.text}`);
console.log(`  stop=${first.stopReason} in=${first.inputTokens} out=${first.outputTokens} ${first.latencyMs} ms\n`);

history.push({ role: "assistant", content: [{ text: first.text }] });
history.push({ role: "user", content: [{ text: "Give one way to cut that cost." }] });
const second = await ask(history);
console.log(`Follow-up -> ${second.text}\n`);

// 2. ConverseStream: print tokens as they arrive.
process.stdout.write("Stream -> ");
const streamed = await askStream("List three AWS services that bill per request, one per line.", (t) => process.stdout.write(t));
console.log(`\n  stop=${streamed.stopReason} out=${streamed.outputTokens}\n`);

// 3. InvokeModel with the model's native body.
console.log(`InvokeModel -> ${await invokeNova("Reply with the single word: ready")}`);
Terminal

npm install @aws-sdk/client-bedrock-runtime
npm install --save-dev tsx typescript @types/node
npm pkg set type=module

AWS_PROFILE=dev AWS_REGION=us-east-1 npx tsx chat.ts
# Model: us.amazon.nova-lite-v1:0
#
# Converse -> A NAT gateway charges an hourly fee for each hour it's provisioned and a per-GB fee for data it processes. ...
#   stop=end_turn in=31 out=48 612 ms
#
# Follow-up -> Route S3 and DynamoDB traffic through gateway VPC endpoints so it bypasses the NAT gateway. ...
#
# Stream -> Amazon S3
# AWS Lambda
# Amazon API Gateway
#   stop=end_turn out=12
#
# InvokeModel -> ready

Model output varies from run to run; token counts and latency are illustrative. If stop shows max_tokens, raise maxTokens or ask for a shorter answer.

Model IDs, inference profiles and model access

The modelId field accepts a base model ID, an inference profile ID or ARN, and a few other resource ARNs such as provisioned throughput. Three rules cover most cases:

  • Base model IDs like amazon.nova-lite-v1:0 run in the Region you call, when that Region serves the model on demand.
  • Cross-Region inference profiles add a geography prefix: us., eu., apac., or global. for models that support it. Bedrock routes the request to one of the profile’s destination Regions. For us.amazon.nova-lite-v1:0 called from us-east-1, those are us-east-1, us-east-2 and us-west-2.
  • Some models only run through a profile. Calling their base ID returns a ValidationException saying on-demand throughput isn’t supported and to retry with an inference profile. The ask() function above turns that into a clearer error.

Each model’s page in the Bedrock documentation lists its IDs, Regions and whether it supports Converse and streaming. Access to serverless models is enabled by default in commercial Regions. The first call to a third-party model sold through AWS Marketplace starts a subscription in the background, which needs aws-marketplace:Subscribe, aws-marketplace:Unsubscribe and aws-marketplace:ViewSubscriptions for that first caller, and Anthropic models also need a one-time use case form per account or organization. Amazon models such as Nova have no Marketplace product, so they skip that step.

Permissions needed

Scope the policy to the models you use. With a cross-Region profile, allow the profile and the base model in every destination Region:

bedrock-invoke-policy.json

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "InvokeNovaLiteThroughUsProfile",
      "Effect": "Allow",
      "Action": [
        "bedrock:InvokeModel",
        "bedrock:InvokeModelWithResponseStream"
      ],
      "Resource": [
        "arn:aws:bedrock:us-east-1:123456789012:inference-profile/us.amazon.nova-lite-v1:0",
        "arn:aws:bedrock:us-east-1::foundation-model/amazon.nova-lite-v1:0",
        "arn:aws:bedrock:us-east-2::foundation-model/amazon.nova-lite-v1:0",
        "arn:aws:bedrock:us-west-2::foundation-model/amazon.nova-lite-v1:0"
      ]
    }
  ]
}

Foundation model ARNs have no account ID. A service control policy that blocks one destination Region makes profile calls fail even if the others are allowed. To confirm the actions your code calls, see how to find the IAM actions your AWS SDK JavaScript code needs, or paste bedrock.ts into the IAM policy generator for TypeScript code. To keep other identities off expensive models, deny bedrock:* on those model ARNs.

Throttling, quotas and cost

Bedrock limits requests and tokens per minute, per model and Region, as account quotas. When you exceed them, calls fail with ThrottlingException. The SDK retries throttling errors on its own; retryMode: "adaptive" also slows the client down after throttles, which suits batch jobs that fire many calls. The guide to configure retries and timeouts in AWS SDK v3 explains attempts and backoff, and monitor AWS service quota usage and get alerts covers watching quotas before they bite. Cross-Region profiles also help, because they spread load across Regions.

On-demand pricing is per input token and per output token, at separate rates that differ from model to model. Check the Amazon Bedrock pricing page for current rates before choosing a model. Record usage from every response; you can publish token counts as CloudWatch custom metrics and alarm on them, and set an AWS budget alert with SDK v3 for the account.

Troubleshooting common Bedrock errors

  • ValidationException: on-demand throughput isn’t supported. Use the model’s inference profile ID, such as us. plus the model ID.
  • ValidationException: the provided model identifier is invalid. A typo, or the model isn’t offered in this Region. Check the model’s page for Regions.
  • AccessDeniedException. Missing bedrock:InvokeModel on the profile or on a destination Region’s model ARN, a missing Marketplace subscription for a third-party model, or an Anthropic model without the use case form. The steps to troubleshoot IAM access denied errors help separate these.
  • ThrottlingException after retries. Lower concurrency, spread load with a cross-Region profile, or request a quota increase.
  • ModelTimeoutException or very slow replies. Long prompts and large maxTokens take longer. Stream the response, or raise the client’s request timeout for long generations.
  • Empty text. The model answered with a non-text block, such as a tool use request, or a guardrail intervened. Check stopReason.

Limits of this approach

The module sends one request at a time and keeps conversation history in memory; a chat product needs to store history, trim it to the model’s context window and summarize old turns. It doesn’t cover tool use, guardrails, prompt caching or batch inference, which all build on the same client. Model output is not deterministic and can be wrong, so validate anything you act on. Don’t put secrets in prompts; load keys from AWS with the guide to get a Secrets Manager secret value with SDK v3 and keep them out of the text you send.

Porting older code? The AWS SDK JavaScript v2 to v3 converter drafts the change, and the guide to migrate a Node.js app from AWS SDK v2 to v3 covers the rest. If you’d rather ask questions about your AWS account than build a model call yourself, ChatWithCloud answers them in plain English from your terminal; how ChatWithCloud works explains which data goes to the model.

Frequently asked questions

What is the difference between Converse and InvokeModel in Bedrock?

Converse uses one request and response shape for every model that supports messages. InvokeModel takes each provider’s native JSON body and returns raw bytes. Both need bedrock:InvokeModel.

How do I stream a Bedrock response in Node.js?

Send ConverseStreamCommand and iterate response.stream with for await. Text arrives in contentBlockDelta.delta.text, the stop reason in messageStop, and token usage in the final metadata event.

Why does my Bedrock model ID work in one Region but not another?

Models aren’t offered in every Region, and some only run through an inference profile. Check the model’s Regional availability, and use a us., eu. or apac. profile ID where needed.

Do I need to request model access in Amazon Bedrock?

Not for most models: access to serverless models is on by default in commercial Regions. Third-party Marketplace models subscribe on first use, which needs Marketplace permissions, and Anthropic models need a one-time use case form.

Related guides

Ask your AWS account in plain English

Your first 15 runs are free, with no OpenAI key needed.

npx chatwithcloud