Photo by Shubham Dhage on Unsplash
To invoke a Bedrock model with AWS SDK v3, install @aws-sdk/client-bedrock-runtime and send a ConverseCommand with a modelId (or inference profile ID such as us.amazon.nova-lite-v1:0) and a messages array. Read the answer from output.message.content. Use ConverseStreamCommand to stream tokens, and InvokeModelCommand only when you need a model’s native request format.
Amazon Bedrock gives you foundation models from Amazon and other providers behind one AWS API, billed to your AWS account and authorized with IAM. The SDK part is short. What trips people up is everything around it: which model ID to pass, why the same ID works in one Region and fails in another, which IAM resources to allow, and what to do when calls start throttling.
This guide is for Node.js and TypeScript developers adding a model call to an app, a Lambda function or a script. You’ll get a typed module that covers the Bedrock Converse API, streaming and the lower-level InvokeModel call with AWS SDK v3, plus the IAM policy and the fixes for the errors you’re most likely to see.
Converse or InvokeModel: which should you use?
ConverseCommand / ConverseStreamCommand |
InvokeModelCommand / InvokeModelWithResponseStreamCommand |
|
|---|---|---|
| Request body | One shape for every model that supports messages: messages, system, inferenceConfig |
The provider’s own JSON, different for each model family |
| Response | Normalized: output, stopReason, usage, metrics |
Raw bytes you decode and parse per model |
| Switching models | Change the modelId |
Rewrite the body and the parser |
| Model-specific options | Through additionalModelRequestFields |
Directly in the body |
| IAM action | bedrock:InvokeModel (Converse), bedrock:InvokeModelWithResponseStream (ConverseStream) |
The same two actions |
Default to Converse. Use InvokeModel for models that don’t take messages, such as embedding models, or when a provider feature isn’t exposed through Converse.
Prerequisites
- Node.js 18 or later, TypeScript and
tsx, with"type": "module"inpackage.json. - The
@aws-sdk/client-bedrock-runtimepackage. Its source, README and changelog are in the client-bedrock-runtime folder of the AWS SDK for JavaScript v3 repository. - Credentials the SDK can resolve from a profile, SSO or a role; the guide to AWS SDK v3 credential providers for profiles, SSO and assume role covers each.
- A Region where your chosen model is available, and model access for your account (see below).
How to invoke a Bedrock model with AWS SDK v3, step by step
- Create one runtime client
new BedrockRuntimeClient({ retryMode: "adaptive" })at module level. The runtime client (bedrock-runtime) sends prompts; the separate@aws-sdk/client-bedrockclient manages models and profiles. - Pick a model or inference profile IDFor example
amazon.nova-lite-v1:0in a Region that serves it directly, or the US cross-Region profileus.amazon.nova-lite-v1:0. - Build the messagesEach message has a
role(userorassistant) and acontentarray of blocks such as{ text }. Put instructions insystem. - Set limits
inferenceConfig.maxTokenscaps the answer and the output-token cost;temperatureandtopPcontrol randomness. - Send and read the resultJoin the text blocks in
output.message.content, and checkstopReason:max_tokensmeans the answer was cut off. - Record usage
usage.inputTokensandusage.outputTokensare what you pay for. Log them per request.
Example: a typed Bedrock module
// bedrock.ts
// Call Amazon Bedrock models with AWS SDK for JavaScript v3: Converse for a full reply,
// ConverseStream for token-by-token output, and InvokeModel for a model's native request body.
import {
BedrockRuntimeClient,
ConverseCommand,
ConverseStreamCommand,
InvokeModelCommand,
ThrottlingException,
ValidationException,
type Message,
} from "@aws-sdk/client-bedrock-runtime";
// Adaptive retry mode slows the client down when Bedrock returns ThrottlingException.
const bedrock = new BedrockRuntimeClient({ retryMode: "adaptive", maxAttempts: 6 });
// A US cross-Region inference profile for Amazon Nova Lite. The plain model ID
// "amazon.nova-lite-v1:0" also works in Regions that serve the model in-Region.
export const MODEL_ID = process.env.BEDROCK_MODEL_ID ?? "us.amazon.nova-lite-v1:0";
export interface Reply {
text: string;
stopReason: string;
inputTokens: number;
outputTokens: number;
latencyMs: number;
}
/** One request, one complete answer. Pass the running history to keep a conversation. */
export async function ask(history: Message[], system = "You are a concise assistant for AWS engineers."): Promise<Reply> {
try {
const res = await bedrock.send(
new ConverseCommand({
modelId: MODEL_ID,
system: [{ text: system }],
messages: history,
inferenceConfig: { maxTokens: 512, temperature: 0.2 },
}),
);
const text = (res.output?.message?.content ?? []).map((block) => block.text ?? "").join("");
return {
text,
stopReason: res.stopReason ?? "unknown",
inputTokens: res.usage?.inputTokens ?? 0,
outputTokens: res.usage?.outputTokens ?? 0,
latencyMs: res.metrics?.latencyMs ?? 0,
};
} catch (err) {
if (err instanceof ValidationException && /inference profile/i.test(err.message)) {
throw new Error(`${MODEL_ID} needs an inference profile ID here, such as us.${MODEL_ID}`, { cause: err });
}
if (err instanceof ThrottlingException) {
throw new Error("Bedrock is still throttling after retries: lower concurrency or request a quota increase", { cause: err });
}
throw err;
}
}
/** Stream the answer, calling onText for each chunk as the model writes it. */
export async function askStream(prompt: string, onText: (chunk: string) => void): Promise<Reply> {
const res = await bedrock.send(
new ConverseStreamCommand({
modelId: MODEL_ID,
messages: [{ role: "user", content: [{ text: prompt }] }],
inferenceConfig: { maxTokens: 512 },
}),
);
const reply: Reply = { text: "", stopReason: "unknown", inputTokens: 0, outputTokens: 0, latencyMs: 0 };
for await (const event of res.stream ?? []) {
if (event.contentBlockDelta?.delta?.text) {
reply.text += event.contentBlockDelta.delta.text;
onText(event.contentBlockDelta.delta.text);
} else if (event.messageStop) {
reply.stopReason = event.messageStop.stopReason ?? "unknown";
} else if (event.metadata) {
reply.inputTokens = event.metadata.usage?.inputTokens ?? 0;
reply.outputTokens = event.metadata.usage?.outputTokens ?? 0;
reply.latencyMs = event.metadata.metrics?.latencyMs ?? 0;
}
}
return reply;
}
/** InvokeModel sends the provider's own JSON. This body is the Amazon Nova format. */
export async function invokeNova(prompt: string): Promise<string> {
const res = await bedrock.send(
new InvokeModelCommand({
modelId: MODEL_ID,
contentType: "application/json",
accept: "application/json",
body: JSON.stringify({
messages: [{ role: "user", content: [{ text: prompt }] }],
inferenceConfig: { maxTokens: 512 },
}),
}),
);
const json = JSON.parse(new TextDecoder().decode(res.body));
return json.output?.message?.content?.[0]?.text ?? "";
}
A caller that holds a two-turn conversation, streams a second answer and makes one native call:
// chat.ts
// Usage: AWS_PROFILE=dev AWS_REGION=us-east-1 npx tsx chat.ts "Explain what a NAT gateway charges for"
import type { Message } from "@aws-sdk/client-bedrock-runtime";
import { ask, askStream, invokeNova, MODEL_ID } from "./bedrock.js";
const question = process.argv[2] ?? "In two sentences, what does an AWS NAT gateway charge for?";
console.log(`Model: ${MODEL_ID}\n`);
// 1. Converse: a two-turn conversation. Converse is stateless, so you send the whole history each time.
const history: Message[] = [{ role: "user", content: [{ text: question }] }];
const first = await ask(history);
console.log(`Converse -> ${first.text}`);
console.log(` stop=${first.stopReason} in=${first.inputTokens} out=${first.outputTokens} ${first.latencyMs} ms\n`);
history.push({ role: "assistant", content: [{ text: first.text }] });
history.push({ role: "user", content: [{ text: "Give one way to cut that cost." }] });
const second = await ask(history);
console.log(`Follow-up -> ${second.text}\n`);
// 2. ConverseStream: print tokens as they arrive.
process.stdout.write("Stream -> ");
const streamed = await askStream("List three AWS services that bill per request, one per line.", (t) => process.stdout.write(t));
console.log(`\n stop=${streamed.stopReason} out=${streamed.outputTokens}\n`);
// 3. InvokeModel with the model's native body.
console.log(`InvokeModel -> ${await invokeNova("Reply with the single word: ready")}`);
npm install @aws-sdk/client-bedrock-runtime
npm install --save-dev tsx typescript @types/node
npm pkg set type=module
AWS_PROFILE=dev AWS_REGION=us-east-1 npx tsx chat.ts
# Model: us.amazon.nova-lite-v1:0
#
# Converse -> A NAT gateway charges an hourly fee for each hour it's provisioned and a per-GB fee for data it processes. ...
# stop=end_turn in=31 out=48 612 ms
#
# Follow-up -> Route S3 and DynamoDB traffic through gateway VPC endpoints so it bypasses the NAT gateway. ...
#
# Stream -> Amazon S3
# AWS Lambda
# Amazon API Gateway
# stop=end_turn out=12
#
# InvokeModel -> ready
Model output varies from run to run; token counts and latency are illustrative. If stop shows max_tokens, raise maxTokens or ask for a shorter answer.
Model IDs, inference profiles and model access
The modelId field accepts a base model ID, an inference profile ID or ARN, and a few other resource ARNs such as provisioned throughput. Three rules cover most cases:
- Base model IDs like
amazon.nova-lite-v1:0run in the Region you call, when that Region serves the model on demand. - Cross-Region inference profiles add a geography prefix:
us.,eu.,apac., orglobal.for models that support it. Bedrock routes the request to one of the profile’s destination Regions. Forus.amazon.nova-lite-v1:0called from us-east-1, those are us-east-1, us-east-2 and us-west-2. - Some models only run through a profile. Calling their base ID returns a
ValidationExceptionsaying on-demand throughput isn’t supported and to retry with an inference profile. Theask()function above turns that into a clearer error.
Each model’s page in the Bedrock documentation lists its IDs, Regions and whether it supports Converse and streaming. Access to serverless models is enabled by default in commercial Regions. The first call to a third-party model sold through AWS Marketplace starts a subscription in the background, which needs aws-marketplace:Subscribe, aws-marketplace:Unsubscribe and aws-marketplace:ViewSubscriptions for that first caller, and Anthropic models also need a one-time use case form per account or organization. Amazon models such as Nova have no Marketplace product, so they skip that step.
Permissions needed
Scope the policy to the models you use. With a cross-Region profile, allow the profile and the base model in every destination Region:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "InvokeNovaLiteThroughUsProfile",
"Effect": "Allow",
"Action": [
"bedrock:InvokeModel",
"bedrock:InvokeModelWithResponseStream"
],
"Resource": [
"arn:aws:bedrock:us-east-1:123456789012:inference-profile/us.amazon.nova-lite-v1:0",
"arn:aws:bedrock:us-east-1::foundation-model/amazon.nova-lite-v1:0",
"arn:aws:bedrock:us-east-2::foundation-model/amazon.nova-lite-v1:0",
"arn:aws:bedrock:us-west-2::foundation-model/amazon.nova-lite-v1:0"
]
}
]
}
Foundation model ARNs have no account ID. A service control policy that blocks one destination Region makes profile calls fail even if the others are allowed. To confirm the actions your code calls, see how to find the IAM actions your AWS SDK JavaScript code needs, or paste bedrock.ts into the IAM policy generator for TypeScript code. To keep other identities off expensive models, deny bedrock:* on those model ARNs.
Throttling, quotas and cost
Bedrock limits requests and tokens per minute, per model and Region, as account quotas. When you exceed them, calls fail with ThrottlingException. The SDK retries throttling errors on its own; retryMode: "adaptive" also slows the client down after throttles, which suits batch jobs that fire many calls. The guide to configure retries and timeouts in AWS SDK v3 explains attempts and backoff, and monitor AWS service quota usage and get alerts covers watching quotas before they bite. Cross-Region profiles also help, because they spread load across Regions.
On-demand pricing is per input token and per output token, at separate rates that differ from model to model. Check the Amazon Bedrock pricing page for current rates before choosing a model. Record usage from every response; you can publish token counts as CloudWatch custom metrics and alarm on them, and set an AWS budget alert with SDK v3 for the account.
Troubleshooting common Bedrock errors
ValidationException: on-demand throughput isn’t supported. Use the model’s inference profile ID, such asus.plus the model ID.ValidationException: the provided model identifier is invalid. A typo, or the model isn’t offered in this Region. Check the model’s page for Regions.AccessDeniedException. Missingbedrock:InvokeModelon the profile or on a destination Region’s model ARN, a missing Marketplace subscription for a third-party model, or an Anthropic model without the use case form. The steps to troubleshoot IAM access denied errors help separate these.ThrottlingExceptionafter retries. Lower concurrency, spread load with a cross-Region profile, or request a quota increase.ModelTimeoutExceptionor very slow replies. Long prompts and largemaxTokenstake longer. Stream the response, or raise the client’s request timeout for long generations.- Empty text. The model answered with a non-text block, such as a tool use request, or a guardrail intervened. Check
stopReason.
Limits of this approach
The module sends one request at a time and keeps conversation history in memory; a chat product needs to store history, trim it to the model’s context window and summarize old turns. It doesn’t cover tool use, guardrails, prompt caching or batch inference, which all build on the same client. Model output is not deterministic and can be wrong, so validate anything you act on. Don’t put secrets in prompts; load keys from AWS with the guide to get a Secrets Manager secret value with SDK v3 and keep them out of the text you send.
Porting older code? The AWS SDK JavaScript v2 to v3 converter drafts the change, and the guide to migrate a Node.js app from AWS SDK v2 to v3 covers the rest. If you’d rather ask questions about your AWS account than build a model call yourself, ChatWithCloud answers them in plain English from your terminal; how ChatWithCloud works explains which data goes to the model.
Frequently asked questions
What is the difference between Converse and InvokeModel in Bedrock?
Converse uses one request and response shape for every model that supports messages. InvokeModel takes each provider’s native JSON body and returns raw bytes. Both need bedrock:InvokeModel.
How do I stream a Bedrock response in Node.js?
Send ConverseStreamCommand and iterate response.stream with for await. Text arrives in contentBlockDelta.delta.text, the stop reason in messageStop, and token usage in the final metadata event.
Why does my Bedrock model ID work in one Region but not another?
Models aren’t offered in every Region, and some only run through an inference profile. Check the model’s Regional availability, and use a us., eu. or apac. profile ID where needed.
Do I need to request model access in Amazon Bedrock?
Not for most models: access to serverless models is on by default in commercial Regions. Third-party Marketplace models subscribe on first use, which needs Marketplace permissions, and Anthropic models need a one-time use case form.
Related guides
Ask your AWS account in plain English
Your first 15 runs are free, with no OpenAI key needed.
npx chatwithcloud