This guide takes you from a working API key to a production AI application: multi-turn conversations, streaming, tool calling, and error handling.
- New here? Start with the User Guide for auth and quick-starts.
- Looking for schemas? See the API Reference.
Setup
Everything below assumes your OpenRouter API key is in the environment:
export OPENROUTER_API_KEY=sk-or-v1-...
Pick whichever SDK your app already uses — the gateway speaks both dialects:
# Anthropic SDK — bare base URL
from anthropic import Anthropic
client = Anthropic(api_key=os.environ["OPENROUTER_API_KEY"],
base_url="https://api.badlandslabs.com")
# OpenAI SDK — base URL + /v1
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENROUTER_API_KEY"],
base_url="https://api.badlandslabs.com/v1")
// Node — same rule: Anthropic takes the bare URL, OpenAI takes /v1
import Anthropic from "@anthropic-ai/sdk";
const anthropic = new Anthropic({
apiKey: process.env.OPENROUTER_API_KEY,
baseURL: "https://api.badlandslabs.com",
});
import OpenAI from "openai";
const openai = new OpenAI({
apiKey: process.env.OPENROUTER_API_KEY,
baseURL: "https://api.badlandslabs.com/v1",
});
Multi-turn conversations
The API is stateless — your app owns the conversation history. Append each assistant reply to the message list and send the whole list every turn:
messages = []
def chat(user_text: str) -> str:
messages.append({"role": "user", "content": user_text})
res = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
system="You are a concise support assistant for Acme Corp.",
messages=messages,
)
reply = res.content[0].text
messages.append({"role": "assistant", "content": reply})
return reply
print(chat("What plans do you offer?"))
print(chat("Which one did you just recommend?")) # remembers turn 1
Budgeting context: POST /v1/messages/count_tokens preserves an upstream count when
supported. If that route returns 404/405, the gateway substitutes an estimate (~4
chars/token). For guaranteed exact budgets, tokenize client-side before sending.
Streaming
Use streaming when the client requires an SSE contract. The gateway buffers and filters the complete upstream stream before releasing its sanitized events, so this endpoint does not promise first-token latency.
# Anthropic SDK
with client.messages.stream(
model="claude-sonnet-5", max_tokens=1024,
messages=[{"role": "user", "content": "Write a haiku about prairies."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
// OpenAI SDK
const stream = await openai.chat.completions.create({
model: "gpt-4o",
stream: true,
messages: [{ role: "user", content: "Write a haiku about prairies." }],
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
Both dialects emit their standard wire protocol (Anthropic SSE events /
OpenAI data: chunks ending in data: [DONE]), so any SSE client works.
You may see : keep-alive comments while filtering is in progress; they are
spec-compliant and SDKs ignore them automatically.
Tool calling
Define tools, let the model decide when to call them, run the tool, and send the result back. A complete round trip on the Anthropic dialect:
tools = [{
"name": "get_weather",
"description": "Current weather for a city",
"input_schema": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
}]
messages = [{"role": "user", "content": "Should I bike to work in Calgary today?"}]
res = client.messages.create(model="claude-sonnet-5", max_tokens=1024,
tools=tools, messages=messages)
while res.stop_reason == "tool_use":
tool_use = next(b for b in res.content if b.type == "tool_use")
result = get_weather(**tool_use.input) # your function
messages.append({"role": "assistant", "content": res.content})
messages.append({"role": "user", "content": [{
"type": "tool_result",
"tool_use_id": tool_use.id,
"content": str(result),
}]})
res = client.messages.create(model="claude-sonnet-5", max_tokens=1024,
tools=tools, messages=messages)
print(res.content[0].text)
On the OpenAI dialect the same loop reads res.choices[0].message.tool_calls,
runs each call, and appends a {"role": "tool", "tool_call_id": ..., "content": ...}
message. Parallel tool calls, forced tools (tool_choice), and streaming tool
deltas all work in both dialects.
Handling reasoning output
The serving model may be a reasoning model. The gateway sends this OpenAI-dialect option upstream automatically to keep the final-answer channel usable:
{ "reasoning": { "exclude": true } }
The gateway remains the final privacy boundary on both dialects: it removes native
Anthropic thinking/redacted_thinking blocks, <think>/<analysis> text, known
provider/model identity mentions (including those split across SSE chunks), and the
resolved upstream model identity. The response model field remains caller-facing.
Errors, rate limits, retries
Errors arrive in your dialect's native shape, so SDK exception types work as usual. The ones to handle in app code:
| Status | Meaning | App behavior |
|---|---|---|
| 401 | Missing or rejected key | Fix/rotate the key — don't retry |
| 402 | Key out of credits | Surface to the user/billing |
| 429 | Rate limited (Retry-After: 60) |
Back off and retry |
| 5xx | Upstream hiccup | Retry with backoff |
Official SDKs already retry 429/5xx with exponential backoff — tune rather than reimplement:
client = Anthropic(api_key=key, base_url="https://api.badlandslabs.com",
max_retries=4, timeout=120.0)
Respect Retry-After if you roll your own HTTP: sleep the given seconds, then
resend the identical request.
Production checklist
- Keys: load from the environment or a secrets manager; never log them or ship them to browsers. The gateway never stores them either.
- Choose streaming for protocol compatibility, not guaranteed first-token latency; identity filtering buffers up to 8 MiB so split values can be sanitized safely.
- Persist conversation history yourself (DB/session store) — the API is stateless.
- Retries on, idempotently: chat requests are safe to resend on 429/5xx.
- Token budgets: the count may be an estimate after a 404/405 compatibility fallback; tokenize client-side when limits are strict.
- Health checks:
GET https://api.badlandslabs.com/health— no auth, JSON status.