Getting started
GlobalRouter speaks the native Anthropic Messages API and an OpenAI-compatible endpoint. In most tools you only change a base URL and an API key.
1. Create an API key
Sign up, then create a key from your dashboard. You can name each key, set a daily spend cap, and disable it at any time.
2. Point your tool at us
Claude Code
export ANTHROPIC_BASE_URL=https://globalrouterai.com
export ANTHROPIC_AUTH_TOKEN=sk-your-keyReload your shell and run claude as usual.
Cursor
Settings → Models → Override OpenAI Base URL. Use https://globalrouterai.com/v1 and paste your key.
Cline / Roo Code
Choose Anthropic as the provider, set the custom base URL to https://globalrouterai.com, and paste your key.
Anthropic SDK
from anthropic import Anthropic
client = Anthropic(
base_url="https://globalrouterai.com",
api_key="sk-your-key",
)
msg = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
)
print(msg.content[0].text)OpenAI SDK
import OpenAI from "openai"
const client = new OpenAI({
baseURL: "https://globalrouterai.com/v1",
apiKey: "sk-your-key",
})
const r = await client.chat.completions.create({
model: "claude-sonnet-5",
messages: [{ role: "user", content: "Hello" }],
})3. Keep prompt caching on
Caching is supported exactly as upstream: mark a stable prefix with cache_control and cache reads bill at one tenth of input. Coding agents resend a large system prompt every turn, so this is usually where most of your bill goes — and where most of your savings come from.
client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
system=[{
"type": "text",
"text": LARGE_STABLE_PROMPT,
"cache_control": {"type": "ephemeral"},
}],
messages=[{"role": "user", "content": question}],
)Check usage.cache_read_input_tokens on the response to confirm you are getting hits.
Available models
| claude-sonnet-5 | Claude Sonnet 5 |
| claude-opus-4-8 | Claude Opus 4.8 |
| claude-fable-5 | Claude Fable 5 |
| claude-sonnet-4-6 | Claude Sonnet 4.6 |
| claude-haiku-4-5 | Claude Haiku 4.5 |
| gpt-5.6-terra | GPT-5.6 Terra |
| gpt-5.6-sol | GPT-5.6 Sol |
| gpt-6-astra | GPT-6 Astra |
| gpt-5.5 | GPT-5.5 |
| glm-5.3 | GLM-5.3 |
| glm-5.3-flash | GLM-5.3 Flash |
| glm-5.2 | GLM-5.2 |
| glm-5.1 | GLM-5.1 |
| glm-5 | GLM-5 |
How this differs from a first-party API
We route your requests through upstream providers rather than holding a direct commercial relationship with the model vendor. That has two practical consequences worth knowing before you build on it.
System-level instructions
Requests carry additional system-level context from the upstream provider. In practice the model may describe itself as a coding assistant, and a custom persona set in your own system prompt may not be fully adopted. If your product depends on a specific persona — or on the model denying that it has tools — test that behavior before relying on it.
What is unaffected
Tool definitions you send are respected: the model will not invoke tools you did not define. Streaming, prompt caching, token accounting, stop reasons, and the Messages API surface all behave as documented.
Who this suits
Coding agents — Claude Code, Cursor, Cline — where the model is meant to be a coding assistant anyway, see no practical difference. General-purpose assistants, roleplay, and persona-driven products should evaluate carefully; this is not what the service is tuned for.
Every account starts with a small balance you control, so you can test your own workload before committing. Unused credits are refundable in full within 30 days.
Errors and capacity
When upstream capacity is constrained we return 503 with a Retry-After header rather than failing silently. Rate limits return 429. If your balance runs out you get 402. Live capacity is published on our status page.
GlobalRouter is an independent third-party API gateway. We are not affiliated with, endorsed by, or sponsored by Anthropic. Requests are routed through upstream providers, and model behavior — including system-level instructions and how the model describes itself — can differ from a first-party API. The service is built and tuned for coding agent workloads; evaluate it for your use case before relying on it.