Getting started

GlobalRouter speaks the native Anthropic Messages API and an OpenAI-compatible endpoint. In most tools you only change a base URL and an API key.

1. Create an API key

Sign up, then create a key from your dashboard. You can name each key, set a daily spend cap, and disable it at any time.

2. Point your tool at us

Claude Code

~/.zshrc
export ANTHROPIC_BASE_URL=https://globalrouterai.com
export ANTHROPIC_AUTH_TOKEN=sk-your-key

Reload your shell and run claude as usual.

Cursor

Settings → Models → Override OpenAI Base URL. Use https://globalrouterai.com/v1 and paste your key.

Cline / Roo Code

Choose Anthropic as the provider, set the custom base URL to https://globalrouterai.com, and paste your key.

Anthropic SDK

python
from anthropic import Anthropic

client = Anthropic(
    base_url="https://globalrouterai.com",
    api_key="sk-your-key",
)

msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}],
)
print(msg.content[0].text)

OpenAI SDK

typescript
import OpenAI from "openai"

const client = new OpenAI({
  baseURL: "https://globalrouterai.com/v1",
  apiKey: "sk-your-key",
})

const r = await client.chat.completions.create({
  model: "claude-sonnet-5",
  messages: [{ role: "user", content: "Hello" }],
})

3. Keep prompt caching on

Caching is supported exactly as upstream: mark a stable prefix with cache_control and cache reads bill at one tenth of input. Coding agents resend a large system prompt every turn, so this is usually where most of your bill goes — and where most of your savings come from.

python
client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    system=[{
        "type": "text",
        "text": LARGE_STABLE_PROMPT,
        "cache_control": {"type": "ephemeral"},
    }],
    messages=[{"role": "user", "content": question}],
)

Check usage.cache_read_input_tokens on the response to confirm you are getting hits.

Available models

claude-sonnet-5Claude Sonnet 5
claude-opus-4-8Claude Opus 4.8
claude-fable-5Claude Fable 5
claude-sonnet-4-6Claude Sonnet 4.6
claude-haiku-4-5Claude Haiku 4.5
gpt-5.6-terraGPT-5.6 Terra
gpt-5.6-solGPT-5.6 Sol
gpt-6-astraGPT-6 Astra
gpt-5.5GPT-5.5
glm-5.3GLM-5.3
glm-5.3-flashGLM-5.3 Flash
glm-5.2GLM-5.2
glm-5.1GLM-5.1
glm-5GLM-5

How this differs from a first-party API

We route your requests through upstream providers rather than holding a direct commercial relationship with the model vendor. That has two practical consequences worth knowing before you build on it.

System-level instructions

Requests carry additional system-level context from the upstream provider. In practice the model may describe itself as a coding assistant, and a custom persona set in your own system prompt may not be fully adopted. If your product depends on a specific persona — or on the model denying that it has tools — test that behavior before relying on it.

What is unaffected

Tool definitions you send are respected: the model will not invoke tools you did not define. Streaming, prompt caching, token accounting, stop reasons, and the Messages API surface all behave as documented.

Who this suits

Coding agents — Claude Code, Cursor, Cline — where the model is meant to be a coding assistant anyway, see no practical difference. General-purpose assistants, roleplay, and persona-driven products should evaluate carefully; this is not what the service is tuned for.

Every account starts with a small balance you control, so you can test your own workload before committing. Unused credits are refundable in full within 30 days.

Errors and capacity

When upstream capacity is constrained we return 503 with a Retry-After header rather than failing silently. Rate limits return 429. If your balance runs out you get 402. Live capacity is published on our status page.

GlobalRouter is an independent third-party API gateway. We are not affiliated with, endorsed by, or sponsored by Anthropic. Requests are routed through upstream providers, and model behavior — including system-level instructions and how the model describes itself — can differ from a first-party API. The service is built and tuned for coding agent workloads; evaluate it for your use case before relying on it.