Docs

OffRouter speaks the OpenAI chat completions API. If your code works with an OpenAI SDK, it works here after two changes: the base URL and the key.

Quickstart

Base URL: /v1. There are two ways to pay for a call, and both reach the same models at the same price.

  • Prepaid key. Deposit once, get a key that starts with off_sk_, and use it as a normal bearer token. Best for anything built on an OpenAI SDK.
  • Pay per call. Send no key. The gateway replies 402 with a price and an x402 client pays it. Best for autonomous agents that hold their own wallet.
curl /v1/chat/completions \
  -H "Authorization: Bearer $OFFROUTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "deepseek-v4-flash", "stream": true,
       "messages": [{"role": "user", "content": "Hello"}]}'

List the models with GET /v1/models, or see them on the models page.

Prepaid keys

A key is created by a deposit and spends down per call. There are no accounts: the key is the account, and the wallet that paid for it is its owner. The easiest way to make one is the dashboard. The same thing over the API:

POST   /v1/keys         {"amount_usd": 5}   # x402 deposit, returns the key once
POST   /v1/keys/topup   {"amount_usd": 5}   # x402 deposit + Bearer, adds balance
GET    /v1/keys/me                          # balance and recent usage
DELETE /v1/keys/me                          # revoke, refund what is left

How a keyed call is charged

Before the call is served, the worst case (your full input plus max_tokens of output) is held against the balance. After it finishes, the hold is settled down to the real token count. A key can never spend more than it holds, even with many calls at once.

If you leave max_tokens out it defaults to 8192. If the balance cannot cover the budget you asked for, the budget is lowered to what it can.

The charge for each call comes back in the X-OffRouter-Charge-USD response header.

Pay per call

With no key, a request is paid for on the spot using x402. The flow is two requests:

POST /v1/chat/completions                       # 402, price in PAYMENT-REQUIRED
POST /v1/chat/completions + PAYMENT-SIGNATURE   # settles, then 200 + PAYMENT-RESPONSE

max_tokens is required here. The price in the 402 is a ceiling: your full input plus max_tokens of output. Your client signs that amount, it settles onchain, and the call is served only after settlement.

Refunds

When the call ends, the real charge is worked out from actual usage. The difference between what you signed and what you used is queued as a refund to the paying wallet, on the network you paid on. A failed call or an empty completion is refunded in full. Small refunds are batched and sent together.

Refunds mean extra onchain transfers. If you make many calls, a prepaid key is cheaper to run and leaves a smaller onchain trail: one deposit instead of a payment and a refund per call.

Client

import { privateKeyToAccount } from "viem/accounts";
import { x402Client, wrapFetchWithPayment } from "@x402/fetch";
import { ExactEvmScheme } from "@x402/evm/exact/client";

const account = privateKeyToAccount(process.env.AGENT_KEY);
const client = x402Client.fromConfig({
  schemes: [{ network: "eip155:4663", client: new ExactEvmScheme(account) }],
});
const pay = wrapFetchWithPayment(fetch, client);

const res = await pay("/v1/chat/completions", {
  method: "POST",
  headers: { "content-type": "application/json" },
  body: JSON.stringify({
    model: "gpt-oss-120b",
    max_tokens: 512,
    messages: [{ role: "user", content: "Hello" }],
  }),
});

A machine-readable description of every paid endpoint is at /.well-known/x402.

Networks

Every 402 lists each network below in accepts[] at the same amount. Pay on whichever one your wallet holds a balance on. Each network settles in its own dollar stablecoin, all with 6 decimals. Payments are a signed transfer authorization (EIP-3009), so the payer needs no gas token.

NetworkCAIP-2AssetPaid fromAsset contract
Loading from the gateway…

This table is read live from GET /networks. A deposit is refunded on the network it arrived on.

Arc and Circle Gateway

On Arc, payments go through Circle Gateway. You pay from a Gateway balance, not from the USDC sitting in your wallet, so deposit into Gateway once before your first call. The 402 entry for Arc names Circle's GatewayWalletBatched contract in extra, and the authorization must stay valid for 7 days because Circle settles in batches.

Circle's client handles all of that: new GatewayClient({ chain: "arc", privateKey }) from @circle-fin/x402-batching/client, then client.deposit("5") once and client.pay(url, options) per call. A plain x402 client works unchanged on Base and Robinhood Chain.

Receipts

Every model runs in a GPU enclave (Intel TDX with a confidential NVIDIA GPU) on hardware we do not operate. Each response carries three headers so you can check that for yourself:

  • X-OffRouter-Receipt: the receipt id. It equals the completion id.
  • X-OffRouter-Receipt-URL: where to fetch the signed receipt, GET /v1/private/receipts/:id. It holds hashes and timestamps, never content, and needs no key.
  • X-OffRouter-Attestation: the enclave's public attestation report. The key that signs the receipt is published in it.

We cannot forge a receipt: it is signed by a key that only exists inside the attested enclave.

What this does and does not cover

The enclave operator cannot read your prompts or outputs. The OffRouter gateway does see the request in memory while forwarding it, and does not store it. Removing the gateway from that path, by encrypting on the client to the enclave key, is planned and not available yet.

MCP server

POST /mcp is a remote MCP server. There is nothing to install: add the URL with a prepaid key as the bearer.

claude mcp add --transport http offrouter /mcp \
  --header "Authorization: Bearer $OFFROUTER_KEY"

Tools: list_models (no key needed), ask, compare (the same prompt to several models) and balance. Tool calls are billed exactly like direct API calls. Pay per call is not available over MCP, because a remote server cannot sign for your wallet.

Errors

API errors use the OpenAI shape, {"error": {"message", "type", "code"}}. Payment problems use the x402 shape with a code.

StatusCodeWhat to do
400max_tokens_requiredSet max_tokens on pay-per-call requests.
401invalid_api_keyCheck the key. Closed keys stop working at once.
402insufficient_balanceTop up the key or lower max_tokens.
402terms_mismatchSign exactly the amount, asset and recipient from the 402.
402payment_replayedThat payment was already used. Sign a new one.
404model_not_foundUse an id from GET /v1/models.
429model_overloadedRetry in a few seconds.
502upstream_errorRetry shortly. A paid call that fails is refunded.