// model gateway LIVE

OpenRouter

One key, a whole catalogue of models, paid per token

OpenRouter lets Fulltrace use models from many providers through one account, paid per token. Task runs use it for every stage, with backup models when one fails, and Studio estimates a run's cost from past runs before it starts. The chat in Studio uses it too, on a key of its own.

03 separate keys 03 role profiles 6h catalogue cache

The Right Model for Each Job

Fulltrace sends each stage of a job to the cheapest model that does it well, keeps repeatable work in plain code, and saves the expensive thinking for where it matters. That only works if switching models is easy. OpenRouter makes it easy, with one endpoint, one key and one bill for the whole catalogue.

Each stage has a route: a preferred model, then backups. If a provider is down or busy, the job moves to the next model and the attempt is recorded. Fix proposals work differently. They use Claude Sonnet 4.5 in a single call, and a network error is retried on that same model. A proposal that is rejected stays rejected, so a weaker model never gets a second try at the same fix. The full story is on how it works, and the Codex profiles that use OpenRouter are on the codex page.

The key follows the secrets rule (G10). OPENROUTER_API_KEY lives in a user environment variable and is read when it is needed. It is never stored in the repo, in config, in queued jobs or in reports.

Where it plugs in
01Codex role profiles: three jobs, three models
02Runtime model routes with failover
03The shared model picker in Studio
04Per-model cost tracking and pricing
05The Chat tab in Studio, on a key of its own
Status

Live and in daily use. The backup routes are active, the catalogue picker is live in Studio, and the Codex sync script deploys the three role profiles.

The Whole Catalogue in One Picker

One shared picker replaced every scattered model box in Studio. It is a searchable, grouped list over OpenRouter's live catalogue, used for the fix-proposal model and for every stage and run override. The logic behind it has its own tests, and it is careful about price: only a model priced at exactly zero is marked free. Some pickers also pin sections for my subscriptions and for models running on my own computer.

model_picker_internalsgroups, fallback layers, endpoint
Curated groups
Claude, Codex/GPT, Other

A mapping in code sorts OpenRouter's providers into three groups, and the current choice is pinned at the top under its own heading.

Never empty
Four layers of fallback

A copy less than six hours old is used first. Otherwise it fetches the live list, then falls back to an older copy, and last to a short built-in list. Offline or mid-outage, the picker still shows a list.

Nothing new to secure
Read-only model details

The extension fetches the catalogue from a public address. It needs no API key, adds no new server endpoint and grants no new power. The picker only writes settings that could already be written.

Estimated First, Counted After

Before I confirm a run, Studio estimates its cost from past runs of the same workflow. One shared module prices Anthropic, OpenAI and OpenRouter models alike, and counts every token against the model that did the work (Studio shows it). One rule holds throughout, G1: the free models, which have tight rate limits, never run a job that checks or changes code. The Chat tab has a third key of its own. The key that runs audits and fixes cannot be used for chat, and the read-only key used for analytics cannot send a message.

Attribution
Every token accounted for

Studio shows usage and cost for every model, per session and all time, with OpenRouter models listed beside Claude. Building it turned up a real bug: cached and reasoning tokens were being counted twice, when they are already part of the total. That is fixed.

Balance
What is left, live

Studio also shows how much OpenRouter credit is left, kept apart from what the subscription plans cover, so I can see there is credit before I launch.