fallbacks field in a request body is removed.
Use fallbacks for agents that must keep answering through a provider outage, or for models that are often overloaded or rate-limited at peak times. They don’t help with requests that are wrong in themselves: a request the provider rejects as invalid, or that a policy blocks, is not sent anywhere else.
Set up fallbacks
1
Add the providers
Each fallback is one of your organization’s providers. Add the ones you need on Providers first (see Supported providers). A provider without a default model can still be a fallback; you choose its model when you add it.
2
Open Model routing
Open the agent, go to the Configuration tab and find Model routing. The primary is at the top, and the fallbacks are listed under Fallbacks, tried in order.
3
Add a fallback
Under Add a fallback, choose a Provider and, if you want another model than the provider’s own, a Model. Choose Add fallback. Repeat for up to three fallbacks.
4
Order them
Move fallbacks up or down with the arrows. The first one is tried first.
5
Choose when to fail over
Under Fail over when the provider, tick the failures that move a request to the next provider (see Which failures fail over). Timeouts and the circuit breaker are under Timeouts and health.
6
Save
Choose Save changes. Xenovia checks every fallback before saving and refuses the whole change if one fails a check (see Save checks).
gpt-5 as the primary and gpt-5-mini as a fallback. That helps when one model is overloaded, not when the whole provider is down, because both use the same account and endpoint.
Choose the order
Each fallback in Model routing is labelled by how close it is to the primary, and warns about capabilities the primary’s model has and it lacks (No image input, No tool calling, Smaller context window). A request that needs a missing capability skips that fallback.- Same model as primary first: the primary’s model on another provider, such as a Claude model served by Anthropic as the primary and the same model on Amazon Bedrock or Google Vertex AI as fallback 1. Answers should match.
- Different model next: another model of the same family.
- Different model family last. Prompts, tool calls and structured output may behave differently, so test the agent on it first. A Simulation run can use another of your providers.
Change or remove the primary
- Change primary keeps the fallbacks, except one with the same provider and model as the new primary, which is removed.
- You can’t detach the primary while the agent has fallbacks. Remove the fallbacks or change the primary first.
- A provider that is a fallback of any agent can’t be deleted until it’s removed from those agents.
Save checks
A save is refused, and nothing changes, when:
Codes without a status are
422. If a saved fallback later becomes unusable (its provider loses its model or credential, it ends up with the same provider and model as another entry, or it no longer matches the primary’s region), the agent’s readiness and Model routing flag it, and the runtime leaves that fallback out until it’s fixed. The rest of the route keeps serving.
Which failures fail over
These never fail over, and reach your application as they would without fallbacks:
- invalid requests:
400,413,422, a prompt too long for the context window, a404that isn’t about the model, and any other4xxexcept429and the codes under Rejects the key or model, which fail over only when you tick it; - Xenovia policy and guardrail blocks, and billing limits;
- a provider’s own content refusal;
- a client that disconnects.
401, 402 or 403 reaches your application as a 502, as without fallbacks, so it’s never mistaken for your own agent key being rejected.
If a fallback answers with an error that doesn’t fail over, such as a 400, that error is the request’s answer and later fallbacks are not tried.
Retries and timeouts
Before moving on, the runtime retries a failed attempt on the same provider at most once: after a short backoff (about half a second) forunreachable, server_error and overloaded, and for rate_limited only when the provider’s Retry-After asks for two seconds or less. Timeouts and rejected keys are not retried on the same provider.
These settings are under Timeouts and health, apart from the failover triggers:
With the defaults, a primary that hangs uses 120 seconds, fallback 1 gets up to 120 seconds, and fallback 2 gets the remaining 60. No attempt after the first starts with less than 10 seconds of the budget left; the fallbacks left are skipped as out of time.
First-token timeout
A provider that accepts a request but never starts answering holds it until the attempt timeout. A first-token timeout notices that early without cutting off long answers. When it’s set, each attempt is streamed from the provider. If no output (text, reasoning, a refusal or a tool call) arrives within that many seconds of the attempt’s start, connecting and queueing included, the attempt is cancelled and counted astimeout. It isn’t retried on the same provider. It fails over when Times out is ticked; otherwise your application gets a 504 that names the first-token timeout. Once output has started, only the attempt timeout and the total budget apply.
Your application sees no difference: the runtime still collects the whole answer, runs response policies on it and returns it the way you asked for it, streamed or not.
Use it for interactive agents, where waiting on a stalled provider costs more than switching. Set it well above the model’s usual time to first token; traces show Time to first token for attempts made under it.
- A first-token timeout counts as a failure of that provider for the circuit breaker below, which every agent using the provider shares. A timeout too tight for a slow model can make other agents skip that provider.
- It doesn’t apply to requests that ask for log probabilities, several choices (
nabove 1), audio output or a background Responses run. Those are still sent as one call.
Skipping a provider that keeps failing
The runtime keeps a circuit breaker for each provider (each runtime instance keeps its own). After five outage failures within 30 seconds, with no successful answer between them, the breaker opens for about a minute, or longer (up to five minutes) when the provider’s429 or 503 asks for it with Retry-After. Then one request is let through to test the provider: if it succeeds, the breaker closes; if not, it opens again. Failures from every agent that uses the provider count.
With Skip a provider that keeps failing on (the default), a provider whose breaker is open, the primary included, is tried last instead of in its place, so requests go straight to the next provider instead of waiting on one that is down. If every other provider fails, it’s still tried. The first fallback tried in its place reports the reason breaker_open.
Fallbacks that are left out
A fallback is skipped for a request, and the next one tried, when:
Capabilities come from a public model directory. A model it doesn’t know is treated as capable. Capability checks never skip the primary, and a request that no fallback could take, because of capabilities or provider-side state, runs on the primary exactly as it would without fallbacks.
Governance
Policies
Fallbacks are held to the agent’s policies twice. When you save. Each fallback’s provider and model are checked against the agent’s enforced policies made from the Approved provider allowlist, Approved model allowlist and Model to provider route lock templates, and against the organization’s enforced baselines made from them. A fallback one of them would block can’t be saved. Policies you wrote yourself, and template policies whose Rego was edited, are not checked at save time. On every attempt. Before calling a fallback, the runtime runs the agent’s request-stage policies again with that fallback’sprovider, provider_adapter, provider_host, model_family and model. If a policy blocks or escalates it, or the evaluation fails, that fallback is skipped and the request moves on to the next one; the request itself is not refused. A redaction applies to that attempt only. These re-checks show in the trace but are not the request’s policy decision: they are left out of block rates, policy alerts and compliance evidence.
Guardrails run once, before the first attempt. Their verdict, and every redaction made before the first attempt, stands for every attempt, so no fallback receives content Xenovia removed.
Policies of agents with fallbacks receive two more fields:
The request stage sees
1 and false (the primary), a fallback’s re-check sees its own attempt and route, and the response stage sees the attempt that answered. To tell a fallback apart, test route_is_fallback, not route_attempt: a fallback tried first because the primary’s breaker is open is attempt 1.
For example, this request-stage policy lets fallbacks use only the organization’s cloud accounts:
Residency
Turn on Keep fallbacks in the primary’s region to refuse fallbacks that serve requests from a different region than the primary. Xenovia detects the region of Amazon Bedrock providers from their AWS region (a model ID startingglobal. counts as Global) and of Google Vertex AI providers from their location. For any other provider, set Data residency (US, EU, Asia Pacific or Global) on the provider in Providers. A provider whose region is unknown never matches, not even another unknown one.
The setting is checked when you save the route. A fallback that stops matching later, because a provider’s region changed, is left out until it’s fixed.
Audit log
Every change to an agent’s fallbacks or routing settings is recorded in the audit log asagent.fallbacks.update, with the route and settings before and after.
What your application sees
Response headers
A request runs through the agent’s fallback route when the agent has at least one fallback that could take it. Every such request gets these headers once a provider was called, on errors as well as successes, including a request the primary answered straight away (
X-Xenovia-Attempts: 1, no X-Xenovia-Fallback-Reason). Browsers can read them (they’re exposed through CORS).
These requests don’t run through the route, and carry none of the headers even though the primary was called:
- requests of agents without fallbacks;
- requests sent with
X-Xenovia-Fallback: off; - requests no fallback could take, because of capabilities or provider-side state (see Fallbacks that are left out);
- Simulation runs on another provider than the agent’s primary, which that provider serves alone.
When every provider fails
A request’s answer is the first success. Only when every provider tried failed with an error the agent fails over on does the runtime return503:
error.code. The body also carries extra_fields, as other runtime errors do. OpenAI SDKs raise a 503 as InternalServerError and retry it twice by default, which runs the whole route again each time; lower max_retries if that takes too long. Agents without fallbacks keep their usual errors.
Turn fallbacks off for one request
SendX-Xenovia-Fallback: off to run a request on the primary only, with the usual single-provider retries. Use it for evaluations that must measure one fixed model.
Monitor fallbacks
Traces
In Sessions, a turn’s model call shows how the route went. Each failed attempt is its own Failed attempt step with its provider, status and failure class. The model call that answered says which attempt it was, why the request moved on and, when a fallback answered, Served by fallback 1, 2 or 3. Skipped fallbacks are listed with their reason. A failed attempt the runtime recovered from doesn’t count as a failed request in overviews. The traces API returnsprovider_id, route_attempt, route_is_fallback, fallback_reason and primary_provider_id on model runs, and failed attempts have run_type llm_attempt.
Routing panel
The agent’s Overview tab has a Routing panel for the last hour, 24 hours or 7 days: the Fallback rate, the requests Served by a fallback, the requests where Every provider failed, requests and failed attempts per provider, and Why requests failed over. The same numbers come fromGET /api/v1/agents/{agent_id}/routing-stats?window=24h (1h, 24h or 7d).
Provider health
On Providers, each provider shows its health over the last 15 minutes:
A failure is an attempt that failed with one of the six classes above, or a request that ended in
fallbacks_exhausted. That includes Rejects the key or model: a rejected key or missing model counts against the provider’s health even while it doesn’t fail over. For requests that didn’t run through a fallback route, a failure is a 429, 500, 502, 503, 504 or 529 from the provider. The same data comes from GET /api/v1/providers/health?window=15m (15m, 1h or 24h).
Alerts
Create these in Settings → Alerts:
Run alert rules can also filter on
provider_id. Run error doesn’t fire for failed attempts the runtime moved on from; the request’s own model call reports how it ended.
Manage fallbacks with the API
PUT /api/v1/agents/{agent_id}/fallbacks replaces an agent’s whole list of fallbacks, and optionally its routing settings. It takes an organization key with write scope (see Authentication) or a signed-in session.
- The order of
fallbacksis the order they’re tried. An empty list removes every fallback. modelis optional;nulluses the provider’s own model.routingis optional. Leave it out to keep the saved settings, send some fields to change only those, or sendnullto reset every setting to its default. Fields are those in Retries and timeouts; onlyfirst_token_timeout_smay benull(off).
fallbacks (each with its position, provider_id, effective model, model_override, residency and the model’s capabilities) and the effective routing. A refused save answers with a detail object holding a code from Save checks and a message, and changes nothing.
To read the route, GET /api/v1/agents/{agent_id} returns fallbacks and routing, and GET /api/v1/agents/{agent_id}/providers lists the primary (role: primary, position: 0) followed by the fallbacks. Providers accept a residency field (us, eu, apac or global; null to detect it).
Costs
- Xenovia counts a request once against your organization’s request allowance, however many attempts it took. A request that ends in
fallbacks_exhaustedisn’t counted. - Xenovia’s LLM cost figures price the call that answered. Failed attempts carry no tokens.
- Providers may still bill for failed attempts. A provider can finish, and charge for, a generation after Xenovia gave up on it at a timeout. With a first-token timeout, output streamed before a provider error may be billed, twice if the attempt was retried on the same provider.
- Prompt caches don’t carry over. A fallback starts with its own cache, so long cached prompts cost more while the agent is failing over, and a fallback’s model has its own price.
Limitations
- Hybrid deployments. A customer-hosted runtime can’t load fallback credentials yet, so it skips every fallback (couldn’t be set up) and serves from the primary.
- Requests can’t set fallbacks. A
fallbacksfield in a request body is removed; only the agent’s route is used. - Provider-side state. A request that refers to state held by the primary’s provider, such as a Responses
previous_response_idorconversation, or a provider file ID, can only fail over to a fallback on that same provider with another model. Other fallbacks are skipped. - Up to three fallbacks per agent.