Skip to main content
An agent sends its model requests to its primary provider first. Fallback providers are up to three more of your organization’s providers, tried in order when the primary fails with an outage: it can’t be reached, it times out, it rate-limits the request, or it answers with a server error or an overload. Your application keeps calling the same agent endpoint. The runtime switches providers behind it and returns one answer. While the primary keeps failing, requests try the fallbacks first and the primary last. Fallbacks are set on the agent, in the console or with the management API. A request can’t choose its own: a fallbacks field in a request body is removed. Use fallbacks for agents that must keep answering through a provider outage, or for models that are often overloaded or rate-limited at peak times. They don’t help with requests that are wrong in themselves: a request the provider rejects as invalid, or that a policy blocks, is not sent anywhere else.

Set up fallbacks

1

Add the providers

Each fallback is one of your organization’s providers. Add the ones you need on Providers first (see Supported providers). A provider without a default model can still be a fallback; you choose its model when you add it.
2

Open Model routing

Open the agent, go to the Configuration tab and find Model routing. The primary is at the top, and the fallbacks are listed under Fallbacks, tried in order.
3

Add a fallback

Under Add a fallback, choose a Provider and, if you want another model than the provider’s own, a Model. Choose Add fallback. Repeat for up to three fallbacks.
4

Order them

Move fallbacks up or down with the arrows. The first one is tried first.
5

Choose when to fail over

Under Fail over when the provider, tick the failures that move a request to the next provider (see Which failures fail over). Timeouts and the circuit breaker are under Timeouts and health.
6

Save

Choose Save changes. Xenovia checks every fallback before saving and refuses the whole change if one fails a check (see Save checks).
You can also add fallbacks when you create an agent: the model step has an optional Fallbacks field, which offers providers that have a default model. Set the order and failover settings on the Configuration tab afterward. The same provider can appear more than once in the route with different models, for example an OpenAI provider serving gpt-5 as the primary and gpt-5-mini as a fallback. That helps when one model is overloaded, not when the whole provider is down, because both use the same account and endpoint.

Choose the order

Each fallback in Model routing is labelled by how close it is to the primary, and warns about capabilities the primary’s model has and it lacks (No image input, No tool calling, Smaller context window). A request that needs a missing capability skips that fallback.
  1. Same model as primary first: the primary’s model on another provider, such as a Claude model served by Anthropic as the primary and the same model on Amazon Bedrock or Google Vertex AI as fallback 1. Answers should match.
  2. Different model next: another model of the same family.
  3. Different model family last. Prompts, tool calls and structured output may behave differently, so test the agent on it first. A Simulation run can use another of your providers.

Change or remove the primary

  • Change primary keeps the fallbacks, except one with the same provider and model as the new primary, which is removed.
  • You can’t detach the primary while the agent has fallbacks. Remove the fallbacks or change the primary first.
  • A provider that is a fallback of any agent can’t be deleted until it’s removed from those agents.

Save checks

A save is refused, and nothing changes, when: Codes without a status are 422. If a saved fallback later becomes unusable (its provider loses its model or credential, it ends up with the same provider and model as another entry, or it no longer matches the primary’s region), the agent’s readiness and Model routing flag it, and the runtime leaves that fallback out until it’s fixed. The rest of the route keeps serving.

Which failures fail over

These never fail over, and reach your application as they would without fallbacks:
  • invalid requests: 400, 413, 422, a prompt too long for the context window, a 404 that isn’t about the model, and any other 4xx except 429 and the codes under Rejects the key or model, which fail over only when you tick it;
  • Xenovia policy and guardrail blocks, and billing limits;
  • a provider’s own content refusal;
  • a client that disconnects.
Rejects the key or model is off by default because a revoked key or a missing model is a configuration problem, not an outage: failing over keeps the agent answering but hides the problem. If you turn it on, add the Request served by a fallback alert (see Alerts) so you notice. While it’s off, a provider’s 401, 402 or 403 reaches your application as a 502, as without fallbacks, so it’s never mistaken for your own agent key being rejected. If a fallback answers with an error that doesn’t fail over, such as a 400, that error is the request’s answer and later fallbacks are not tried.

Retries and timeouts

Before moving on, the runtime retries a failed attempt on the same provider at most once: after a short backoff (about half a second) for unreachable, server_error and overloaded, and for rate_limited only when the provider’s Retry-After asks for two seconds or less. Timeouts and rejected keys are not retried on the same provider. These settings are under Timeouts and health, apart from the failover triggers: With the defaults, a primary that hangs uses 120 seconds, fallback 1 gets up to 120 seconds, and fallback 2 gets the remaining 60. No attempt after the first starts with less than 10 seconds of the budget left; the fallbacks left are skipped as out of time.
Adding a fallback changes the agent’s timeouts. Without fallbacks, a call to the provider can run for 300 seconds, with up to two retries. With fallbacks, each attempt ends at the attempt timeout, 120 seconds by default. If the agent writes long answers or uses a reasoning model that can think for minutes, raise Attempt timeout and Total budget before you add fallbacks, or long requests will time out and fail over. Keep your HTTP client’s timeout above the total budget, or your application gives up while a fallback is still answering.

First-token timeout

A provider that accepts a request but never starts answering holds it until the attempt timeout. A first-token timeout notices that early without cutting off long answers. When it’s set, each attempt is streamed from the provider. If no output (text, reasoning, a refusal or a tool call) arrives within that many seconds of the attempt’s start, connecting and queueing included, the attempt is cancelled and counted as timeout. It isn’t retried on the same provider. It fails over when Times out is ticked; otherwise your application gets a 504 that names the first-token timeout. Once output has started, only the attempt timeout and the total budget apply. Your application sees no difference: the runtime still collects the whole answer, runs response policies on it and returns it the way you asked for it, streamed or not. Use it for interactive agents, where waiting on a stalled provider costs more than switching. Set it well above the model’s usual time to first token; traces show Time to first token for attempts made under it.
A reasoning model that thinks without streaming its thoughts sends nothing until it’s done thinking, and that time counts against the first-token timeout. Leave it off for such models, or set it above their longest thinking time.
  • A first-token timeout counts as a failure of that provider for the circuit breaker below, which every agent using the provider shares. A timeout too tight for a slow model can make other agents skip that provider.
  • It doesn’t apply to requests that ask for log probabilities, several choices (n above 1), audio output or a background Responses run. Those are still sent as one call.

Skipping a provider that keeps failing

The runtime keeps a circuit breaker for each provider (each runtime instance keeps its own). After five outage failures within 30 seconds, with no successful answer between them, the breaker opens for about a minute, or longer (up to five minutes) when the provider’s 429 or 503 asks for it with Retry-After. Then one request is let through to test the provider: if it succeeds, the breaker closes; if not, it opens again. Failures from every agent that uses the provider count. With Skip a provider that keeps failing on (the default), a provider whose breaker is open, the primary included, is tried last instead of in its place, so requests go straight to the next provider instead of waiting on one that is down. If every other provider fails, it’s still tried. The first fallback tried in its place reports the reason breaker_open.

Fallbacks that are left out

A fallback is skipped for a request, and the next one tried, when: Capabilities come from a public model directory. A model it doesn’t know is treated as capable. Capability checks never skip the primary, and a request that no fallback could take, because of capabilities or provider-side state, runs on the primary exactly as it would without fallbacks.

Governance

Policies

Fallbacks are held to the agent’s policies twice. When you save. Each fallback’s provider and model are checked against the agent’s enforced policies made from the Approved provider allowlist, Approved model allowlist and Model to provider route lock templates, and against the organization’s enforced baselines made from them. A fallback one of them would block can’t be saved. Policies you wrote yourself, and template policies whose Rego was edited, are not checked at save time. On every attempt. Before calling a fallback, the runtime runs the agent’s request-stage policies again with that fallback’s provider, provider_adapter, provider_host, model_family and model. If a policy blocks or escalates it, or the evaluation fails, that fallback is skipped and the request moves on to the next one; the request itself is not refused. A redaction applies to that attempt only. These re-checks show in the trace but are not the request’s policy decision: they are left out of block rates, policy alerts and compliance evidence. Guardrails run once, before the first attempt. Their verdict, and every redaction made before the first attempt, stands for every attempt, so no fallback receives content Xenovia removed. Policies of agents with fallbacks receive two more fields: The request stage sees 1 and false (the primary), a fallback’s re-check sees its own attempt and route, and the response stage sees the attempt that answered. To tell a fallback apart, test route_is_fallback, not route_attempt: a fallback tried first because the primary’s breaker is open is attempt 1. For example, this request-stage policy lets fallbacks use only the organization’s cloud accounts:
A fallback on any other provider is skipped and shows as blocked by policy in the trace. The primary is never affected. See Policies and Enforcement for the decision contract.

Residency

Turn on Keep fallbacks in the primary’s region to refuse fallbacks that serve requests from a different region than the primary. Xenovia detects the region of Amazon Bedrock providers from their AWS region (a model ID starting global. counts as Global) and of Google Vertex AI providers from their location. For any other provider, set Data residency (US, EU, Asia Pacific or Global) on the provider in Providers. A provider whose region is unknown never matches, not even another unknown one. The setting is checked when you save the route. A fallback that stops matching later, because a provider’s region changed, is left out until it’s fixed.

Audit log

Every change to an agent’s fallbacks or routing settings is recorded in the audit log as agent.fallbacks.update, with the route and settings before and after.

What your application sees

Response headers

A request runs through the agent’s fallback route when the agent has at least one fallback that could take it. Every such request gets these headers once a provider was called, on errors as well as successes, including a request the primary answered straight away (X-Xenovia-Attempts: 1, no X-Xenovia-Fallback-Reason). Browsers can read them (they’re exposed through CORS). These requests don’t run through the route, and carry none of the headers even though the primary was called:
  • requests of agents without fallbacks;
  • requests sent with X-Xenovia-Fallback: off;
  • requests no fallback could take, because of capabilities or provider-side state (see Fallbacks that are left out);
  • Simulation runs on another provider than the agent’s primary, which that provider serves alone.
With the OpenAI Python SDK:

When every provider fails

A request’s answer is the first success. Only when every provider tried failed with an error the agent fails over on does the runtime return 503:
Match on error.code. The body also carries extra_fields, as other runtime errors do. OpenAI SDKs raise a 503 as InternalServerError and retry it twice by default, which runs the whole route again each time; lower max_retries if that takes too long. Agents without fallbacks keep their usual errors.

Turn fallbacks off for one request

Send X-Xenovia-Fallback: off to run a request on the primary only, with the usual single-provider retries. Use it for evaluations that must measure one fixed model.

Monitor fallbacks

Traces

In Sessions, a turn’s model call shows how the route went. Each failed attempt is its own Failed attempt step with its provider, status and failure class. The model call that answered says which attempt it was, why the request moved on and, when a fallback answered, Served by fallback 1, 2 or 3. Skipped fallbacks are listed with their reason. A failed attempt the runtime recovered from doesn’t count as a failed request in overviews. The traces API returns provider_id, route_attempt, route_is_fallback, fallback_reason and primary_provider_id on model runs, and failed attempts have run_type llm_attempt.

Routing panel

The agent’s Overview tab has a Routing panel for the last hour, 24 hours or 7 days: the Fallback rate, the requests Served by a fallback, the requests where Every provider failed, requests and failed attempts per provider, and Why requests failed over. The same numbers come from GET /api/v1/agents/{agent_id}/routing-stats?window=24h (1h, 24h or 7d).

Provider health

On Providers, each provider shows its health over the last 15 minutes: A failure is an attempt that failed with one of the six classes above, or a request that ended in fallbacks_exhausted. That includes Rejects the key or model: a rejected key or missing model counts against the provider’s health even while it doesn’t fail over. For requests that didn’t run through a fallback route, a failure is a 429, 500, 502, 503, 504 or 529 from the provider. The same data comes from GET /api/v1/providers/health?window=15m (15m, 1h or 24h).

Alerts

Create these in Settings → Alerts: Run alert rules can also filter on provider_id. Run error doesn’t fire for failed attempts the runtime moved on from; the request’s own model call reports how it ended.

Manage fallbacks with the API

PUT /api/v1/agents/{agent_id}/fallbacks replaces an agent’s whole list of fallbacks, and optionally its routing settings. It takes an organization key with write scope (see Authentication) or a signed-in session.
  • The order of fallbacks is the order they’re tried. An empty list removes every fallback.
  • model is optional; null uses the provider’s own model.
  • routing is optional. Leave it out to keep the saved settings, send some fields to change only those, or send null to reset every setting to its default. Fields are those in Retries and timeouts; only first_token_timeout_s may be null (off).
The response is the agent, with fallbacks (each with its position, provider_id, effective model, model_override, residency and the model’s capabilities) and the effective routing. A refused save answers with a detail object holding a code from Save checks and a message, and changes nothing. To read the route, GET /api/v1/agents/{agent_id} returns fallbacks and routing, and GET /api/v1/agents/{agent_id}/providers lists the primary (role: primary, position: 0) followed by the fallbacks. Providers accept a residency field (us, eu, apac or global; null to detect it).

Costs

  • Xenovia counts a request once against your organization’s request allowance, however many attempts it took. A request that ends in fallbacks_exhausted isn’t counted.
  • Xenovia’s LLM cost figures price the call that answered. Failed attempts carry no tokens.
  • Providers may still bill for failed attempts. A provider can finish, and charge for, a generation after Xenovia gave up on it at a timeout. With a first-token timeout, output streamed before a provider error may be billed, twice if the attempt was retried on the same provider.
  • Prompt caches don’t carry over. A fallback starts with its own cache, so long cached prompts cost more while the agent is failing over, and a fallback’s model has its own price.

Limitations

  • Hybrid deployments. A customer-hosted runtime can’t load fallback credentials yet, so it skips every fallback (couldn’t be set up) and serves from the primary.
  • Requests can’t set fallbacks. A fallbacks field in a request body is removed; only the agent’s route is used.
  • Provider-side state. A request that refers to state held by the primary’s provider, such as a Responses previous_response_id or conversation, or a provider file ID, can only fail over to a fallback on that same provider with another model. Other fallbacks are skipped.
  • Up to three fallbacks per agent.