AI API Fallbacks: Handle Outages Without Retry Loops
An unavailable model should not trigger an endless loop. Separate balance problems from outages, test alternatives and put firm limits on retries.
By omnirouter
AI API fallbacks should keep a small problem from becoming a large bill or a confusing user experience. The goal is not to hide every failure. It is to retry only when another attempt makes sense, switch only to a tested alternative and stop when the remaining options are no longer useful.
That matters when you buy budget access. Upstream capacity and model availability can change, especially for frontier models. Build around a tested open-weight default where it fits your workload, and make failure an explicit state in your application.
Provider failover is not your entire recovery plan
Omnirouter lists automatic failover across providers on its pricing page. That does not mean your application can assume every request will eventually succeed. The service terms explicitly allow models, providers and capacity to change and do not guarantee uninterrupted availability or a particular response time.
Gateway-level provider failover and application-level model fallback solve different problems. A gateway may change the provider serving a model. Your application may decide to use another model after a failed or unacceptable result. Do not assume that switching models preserves behavior, tool support or output formatting.
Classify the error before retrying
Begin with the errors documented in the Omnirouter API guide:
| Response | Suggested application behavior |
|---|---|
| 402 | Stop and surface a balance problem. Repeating the same request does not add credit. |
| 404 | Inspect the model ID and account availability. Do not repeatedly request an unknown or disabled model. |
| 503 | Treat it as a candidate for a bounded retry or a tested fallback, within your deadline. |
These are starting rules, not a complete description of every error a client can encounter. Read the response body and preserve the request_id. If a client reports a timeout without a response, the outcome may be unknown: a timeout is not proof that no work happened upstream.
The documentation says a gateway-unavailable 503 request is not charged. Do not generalize that statement to every transport timeout, partial stream or application cancellation. Inspect your usage records when the result is uncertain.
Set three limits, not just a retry count
Define an attempt limit, a wall-clock deadline and an application-side spend allowance. A limit on any one of them can end the operation. These are controls to implement in your application, not claims about built-in Omnirouter settings.
For a small interactive task, an example policy could allow one original attempt, one retry and at most one attempt on an alternative model. That is an illustrative policy, not a universal recommendation. A background task may need a different deadline; an expensive request may justify no automatic retry.
Use increasing delays with jitter for retryable failures. Respect any documented server backoff guidance your client receives. Check whether your SDK already retries so two layers do not multiply the number of attempts. Avoid having every worker retry at exactly the same instant.
A useful decision sequence is:
- Is the error understood and eligible for retry?
- Is there enough time and budget for another attempt?
- Has this operation already consumed its attempt allowance?
- If retry is no longer appropriate, is a tested fallback available?
- Otherwise, stop and return a clear error.
Validate the fallback before an outage
Pick an available candidate from GLM, DeepSeek, Qwen or Kimi when it meets your task requirements. Test the exact model ID and endpoint rather than relying on the family name. No family is an uptime guarantee, and two model names may still depend on shared infrastructure.
Keep a small acceptance suite covering your required output format, language, context size and any supported tool behavior. A fallback that responds quickly but returns unusable data is not a successful recovery.
Make model switching visible when it affects the user’s requested behavior or cost. Do not silently substitute a weaker workflow for a feature your interface promises. If the user explicitly selected a model, define whether switching is allowed before it becomes necessary.
Treat partial streams and tool actions carefully
If a stream breaks halfway through, do not blindly concatenate a second answer onto the first. Decide whether to discard the partial response, offer a restart or present it as incomplete. A new generation can disagree with the earlier fragment.
A repeated model request can also produce repeated tool instructions. Protect actions such as sending messages or creating records with application-level deduplication and appropriate approvals. An instruction generated twice must not automatically execute twice. Do not assume an undocumented idempotency header will solve this for you.
Log enough to diagnose, not enough to leak
Record the request ID, model ID, endpoint, timestamp, outcome and elapsed time. Keep API keys out of logs. Avoid storing full prompts and outputs by default when metadata is sufficient; review the privacy policy before designing your own capture policy.
For persistent failures, send Telegram support a request ID and a sanitized description. An actionable report is more useful than a screenshot containing credentials.
A good fallback plan is deliberately finite. Start with a tested default, keep one proven alternative and make the stop condition obvious. Check current model availability before wiring those choices into a long-running job.