A production failover chain needs an error policy
Retrying a bad API key five times only makes an outage slower. A routing layer should identify which errors are transient enough to retry and which should move immediately to the next candidate or fail fast.
Timeout budgets matter too. If every provider receives a long timeout, a three-model chain can create a very slow user-visible failure. Bound the total strategy, not only each individual HTTP request.
Fallback models are not always equivalent
Switching between providers that serve the same underlying model can preserve behavior more closely than switching model families. Cross-model failover may change context limits, tool schemas, reasoning behavior, image support or output style.
For critical paths, define fallbacks by capability and test representative prompts. Reliability is not improved if the backup responds but silently breaks the application's contract.
What your telemetry should record
At minimum, store the requested model, served model, provider, response status, latency, retry count, fallback reason and cost. That turns failover from invisible middleware into something an operator can debug.
A useful control plane also separates the initial request from each upstream attempt so you can answer whether latency came from the final provider or from retries that happened first.