Gateway Routing And Reliability
Provon separates model intent from upstream execution. A request first resolves its model value into eligible Provider Key targets, then orders those targets using project policy a
On this page Browse sections
Capability filtering happens before scoring or recovery. Provon never routes a request to a target that cannot satisfy its endpoint and requested features.
Resolution Precedence#
| Request model | Resolver | Typical use |
|---|---|---|
<provider>/<model> |
Direct Provider | Explicit, portable application integration |
<model> |
Model Policies, then registry inference | Stable logical model names owned by a project |
provon/auto |
Auto Router | Dynamic selection across mapped BYOK targets |
Direct Provider#
openai/gpt-5-mini pins provider resolution to openai. Provon still selects among eligible
Provider Keys and credentials for that provider, but it does not substitute a different provider
through project Model Policies.
Use this form as the default when the application should make provider intent explicit.
Model Policies#
A plain model such as support-model is matched against project Model Policies. A policy can define:
- one primary Provider Key or provider;
- ordered same-model recovery targets;
- ordered cross-model substitution targets.
Model patterns support *. Rules are evaluated in configured order, and the first matching rule
wins.
If no policy matches, provider registry inference uses known model namespaces and patterns. Treat registry inference as a convenience, not a substitute for an explicit production policy.
Auto Router#
provon/auto asks Provon to select from enabled model mappings on BYOK Provider Keys. Auto Router:
- filters mappings by requested endpoint and features;
- applies the optional allowed-model patterns;
- excludes unavailable targets;
- scores eligible targets using priority, latency, health, and estimated cost;
- uses round-robin ordering to break ties.
The Selection bias control moves the scoring balance between quality-oriented runtime signals and estimated cost. It does not override capability requirements or hard usage policies.
Provider Keys#
A Provider Key is the routable upstream target. It contains:
- provider and display name;
- enabled state and priority;
- upstream credential;
- optional endpoint override;
- optional request timeout;
- optional header overrides;
- zero or more model mappings.
Within a provider, drag Provider Keys in Providers to set their priority. Matching requests use the first eligible target after policy, health, and load-balancing decisions.
Model Mapping Modes#
| Mode | Behavior |
|---|---|
| All | Route models through the key without explicit mappings |
| Specific | Route only model IDs declared by enabled mappings |
Each mapping can declare:
- public AI Gateway model ID;
- upstream model ID;
- eligible endpoint families;
- required features;
- context window and maximum output metadata;
- pricing profile metadata.
Model mappings are routing truth. Pricing definitions only enrich estimates and traces; they do not make a model routable.
Configure Model Policies#
In the Workbench:
- Connect the required Provider Keys under Providers.
- Open Routing.
- Under Model Policies, add a rule for the plain model pattern.
- Choose the Primary Provider key.
- Optionally enable Same model recovery and order alternate targets.
- Optionally enable Model substitution and declare each alternate model explicitly.
- Save the policy.
Example intent:
support-model
-> OpenAI primary / gpt-5-mini
-> Azure recovery / gpt-5-mini
-> Anthropic substitution / claude-sonnetSame-model recovery preserves the requested model identity across targets. Model substitution changes model behavior and must be enabled explicitly.
Configure Auto Router#
In Routing:
- Enable Auto Router.
- Leave Allowed models empty to consider all enabled mappings, or add wildcard patterns such as
openai/*andanthropic/claude-*. - Set Selection bias between Quality and Cost.
- Save the configuration.
- Send requests with
"model": "provon/auto".
Before relying on Auto Router, verify that each candidate mapping declares accurate endpoints, features, and upstream model IDs.
Reliability Layers#
Provon distinguishes several failure-handling mechanisms:
| Mechanism | Scope | Behavior |
|---|---|---|
| Request timeout | One Provider Key attempt | Abandons an attempt after its configured duration |
| Retry | Same upstream target | Repeats configured status or network failures with bounded delay |
| Fallback | Ordered target chain | Moves to the next eligible target after configured failures |
| Circuit breaker | Across requests per target | Temporarily skips a repeatedly failing target |
| Affinity | Across related requests | Reuses a prior healthy target for stable session behavior |
| Target health | Across requests per target | Removes unavailable targets from health-aware routing |
Retries and fallback are disabled unless configured for the target. This avoids silently repeating non-idempotent work or changing providers after application errors.
Retry And Fallback Guidance#
- Retry only transient network failures,
429, and selected5xxresponses. - Keep retry counts and total request deadlines bounded.
- Honor
Retry-Afteronly within the configured maximum delay. - Use same-model fallback before cross-model substitution.
- Do not fallback on authentication errors, invalid requests, or model permission failures.
- Verify that every fallback target supports the endpoint and requested features.
- Inspect attempt spans after changing recovery policy.
For streaming calls, recovery is only safe before a response has been committed to the client.
Capability-Aware Selection#
Provon derives request requirements from the endpoint and body, including:
- streaming;
- tool calling;
- JSON mode and structured outputs;
- vision or audio input/output;
- reasoning.
Query eligible models before deploying a new request shape:
curl \
"$PROVON_BASE_URL/v1/gateway/model-selection?endpoint=chat-completions&features=streaming,tool-calling" \
-H "Authorization: Bearer $PROVON_API_KEY"See Gateway API for the full discovery contract.
Verify Routing#
After a test request, open its trace and verify:
- requested, Gateway, and upstream model IDs;
- selected Provider Key and credential fingerprint;
- target priority and routing source;
- fallback index and retry count;
- upstream status, duration, and error classification;
- circuit, health, or guardrail events that changed execution.
The final HTTP response shows the winning attempt. The trace explains why that attempt ran.