Dev docs brief — Model Routing (Quota Control + High Availability)
For: developer-docs writer. Goal: publish the developer-facing docs for the model-routing release. Two pieces: Quota Control (already live — light touch) and High Availability (new — write this). Voice: match the existing Token Quotas page — one-line value prop, reference tables, a short "how it's calculated" explainer, config path, one tip. Customer/developer audience, not internal support.
Piece 1 — Quota Control (already documented)
The page exists at /user-guides/quotas/ (Token Quotas). No new article needed. Two small updates for the release:
- Add a cross-link from Token Quotas → the new High Availability page (and vice versa), framed as the two halves of model routing: spend vs. uptime.
- Optional: a one-line intro note that quotas and failover are configured in the same routing area.
Piece 2 — High Availability (new article — write this)
A ready-to-edit draft already exists at deliverables/docs/model-high-availability.md. Use it as the starting point; the required content and facts are below so it can be verified against the product.
Suggested path: /user-guides/high-availability/ (sibling to Token Quotas).
Title: Model High Availability.
Source of truth
Internal Eng wiki, "LLM Failover — How Model & Key Fallback Works" (Guy Losica, Eng Support). This brief translates that internal triage doc into a customer-facing page — drop the escalation columns, "support takeaway," and eng-triage guidance.
Required sections
Value + framing — the agent keeps answering when a provider throttles, runs out of credit, or has an outage; failover is automatic and usually invisible. Position as the uptime half of model routing; cross-link Token Quotas.
How failover works — an ordered list of options, each a model + API key (and region where relevant); send to the first healthy one; on a recoverable error, move to the next automatically. Key concept: "unavailable" is scoped to a specific model + key + region, never a whole provider forever.
Fallback order
- Key order (per model): customer BYOK → another product-level customer key → Foldspace platform key (when platform fallback is allowed).
- Model order: start on the agent's configured model; if all its keys are down, try one peer model in the same capability tier on the other provider (US / global). EU stays on Gemini — no cross-provider hop.
Model pairs table — reproduce exactly:
Tier Google Gemini OpenAI peer Lite Gemini 3.1 Flash Lite GPT-5.6 Luna Standard Gemini 3.5 Flash GPT-5.6 Terra Pro Gemini 3.1 Pro Preview GPT-5.6 Sol Typical full order — the US, BYOK + platform-fallback walk:
# What we try Meaning 1 Requested model + customer BYOK Primary path 2 Requested model + other customer key Same model, different customer key 3 Requested model + platform key Same model, Foldspace key 4 Peer model + customer BYOK Other provider, same tier 5 Peer model + platform key Other provider on Foldspace key Note the streaming rule: failover happens only before the first tokens reach the user; once streaming starts, the option is committed. Flaky errors mid-request may retry briefly on the same option before hopping.
Error types (customer view) — trim the internal 7-row table to what a customer acts on. Columns: Error type · What it means · What failover does · What you should do.
Error type Meaning Failover Customer action Capacity Provider shared capacity exhausted (429) Brief retry, then next option; parks after repeats Usually nothing; transient Unavailable Provider timeout / outage (408 / 425 / 5xx) Brief retry, then next option Nothing; transient Billing / spend cap Customer provider quota / billing exhausted Next key or model; platform fallback may keep chat working Check provider billing / quota Rate limited Key hit RPM / TPM (non-billing 429) ~2-min cooldown on the key, use another key / model Raise limits or reduce traffic Auth Invalid / unauthorized key (401 / 403) None — no silent platform fallback Fix / rotate BYOK; chat fails until corrected Client error Bad request / unsupported model (400 / 404 / 422) None — config issue, not capacity Review model config on the agent Call out separately that Auth failures are intentionally fail-loud — Foldspace will not spend platform budget in the background to cover a broken customer key.
"Already unavailable" behavior — recently-parked model+key combos are skipped: customer key down + platform OK → same model continues on platform key; all keys for a model down → try the peer model; nothing healthy left → the turn fails with a technical-difficulty message.
Confirm before publish
- Model-name exposure: the pair table names OpenAI peers (GPT-5.6 Luna / Terra / Sol). Confirm these are cleared for public docs, or publish tiers only.
- Approved providers: the internal wiki's billing row mentions "Anthropic billing" as an example; approved providers are OpenAI + Gemini only — leave Anthropic out of the customer doc.
- Cooldown value: internal doc says ~2 minutes default — confirm before stating it publicly.
Explicitly out of scope
Provisioned throughput / pay-ahead capacity is not part of this release. Do not document it.