Skip to content

docs(fallbacks): a fallback hop drops the encrypted reasoning its target cannot decrypt and keeps the summary - #2205

Merged
mateo-berri merged 1 commit into
mainfrom
litellm_docs_fallback_hop_drops_foreign_encrypted_reasoning
Oct 9, 2026
Merged

mateo-berri merged 1 commit into
mainfrom
litellm_docs_fallback_hop_drops_foreign_encrypted_reasoning

Conversation

@mateo-berri

@mateo-berri mateo-berri commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

TLDR

BerriAI/litellm#45393 made every fallback hop drop the encrypted reasoning its target cannot decrypt and keep the summary: an order-based hop to the next order, a configured fallbacks hop to another model group, and the retry of a Responses stream that broke mid-stream. The docs only described that strip inside the opt-in encrypted_content_affinity check, so a reader setting up an ordered group of two providers could not tell that a hop survives on its own. This PR adds it to the four pages that explain fallbacks and order-based routing, and says why the check still belongs on for such a group. The code change is in main and will be in v1.106.0-rc.1 (Sat, Oct 10), with stable v1.106.0 on Sat, Oct 17

User Flow

Before

A proxy admin reading the docs is left believing a hop to another provider replays the failed deployment's encrypted reasoning and fails with invalid_encrypted_content unless encrypted_content_affinity is on

  1. Open https://docs.litellm.ai/docs/proxy/reliability#explanation. The list names the three fallback types and says nothing about what a fallback does with a conversation that carries encrypted reasoning from the failed deployment
  2. Open https://docs.litellm.ai/docs/proxy/load_balancing#how-order-based-fallback-works. The section ends at "If all order levels are exhausted, the router falls through to any configured model-level fallbacks", with no word on reasoning items
  3. Open https://docs.litellm.ai/docs/routing#deployment-ordering-priority. The "When a request to an order=1 deployment fails" paragraph ends the same way
  4. Open https://docs.litellm.ai/docs/proxy/load_balancing#special-considerations-for-responses-api. The last paragraph puts the drop under the affinity check ("When that deployment is not in the healthy pool for a follow-up ..."), so the reader takes the check as the only way a follow-up survives a deployment change
  5. Open https://docs.litellm.ai/docs/response_api#when-the-originating-deployment-cannot-serve-the-turn. The section describes the degraded form (summary kept, encrypted payload dropped) as the check's own behavior. Nothing says a fallback hop does the same with the check off, that there is no per-deployment switch, which deployment a hop attributes an unmarked item to, or what a hop answered through v1.105.x

After

The same pages say every fallback hop drops the reasoning its target cannot decrypt and keeps the summary, with no switch to flip, and why encrypted_content_affinity still belongs on an ordered group of two providers

  1. Open https://docs.litellm.ai/docs/proxy/reliability#explanation. Below the three-type list, a new paragraph says a fallback to another provider or another API key drops the encrypted reasoning the target cannot decrypt and keeps each Responses reasoning item's summary (a /v1/messages thinking block is dropped whole), that this runs on all three fallback types, on an order-based fallback, and on the fallback of a stream that broke mid-stream, that there is no per-deployment switch, and when to turn the check on as well
  2. Open https://docs.litellm.ai/docs/proxy/load_balancing#how-order-based-fallback-works. After the model-level fallbacks sentence, one sentence says a hop to a deployment that cannot decrypt the failed deployment's reasoning drops those items and keeps their summaries, with a link to the fallbacks page
  3. Open https://docs.litellm.ai/docs/routing#deployment-ordering-priority. The same sentence follows the order paragraph
  4. Open https://docs.litellm.ai/docs/proxy/load_balancing#special-considerations-for-responses-api. A new last paragraph says a fallback hop does the same whether or not the check is on, names the three hop kinds, says there is no per-deployment switch, and says why to keep the check on for an ordered group of two providers
  5. Open https://docs.litellm.ai/docs/response_api#when-the-originating-deployment-cannot-serve-the-turn. Two new paragraphs after "Only a peer keeps the reasoning ..." say what a hop strips with or without the check, how a hop attributes unmarked items when the check is off (to the deployment that just failed, so a hop to the same api_base and api_key keeps them, any other hop drops them, and the next turn pays one failed call at order: 1 first), that such a hop logs only its Falling back to model_group line, why the check still belongs on, and that releases through v1.105.x answered 400 (invalid_encrypted_content on OpenAI, invalid encrypted reasoning on Bedrock) or, with the check on, 429 No deployments available

Changes

docs/proxy/reliability.md gets the paragraph right after the three-type list in "Explanation", since that is where a reader learns what a fallback is. docs/proxy/load_balancing.md gets one sentence at the end of "How order-based fallback works" and a closing paragraph in "Special Considerations for Responses API". docs/routing.md gets the same one sentence at the end of "Deployment Ordering (Priority)". docs/response_api.md gets the two paragraphs in "When the originating deployment cannot serve the turn", where the affinity check's own strip is described, so the hop strip sits next to it. Every fact is checked against the PR's diff and its QA table: the strip runs for every hop kind, there is no per-deployment switch, a hop logs no line of its own, and before the fix a hop answered 400 (or 429 with the check on)

Both docs lints pass at the tip: node scripts/check-writing-style.js docs blog release_notes and python3 scripts/check-docs.py docs ("No structural problems found")

Screenshots / Proof of Fix

Rendered with the docs dev server (npm start, hot reload) at the branch tip, viewport 1280x900. Each before shot is the same page at main, framed at the same paragraph. To reproduce, start the dev server and open the paths below on it

reliability, "Explanation"

Path /docs/proxy/reliability#explanation. Look right below the fallbacks list item. Before, the section ends at the list. After, the paragraph starting "A fallback to a deployment on another provider" follows it

Before

Before: before-reliability-explanation
After

After: after-reliability-explanation

load_balancing, "How order-based fallback works"

Path /docs/proxy/load_balancing#how-order-based-fallback-works. Look right below "If all order levels are exhausted ...". After, the sentence starting "A hop to a deployment that cannot decrypt" follows it

Before

Before: before-load-balancing-order
After

After: after-load-balancing-order

routing, "Deployment Ordering (Priority)"

Path /docs/routing#deployment-ordering-priority. Look right below the "When a request to an order=1 deployment fails" paragraph. After, the same sentence follows it

Before

Before: before-routing-ordering
After

After: after-routing-ordering

load_balancing, "Special Considerations for Responses API"

Path /docs/proxy/load_balancing#special-considerations-for-responses-api. Look at the end of the section, above "Learn more about Encrypted Content Affinity". After, the paragraph starting "A fallback hop does the same whether or not the check is on" closes the section

Before

Before: before-load-balancing-responses
After

After: after-load-balancing-responses

response_api, "When the originating deployment cannot serve the turn"

Path /docs/response_api#when-the-originating-deployment-cannot-serve-the-turn. Look right below the "Only a peer keeps the reasoning" paragraph. After, the two paragraphs starting "The same strip runs on every fallback hop" and "Turn the check on for an ordered group of two providers all the same" sit between it and "The check can be turned on and off"

Before

Before: before-response-api-cannot-serve
After

After: after-response-api-cannot-serve


Note

Low Risk
Documentation-only changes with no runtime or configuration impact.

Overview
Documents behavior that already shipped in main (BerriAI/litellm#45393): every fallback hop drops encrypted reasoning the target deployment cannot decrypt and keeps readable summaries, so turns succeed instead of failing with invalid_encrypted_content.

The updates spread that fact across four docs where admins learn about fallbacks and ordering—reliability.md (after the three fallback types), load_balancing.md and routing.md (order-based hops), and response_api.md / load-balancing Responses section (alongside encrypted_content_affinity). They clarify this applies with or without the affinity check, covers order hops, model-group fallbacks, and mid-stream Responses retries, and has no per-deployment toggle.

The response API page adds detail on hop attribution when the check is off, logging differences, and why affinity still matters for multi-provider order groups; it notes pre–v1.106 behavior (400/429 on hops).

Reviewed by Cursor Bugbot for commit 46bffe5. Bugbot is set up for automated code reviews on this repo. Configure here.

@vercel

vercel Bot commented Oct 9, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
litellm Ready Ready Preview Oct 9, 2026 5:13am UTC

Request Review

@mateo-berri

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit 46bffe5. Configure here.

@mateo-berri
mateo-berri merged commit 6a8ba43 into main Oct 9, 2026
6 checks passed
@mateo-berri
mateo-berri deleted the litellm_docs_fallback_hop_drops_foreign_encrypted_reasoning branch October 9, 2026 05:23

This branch was successfully deployed

1 active deployment
Preview — 46bffe5c Deployed Oct 9, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant