Skip to content

feat(protocol): preserve embedded host errors - #922

Draft
afourniernv wants to merge 1 commit into
NVIDIA-NeMo:mainfrom
afourniernv:codex/litellm-host-errors
Draft

afourniernv wants to merge 1 commit into
NVIDIA-NeMo:mainfrom
afourniernv:codex/litellm-host-errors

Conversation

@afourniernv

Copy link
Copy Markdown
Contributor

What

Adds LlmClientError::Host for errors owned by an embedded Rust host.

The original error stays boxed so the host can recover its native type. Switchyard records it as a host error without putting the source text into telemetry.

Why

An embedded client sometimes needs to pass a host error through Switchyard and recover it at the outer boundary. General(String) loses the type, while Ffi describes a different boundary.

Notes for reviewers

Start with crates/protocol/src/client.rs. The remaining changes keep the new category safe and consistent in libsy, client-call telemetry, and runner failure summaries.

This adds one variant to the existing #[non_exhaustive] LlmClientError enum. No existing methods, variants, or signatures change. The error is terminal unless the host maps it to one of Switchyard's existing retryable categories.

Related: #912 adds a separate runner-construction API for embedded hosts. The two PRs can merge independently.

Validated with:

  • cargo fmt --all --check
  • cargo test -p switchyard-libsy boxed_sources_are_reduced_to_their_class
  • cargo test -p switchyard-llm-client --test observability host_error_source_is_redacted_from_the_client_call_span
  • cargo test -p switchyard-runner failure::tests
  • cargo test -p switchyard-llm-client fallback_only_accepts_context_policy_and_unavailable_failures

Signed-off-by: Alex Fournier <afournier@nvidia.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant