Repository navigation
fix(libsy): cancel abandoned offloaded calls - #948
nachiketb-nvidia wants to merge 1 commit into
Conversation
Signed-off-by: nachiketb <nachiketb@nvidia.com>
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (20)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. WalkthroughRust and Python model and decision response methods now accept asynchronous work. The implementation cancels pending work when the algorithm stops waiting. Client integrations, tests, examples, and documentation now use or describe the asynchronous response contract. ChangesAsynchronous Responses
Priority: ➖ Normal Estimated code review effort: 4 (Complex) | ~45 minutes Merge Risk: ⚪ Minimal · up to The change makes abandoned model and decision calls cancel their host-side work, and the callers, tests, and docs are updated to match. No concrete merge-blocking risk was identified. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 42.86% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 56 functions across 16 files. (4 skipped: 4 unsupported.)
A rabbit taps a future into flight. Comment |
What
Let an algorithm cancel one model or decision call without cancelling its other calls or ending the run.
CallModel::respondandCallDecision::respondnow accept the work future and are async. They stop that work when the algorithm drops its waiting call future.respond, rather than finishing the request first. Cancelled calls return normally instead of failing the run withResponseDropped.ModelCall.respondandDecisionCall.respondaccept an awaitable. The binding cancels the Python task and awaits its cleanup when the algorithm stops waiting. Pythonfailis now awaitable too.Algorithm,Driver,run_stream, anddrivekeep their existing signatures. No new dependency or cancellation-token API is needed.Why
Parallel calls already work, but abandoning one call did not stop its HTTP work while the algorithm continued. The host awaited the request before replying, so the reply channel could only detect cancellation after the work finished. A late reply could then fail the entire run.
Passing the future to
respondlets the existing reply channel control that call's lifetime. This supports racing models, abandoning a slow call, and continuing with another call.API Migration
This changes the low-level host API. Pass the work itself, not its awaited result.
Rust:
Python:
Custom stream consumers must keep polling handlers concurrently, then cancel and await outstanding tasks when the run ends. The Python example shows this cleanup. Dropping a spawned task handle alone does not cancel it.
Model errors follow
recover_errorsinsiderespond: recoverable errors go back to the algorithm; other errors return to the host. The standard HTTP client's existing policy stays the same. Decision errors still go back to the algorithm for fallback.Cancellation from the Algorithm
The algorithm cancels a call by dropping its waiting future. There is no separate
cancel()method. This works for bothDriver::call_modelandDriver::call_decision.Race two calls
tokio::select!polls both calls concurrently. When one finishes, it drops the other call future and signals the host to stop that work:This races the first completion, not the first successful answer. These examples use
?to propagate errors.Cancel a decision after a timeout
A timeout drops the decision-call future if its deadline expires. The algorithm can then apply its own fallback:
Cancel based on another model's output
The algorithm can inspect the fast response before deciding whether to keep the strong call. Here,
is_acceptableis the algorithm's own check, not a libsy API:Both calls run concurrently. If the fast answer passes the check, the strong call is cancelled. If it does not, the strong call keeps running. If strong finishes first, fast is cancelled. The selected model and response can then form the normal
RoutingOutcome.Use a decision to cancel a model call
A decision call can also run alongside a speculative model call. In this example, the algorithm starts strong while the judge decides whether fast is sufficient.
choose_fastis the algorithm's interpretation of the decision response:If the judge chooses fast first, the algorithm cancels strong and calls fast. If it chooses strong, the existing strong call continues. If strong finishes first, its answer is used and the judge is cancelled.
When selecting on
&mut future, the future remains owned by the algorithm. Explicitly drop the owned future or leave its scope to cancel it. TheBox::pinexamples above make that ownership clear; dropping only a borrowedPin<&mut _>does not cancel the underlying future.Notes for reviewers
crates/libsy/src/core/algorithm.rs: reply-channel closure drops the work future; a late reply after cancellation no longer aborts unrelated work.crates/libsy-llm-client/src/run.rsnext: the cancellable future owns the actual provider call and its completed-call observations.crates/switchyard-py/src/libsy_bindings.rsandswitchyard_rust/_libsy_async.pytogether. Dropping the Rust adapter alone does not cancel an asyncio task, so the small Python helper owns cancellation and cleanup.LoC Breakdown
Diff line counts across 20 files. These include comments and blank lines. Inline Rust test changes are counted under tests, not runtime implementation.
Runtime implementation totals 203 added / 82 removed (+121 net) across Rust and Python. The two new cancellation test files account for 250 added lines: 132 Rust and 118 Python. Existing test changes mainly migrate callers to the async response API and preserve the tested error behavior.
Validation
cargo test --workspace --lockedcargo clippy --workspace --all-targets --locked -- -D warningsprefill-routerenabledprefill-routerenabledcargo fmt --all --checkmkdocs build --strictLive checks against NVIDIA Inference with GPT-OSS 20B and 120B also passed: concurrent completion, first-response race, whole-run cancellation, and cancellation of one call while the algorithm continues. In the continuation check, the abandoned HTTP future dropped within about 1 ms and the next call completed successfully. These checks verify local future cleanup, not upstream compute or billing cancellation.
Summary by CodeRabbit
New Features
Documentation