Repository navigation
fix: Gateway: wizard page with workspaces has flickering workspace status icons and buttons (CRW-13522) - #382
Conversation
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configuration
Comment |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #382 +/- ##
==========================================
+ Coverage 0.00% 41.55% +41.55%
==========================================
Files 4 124 +120
Lines 26 5432 +5406
Branches 0 1036 +1036
==========================================
+ Hits 0 2257 +2257
- Misses 26 2878 +2852
- Partials 0 297 +297 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
1d9b67d to
5475ca0
Compare
5475ca0 to
2b3b54e
Compare
adietish
left a comment
There was a problem hiding this comment.
Great fix, thanks!
Found another issue that I committed on top and a minor refactoring where I split DevWorkspaceWatch.watchLoop into separate method to get better readable.
…atus icons and buttons (CRW-13522) The DevWorkspace watch resumed from the resourceVersion it was originally started with instead of the latest one it had consumed, and silently retried forever on a 410/Expired response with zero logging anywhere in the failure path. This left the wizard's workspace list frozen (e.g. stuck on "Starting") until the user noticed and clicked Refresh. * Track the resourceVersion of the last consumed watch event and resume from it on reconnect, instead of the stale value start() was called with. * Detect resourceVersion expiry in both forms the k8s client surfaces it (a thrown ApiException(410) and an in-stream "ERROR" event carrying a V1Status) and recover via a relist-and-resume path (new DevWorkspaceListener.onReset callback), mirroring what the manual Refresh button already did, instead of looping on a dead resourceVersion. * Add a resourceVersion-ordering guard when applying updates to the table model, so an out-of-order/stale redelivery (e.g. a relist served from a lagging apiserver replica) can never regress an already-displayed state; a genuinely newer resourceVersion is always applied, even if its phase looks like a "regression" - only ordering is guarded, never phase semantics, so real server-side changes are never hidden. * Log every branch of the watch's error/retry/relist handling, and every consumed ADDED/MODIFIED/DELETED event, so this class of issue is diagnosable from idea.log without a code-reading investigation. Separately investigated and ruled out as a client-side issue: a Running -> Starting -> Running status flicker observed on some workspaces right after they finish starting. Debug-log capture confirmed this is the devworkspace-operator itself writing a genuinely newer, correctly ordered phase value (strictly increasing resourceVersion, no reconnect/relist involved) - the watch is correctly relaying real server state, so there is nothing to fix on the plugin side for that symptom. Fixes: https://redhat.atlassian.net/browse/CRW-13522 Signed-off-by: Victor Rubezhny <vrubezhny@redhat.com> Assisted-By: Claude: Sonnet 5 <noreply@anthropic.com>
Watch recovery must not reuse listWithResult's swallow-as-empty path. A dedicated throwing LIST lets relistAndReconcile's catch skip onReset, so a permission flap cannot clear the workspace rows. Signed-off-by: Andre Dietisheim <adietish@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com>
…W-13522) watchLoop() had grown to ~90 lines with the event loop body and the ApiException handling inlined. Extract the per-event handling into handleEvent() (returning a small EventOutcome so the loop's break-on-ERROR stays explicit) and the catch body into handleApiException(). Pure readability change: same log messages, same relistAndReconcile()/ dispatchToListener() calls, same resourceVersion tracking, no behavior change. Signed-off-by: Andre Dietisheim <adietish@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com>
c593d33 to
770bdc8
Compare
|
here's an explanatory summary using the CRW-13522 — DevWorkspace watch: 410 recovery, ordering guard, loggingThree commits, one class of bug. The wizard's workspace list froze (stuck on
The bug: a watch that resumed from a dead addressThe loop always reconnected with the RV it was started with, and every start(rv)
loop:
- watch from rv ← same stale RV, every reconnect
+ watch from currentRv ← last RV actually consumed
per event:
- dispatch to listener
+ dispatch to listener
+ currentRv = event.rv ← track stream progress
on ApiException(403/404): stop
- on ApiException(other): retry ← silent, forever
+ on ApiException(410): relistAndReconcile() ← log + recover
+ on ApiException(other): retry ← warn
+ on stream end: reconnect ← debug-log
|
fixes https://redhat.atlassian.net/browse/CRW-13522
Summary
The DevWorkspace watch in the "Connect to Dev Spaces" wizard resumed from
the
resourceVersionit was originally started with instead of the latestone it had actually consumed, and silently retried forever on a 410/Expired
response with zero logging anywhere in the failure path. In practice this
left the wizard's workspace list frozen (e.g. stuck on "Starting") until
the user noticed and clicked Refresh.
from it on reconnect, instead of the stale value
start()was called with.(a thrown
ApiException(410)and an in-stream"ERROR"event carrying aV1Status) and recover via a relist-and-resume path (newDevWorkspaceListener.onResetcallback), mirroring what the manualRefresh button already did, instead of looping on a dead resourceVersion.
model, so an out-of-order/stale redelivery (e.g. a relist served from a
lagging apiserver replica) can never regress an already-displayed state;
a genuinely newer resourceVersion is always applied, even if its phase
looks like a "regression" — only ordering is guarded, never phase
semantics, so real server-side changes are never hidden.
consumed ADDED/MODIFIED/DELETED event, so this class of issue is
diagnosable from
idea.logwithout a code-reading investigation.Separately investigated and ruled out as a client-side issue: a
Running→Starting→Runningstatus flicker observed on someworkspaces right after they finish starting. A debug-log capture confirmed
this is the devworkspace-operator itself writing a genuinely newer,
correctly ordered phase value (strictly increasing resourceVersion, no
reconnect/relist involved) — the watch is correctly relaying real server
state, so there's nothing to fix on the plugin side for that symptom.
Test plan
./gradlew test— 567 passed, 0 failed, 4 pre-existing skipsevents, 410-via-
ApiExceptionand 410-via-in-stream-ERROR-eventboth trigger relist-and-resume, 403/404 still stop the watch
permanently,
onResetreconciliation (add/update/remove, namespaceisolation), resourceVersion-ordering guard (stale update dropped,
newer update applied even with a "regressed" phase, missing/
unparseable resourceVersion fails open)
cluster and confirmed the watch now recovers automatically
Fixes: https://redhat.atlassian.net/browse/CRW-13522