Skip to content

huntsman-e2e-tests: test_nn flakily fails with "resource group already exists" from a reused Compose DB volume #502

Description

@20001020ycx

Bug

huntsman-e2e-tests intermittently fails during test_nn at the very first job, with:

Error: neural-network job 0 in batch 0 with seed 0 failed

Caused by:
    resource group already exists
test test_nn ... FAILED

Observed on both linux/amd64 and linux/arm64 self-hosted runners in the same CI run (e.g. PR #498: amd64 job, arm64 job). The change under test in that PR was unrelated to resource-group creation, so the failure is spurious.

Why this is a real (latent) bug, not a one-off

The error is not a concurrency race inside the test:

  • SpiderTestDriver is a process-wide singleton (OnceCell, tests/huntsman/e2e/src/test_driver.rs).
  • SpiderTestDriver::resolve_resource_group holds a Mutex over its resource_groups cache for the entire create-and-insert, so SpiderClient::add_resource_group("e2e-nn") is invoked exactly once per process, even though test_nn submits 24 jobs (NUM_BATCHES = 3 × NUM_JOBS_PER_BATCH = 8) all under the same external id e2e-nn.

So add_resource_group receiving ResourceGroupAlreadyExists on its single call means e2e-nn was already present in the database before the test started — i.e. leaked state from a previous run.

Root cause

The test:huntsman-e2e-tests task (taskfiles/test.yaml) tears down the Compose stack and its volumes only via a deferred step:

  • defer: :docker:compose:local:clean → docker compose --project-name spider-local ... down --volumes --remove-orphans (taskfiles/docker.yaml).

This deferred cleanup runs after the test. There is no volume wipe before :docker:compose:local:up. Consequently:

  1. A prior run on the same self-hosted runner is hard-cancelled/killed (or the runner dies) before its deferred compose:local:clean executes.
  2. The named volume spider-local_spider-database-data — containing the e2e-nn resource group row from that run — survives.
  3. The next run's docker:compose:local:up reuses the existing named volume (Compose does not recreate named volumes that already exist).
  4. test_nn's first add_resource_group("e2e-nn") hits the pre-existing row and fails, because registration is not idempotent: storage returns ResourceGroupAlreadyExists (components/spider-storage/src/grpc.rs, components/spider-storage/src/db/error.rs), surfaced to the client as resource group already exists (components/spider-client/src/error.rs).

This is environmental / infrastructure flakiness in the e2e harness, independent of the code under test. The e2e driver, nn.rs, and the e2e/compose taskfiles are identical across recent history, so any PR's e2e run can hit this whenever a preceding run leaves the volume behind.

Impact

  • Spurious huntsman-e2e-tests failures block PR merges and require manual re-runs.
  • Erodes trust in the e2e signal (a red e2e no longer reliably means "the PR broke something").

Spider version

main @ 737e85a423e552591fcf0acdf91c2b9400a1aa90 (the affected harness code predates this and is unchanged across recent history).

Environment

  • GitHub Actions self-hosted runners: linux/amd64 and linux/arm64, ubuntu-noble, Docker.
  • Docker Compose project spider-local; database image mariadb:10.11.16; named volume spider-local_spider-database-data.

Reproduction steps

  1. Run task test:huntsman-e2e-tests locally (or trigger the CI job).
  2. Hard-cancel/kill it after :docker:compose:local:up has started the stack but before the deferred :docker:compose:local:clean runs (this simulates a cancelled CI job). Confirm the volume survives: docker volume ls | grep spider-local_spider-database-data.
  3. Run task test:huntsman-e2e-tests again. test_nn fails at job 0 with resource group already exists.

(The DB volume persisting between runs is the deterministic trigger; on shared self-hosted runners this happens naturally when a run is cancelled.)

Proposed fixes

Any one of these resolves the flake; combining a harness-hygiene fix with an idempotency fix gives defense-in-depth:

  1. Wipe the Compose volume before up. Add an eager :docker:compose:local:clean (down --volumes --remove-orphans) step before :docker:compose:local:up in test:huntsman-e2e-tests, in addition to the existing deferred cleanup. Cheap and targeted.
  2. Make the e2e driver tolerant of pre-existing groups. In resolve_resource_group, treat ResourceGroupAlreadyExists as success by resolving the existing id (get-or-create), so any leaked state is harmless. (Needs a client API to look up an existing external resource-group id.)
  3. Use a unique external resource-group id per run (e.g. suffix e2e-nn with a run id/UUID) so a stale volume can never collide.
  4. Runner hygiene: prune Docker volumes between jobs on the self-hosted runners.

Recommend (1) as the immediate fix, and (2) or (3) for robustness against any residual leaked state.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions