Bug
huntsman-e2e-tests intermittently fails during test_nn at the very first job, with:
Error: neural-network job 0 in batch 0 with seed 0 failed
Caused by:
resource group already exists
test test_nn ... FAILED
Observed on both linux/amd64 and linux/arm64 self-hosted runners in the same CI run (e.g. PR #498: amd64 job, arm64 job). The change under test in that PR was unrelated to resource-group creation, so the failure is spurious.
Why this is a real (latent) bug, not a one-off
The error is not a concurrency race inside the test:
SpiderTestDriver is a process-wide singleton (OnceCell, tests/huntsman/e2e/src/test_driver.rs).
SpiderTestDriver::resolve_resource_group holds a Mutex over its resource_groups cache for the entire create-and-insert, so SpiderClient::add_resource_group("e2e-nn") is invoked exactly once per process, even though test_nn submits 24 jobs (NUM_BATCHES = 3 × NUM_JOBS_PER_BATCH = 8) all under the same external id e2e-nn.
So add_resource_group receiving ResourceGroupAlreadyExists on its single call means e2e-nn was already present in the database before the test started — i.e. leaked state from a previous run.
Root cause
The test:huntsman-e2e-tests task (taskfiles/test.yaml) tears down the Compose stack and its volumes only via a deferred step:
defer: :docker:compose:local:clean → docker compose --project-name spider-local ... down --volumes --remove-orphans (taskfiles/docker.yaml).
This deferred cleanup runs after the test. There is no volume wipe before :docker:compose:local:up. Consequently:
- A prior run on the same self-hosted runner is hard-cancelled/killed (or the runner dies) before its deferred
compose:local:clean executes.
- The named volume
spider-local_spider-database-data — containing the e2e-nn resource group row from that run — survives.
- The next run's
docker:compose:local:up reuses the existing named volume (Compose does not recreate named volumes that already exist).
test_nn's first add_resource_group("e2e-nn") hits the pre-existing row and fails, because registration is not idempotent: storage returns ResourceGroupAlreadyExists (components/spider-storage/src/grpc.rs, components/spider-storage/src/db/error.rs), surfaced to the client as resource group already exists (components/spider-client/src/error.rs).
This is environmental / infrastructure flakiness in the e2e harness, independent of the code under test. The e2e driver, nn.rs, and the e2e/compose taskfiles are identical across recent history, so any PR's e2e run can hit this whenever a preceding run leaves the volume behind.
Impact
- Spurious
huntsman-e2e-tests failures block PR merges and require manual re-runs.
- Erodes trust in the e2e signal (a red e2e no longer reliably means "the PR broke something").
Spider version
main @ 737e85a423e552591fcf0acdf91c2b9400a1aa90 (the affected harness code predates this and is unchanged across recent history).
Environment
- GitHub Actions self-hosted runners:
linux/amd64 and linux/arm64, ubuntu-noble, Docker.
- Docker Compose project
spider-local; database image mariadb:10.11.16; named volume spider-local_spider-database-data.
Reproduction steps
- Run
task test:huntsman-e2e-tests locally (or trigger the CI job).
- Hard-cancel/kill it after
:docker:compose:local:up has started the stack but before the deferred :docker:compose:local:clean runs (this simulates a cancelled CI job). Confirm the volume survives: docker volume ls | grep spider-local_spider-database-data.
- Run
task test:huntsman-e2e-tests again. test_nn fails at job 0 with resource group already exists.
(The DB volume persisting between runs is the deterministic trigger; on shared self-hosted runners this happens naturally when a run is cancelled.)
Proposed fixes
Any one of these resolves the flake; combining a harness-hygiene fix with an idempotency fix gives defense-in-depth:
- Wipe the Compose volume before
up. Add an eager :docker:compose:local:clean (down --volumes --remove-orphans) step before :docker:compose:local:up in test:huntsman-e2e-tests, in addition to the existing deferred cleanup. Cheap and targeted.
- Make the e2e driver tolerant of pre-existing groups. In
resolve_resource_group, treat ResourceGroupAlreadyExists as success by resolving the existing id (get-or-create), so any leaked state is harmless. (Needs a client API to look up an existing external resource-group id.)
- Use a unique external resource-group id per run (e.g. suffix
e2e-nn with a run id/UUID) so a stale volume can never collide.
- Runner hygiene: prune Docker volumes between jobs on the self-hosted runners.
Recommend (1) as the immediate fix, and (2) or (3) for robustness against any residual leaked state.
Bug
huntsman-e2e-testsintermittently fails duringtest_nnat the very first job, with:Observed on both
linux/amd64andlinux/arm64self-hosted runners in the same CI run (e.g. PR #498: amd64 job, arm64 job). The change under test in that PR was unrelated to resource-group creation, so the failure is spurious.Why this is a real (latent) bug, not a one-off
The error is not a concurrency race inside the test:
SpiderTestDriveris a process-wide singleton (OnceCell,tests/huntsman/e2e/src/test_driver.rs).SpiderTestDriver::resolve_resource_groupholds aMutexover itsresource_groupscache for the entire create-and-insert, soSpiderClient::add_resource_group("e2e-nn")is invoked exactly once per process, even thoughtest_nnsubmits 24 jobs (NUM_BATCHES = 3×NUM_JOBS_PER_BATCH = 8) all under the same external ide2e-nn.So
add_resource_groupreceivingResourceGroupAlreadyExistson its single call meanse2e-nnwas already present in the database before the test started — i.e. leaked state from a previous run.Root cause
The
test:huntsman-e2e-teststask (taskfiles/test.yaml) tears down the Compose stack and its volumes only via a deferred step:defer: :docker:compose:local:clean→docker compose --project-name spider-local ... down --volumes --remove-orphans(taskfiles/docker.yaml).This deferred cleanup runs after the test. There is no volume wipe before
:docker:compose:local:up. Consequently:compose:local:cleanexecutes.spider-local_spider-database-data— containing thee2e-nnresource group row from that run — survives.docker:compose:local:upreuses the existing named volume (Compose does not recreate named volumes that already exist).test_nn's firstadd_resource_group("e2e-nn")hits the pre-existing row and fails, because registration is not idempotent: storage returnsResourceGroupAlreadyExists(components/spider-storage/src/grpc.rs,components/spider-storage/src/db/error.rs), surfaced to the client asresource group already exists(components/spider-client/src/error.rs).This is environmental / infrastructure flakiness in the e2e harness, independent of the code under test. The e2e driver,
nn.rs, and the e2e/compose taskfiles are identical across recent history, so any PR's e2e run can hit this whenever a preceding run leaves the volume behind.Impact
huntsman-e2e-testsfailures block PR merges and require manual re-runs.Spider version
main@737e85a423e552591fcf0acdf91c2b9400a1aa90(the affected harness code predates this and is unchanged across recent history).Environment
linux/amd64andlinux/arm64,ubuntu-noble, Docker.spider-local; database imagemariadb:10.11.16; named volumespider-local_spider-database-data.Reproduction steps
task test:huntsman-e2e-testslocally (or trigger the CI job).:docker:compose:local:uphas started the stack but before the deferred:docker:compose:local:cleanruns (this simulates a cancelled CI job). Confirm the volume survives:docker volume ls | grep spider-local_spider-database-data.task test:huntsman-e2e-testsagain.test_nnfails at job 0 withresource group already exists.(The DB volume persisting between runs is the deterministic trigger; on shared self-hosted runners this happens naturally when a run is cancelled.)
Proposed fixes
Any one of these resolves the flake; combining a harness-hygiene fix with an idempotency fix gives defense-in-depth:
up. Add an eager:docker:compose:local:clean(down --volumes --remove-orphans) step before:docker:compose:local:upintest:huntsman-e2e-tests, in addition to the existing deferred cleanup. Cheap and targeted.resolve_resource_group, treatResourceGroupAlreadyExistsas success by resolving the existing id (get-or-create), so any leaked state is harmless. (Needs a client API to look up an existing external resource-group id.)e2e-nnwith a run id/UUID) so a stale volume can never collide.Recommend (1) as the immediate fix, and (2) or (3) for robustness against any residual leaked state.