The reductions of host-accessible arrays testset added in #649 fails intermittently on the ALCF (Aurora, LTS stack) CI with Julia 1.13:
Error in testset reductions of host-accessible arrays:
Test Failed at test/array.jl:116
Expression: sum(a) == 1024
Evaluated: 0.0f0 == 1024
Error in testset reductions of host-accessible arrays:
Test Failed at test/array.jl:117
Expression: maximum(a) == 1
Evaluated: 0.0f0 == 1
So sum/maximum sometimes still return 0 on a SharedBuffer/HostBuffer-backed array: the symptom #649 was meant to fix. Two of the four tests in the loop fail, i.e. presumably both tests for one of the two buffer types (the output doesn't say which).
Occurrences
All on test:lts:1.13; the corresponding test:lts:1.10 jobs passed every time:
| Branch |
Commit |
ALCF job |
Result |
tb/cooperative-wait (→ #654) |
72a9354 |
277735 |
failed |
tb/cooperative-wait (→ #654) |
72a9354 |
277924 |
passed (re-run, same commit) |
tb/cooperative-wait (→ #654) |
de2d8c1 |
278057 |
failed |
main |
5a3557a |
279092 |
passed |
| #655 (CompatHelper AcceleratedKernels 0.5) |
782a97d |
279579 |
failed |
#655 only touches accumulate/cumsum/cumprod, which sum/maximum don't use, and it resolves the same GPUArrays (11.5.15) and KernelAbstractions (0.9.43) as the passing main run. So this looks like a race, not a regression from any one change. It first appears on the branch that introduced cooperative synchronization (#654).
Environment: Julia 1.13.1, NEO 26.18.38308, IGC 2.34.4, Intel Data Center GPU Max 1550, test workers spread over 12 GPUs (ONEAPI_TEST_SPREAD_GPUS=1), run with --quickfail.
Notes
The
reductions of host-accessible arraystestset added in #649 fails intermittently on the ALCF (Aurora, LTS stack) CI with Julia 1.13:So
sum/maximumsometimes still return 0 on aSharedBuffer/HostBuffer-backed array: the symptom #649 was meant to fix. Two of the four tests in the loop fail, i.e. presumably both tests for one of the two buffer types (the output doesn't say which).Occurrences
All on
test:lts:1.13; the correspondingtest:lts:1.10jobs passed every time:tb/cooperative-wait(→ #654)tb/cooperative-wait(→ #654)tb/cooperative-wait(→ #654)main#655 only touches
accumulate/cumsum/cumprod, whichsum/maximumdon't use, and it resolves the same GPUArrays (11.5.15) and KernelAbstractions (0.9.43) as the passingmainrun. So this looks like a race, not a regression from any one change. It first appears on the branch that introduced cooperative synchronization (#654).Environment: Julia 1.13.1, NEO 26.18.38308, IGC 2.34.4, Intel Data Center GPU Max 1550, test workers spread over 12 GPUs (
ONEAPI_TEST_SPREAD_GPUS=1), run with--quickfail.Notes