Samples-to-converge for dlrmv4 has a heavy slow tail: about 12 % of random seeds need 80 to 106 M samples at GBS 8192 instead of the usual 66 to 76 M. The seed alone decides this; re-running a seed reproduces its crossing exactly, and an independent implementation reproduces the reference's per-seed crossings (13/20 of the RCP seeds on the same eval-grid step, 18/20 within one step, Spearman 0.89). The 20-run RCP set contains no tail run and sits well below what a representative seed set gives:
| runs |
mean |
olympic mean |
std |
runs ≥ 80 M |
| current RCP set, 20 reference runs (AMD) |
69.27 M |
69.07 M |
3.81 M |
1 / 20 |
| the same 20 RCP seeds, our implementation |
68.80 M |
68.80 M |
3.17 M |
0 / 20 |
| 179 random seeds, our implementation |
72.31 M |
72.18 M |
6.73 M |
22 / 179 |
| reference code, 20 seeds stratified over that distribution |
73.29 M |
73.02 M (est.) |
6.8 M |
2 / 20 |
Under the 179-seed distribution, a 20-run olympic mean at or below 69.07 M occurs with probability 1.6 %. The 20 RCP seeds themselves are fast on the independent implementation too (olympic 68.80 M), so this is the seeds, not the platform.
Effects: a correct submission reads "slower than the reference" 93 % of the time (median +3.4 %), and the checker's own t-test would call 40 % of correct submissions significantly slower. The fast-side compliance floor is barely affected (66.3 → about 67 M).
Here's a histogram of 179 runs vs 20 RCP runs showing a long tail that is not captured by the RCPs:

Samples-to-converge for dlrmv4 has a heavy slow tail: about 12 % of random seeds need 80 to 106 M samples at GBS 8192 instead of the usual 66 to 76 M. The seed alone decides this; re-running a seed reproduces its crossing exactly, and an independent implementation reproduces the reference's per-seed crossings (13/20 of the RCP seeds on the same eval-grid step, 18/20 within one step, Spearman 0.89). The 20-run RCP set contains no tail run and sits well below what a representative seed set gives:
Under the 179-seed distribution, a 20-run olympic mean at or below 69.07 M occurs with probability 1.6 %. The 20 RCP seeds themselves are fast on the independent implementation too (olympic 68.80 M), so this is the seeds, not the platform.
Effects: a correct submission reads "slower than the reference" 93 % of the time (median +3.4 %), and the checker's own t-test would call 40 % of correct submissions significantly slower. The fast-side compliance floor is barely affected (66.3 → about 67 M).
Here's a histogram of 179 runs vs 20 RCP runs showing a long tail that is not captured by the RCPs: