Skip to content

DLRMv4 RCPs (GBS 8192) appear to be an unrepresentative seed sample: mean ~4 % low, std ~half, no slow-tail runs #908

Description

@mmarcinkiewicz

Samples-to-converge for dlrmv4 has a heavy slow tail: about 12 % of random seeds need 80 to 106 M samples at GBS 8192 instead of the usual 66 to 76 M. The seed alone decides this; re-running a seed reproduces its crossing exactly, and an independent implementation reproduces the reference's per-seed crossings (13/20 of the RCP seeds on the same eval-grid step, 18/20 within one step, Spearman 0.89). The 20-run RCP set contains no tail run and sits well below what a representative seed set gives:

runs mean olympic mean std runs ≥ 80 M
current RCP set, 20 reference runs (AMD) 69.27 M 69.07 M 3.81 M 1 / 20
the same 20 RCP seeds, our implementation 68.80 M 68.80 M 3.17 M 0 / 20
179 random seeds, our implementation 72.31 M 72.18 M 6.73 M 22 / 179
reference code, 20 seeds stratified over that distribution 73.29 M 73.02 M (est.) 6.8 M 2 / 20

Under the 179-seed distribution, a 20-run olympic mean at or below 69.07 M occurs with probability 1.6 %. The 20 RCP seeds themselves are fast on the independent implementation too (olympic 68.80 M), so this is the seeds, not the platform.

Effects: a correct submission reads "slower than the reference" 93 % of the time (median +3.4 %), and the checker's own t-test would call 40 % of correct submissions significantly slower. The fast-side compliance floor is barely affected (66.3 → about 67 M).

Here's a histogram of 179 runs vs 20 RCP runs showing a long tail that is not captured by the RCPs:

Image

Activity

  1. mmarcinkiewicz commented on Sep 24, 2026

    @mmarcinkiewicz
    ContributorAuthor

    Read from the 20 GBS-8192 RCP logs (mllog timestamps, eval points, sample counts). Every log runs to a fixed wall time and keeps training and evaluating well past its 0.75 crossing, so the stop criterion was a job time limit, not convergence. Three campaigns are visible:

    campaign date (UTC) logs wall time crossing: mean / min / max log stops at (samples) headroom past crossing throughput
    A 07-21 10 8.0 h 69.7 M / 61.9 / 80.3 216 to 227 M ≥ 140 M 7.6k samples/s
    B 07-24, 07-25 7 3.0 h 68.5 M / 66.5 / 73.4 73.4 to 80.3 M 4.6 M (two eval steps) 7.0k samples/s
    C 07-26 3 8.0 h 68.0 M / 66.5 / 71.1 211 to 220 M ≥ 140 M 7.6k samples/s

    Facts: in campaign B the 3 h limit ends each log at 73 to 80 M samples, so a seed in that campaign that needed more than roughly 75 to 80 M could not have produced a crossing, and a run without a crossing cannot become an RCP entry. Campaign B's seven crossings all lie at or below 73.4 M; the two 8 h campaigns include an 80.3 M run. Campaigns A and C had 140 M of headroom and could have recorded any seed up to about 210 M.

    It kind of seems like runs B were 10 originally, but due to small walltime three of them didn't converge in time and got cancelled (the three stragglers in line with expectations - that would skew the distribution towards more "real" one). Then, runs C added the missing three runs, with no stragglers. The walltime got also increased - I guess you figured out 80M samples might have been not enough? Can't say, please investigate - find the timeouted logs and rerun the same seeds with a longer walltime.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions