Skip to content

Fix production problems (coming from preemption at SLAC) - #2130

Merged
tvami merged 4 commits into
trunkfrom
iss2129-fix-en-production-problems
Sep 14, 2026
Merged

tvami merged 4 commits into
trunkfrom
iss2129-fix-en-production-problems

Conversation

@tvami

@tvami tvami commented Sep 11, 2026

Copy link
Copy Markdown
Member

I am updating ldmx-sw, here are the details.

What are the issues that this addresses?

Resolves #2129

Check List

  • I successfully compiled ldmx-sw with my developments.
  • I read, understood and follow the coding rules.
  • I ran my developments and the following shows that they are successful.

@github-actions

Copy link
Copy Markdown
Contributor

Validation Results

All validation samples passed! ✅

Sample Status
cascade_history ✅ PASS
deep_ecal_gun ✅ PASS
eat_signal ✅ PASS
ecal_pn ✅ PASS
hcal ✅ PASS
inclusive ✅ PASS
it_pileup ✅ PASS
kaon_enhanced ✅ PASS
reduced_ldmx ✅ PASS
signal ✅ PASS
signal_target_al ✅ PASS
target_genie ✅ PASS
target_pn_lyso ✅ PASS
target_ti_en ✅ PASS
wab_lhe ✅ PASS

cascade_history:

  • Text Differences Between Logs (4 lines differ)
  • Log character count differs by 0% (gold=4450280, new=4450298); within tolerance
  • Timing for cascade_history: gold=6031s, new=5469s
  • Timing within 9% of gold (tolerance 10%)

deep_ecal_gun:

  • Text Differences Between Logs (24 lines differ)
  • Log character count differs by 0% (gold=11769584, new=11769602); within tolerance
  • Timing anomaly for deep_ecal_gun: new=2781s vs gold=4108s (-32%, tolerance 10%)
  • Timing for deep_ecal_gun: gold=4108s, new=2781s

eat_signal:

  • Timing regression for eat_signal: new=1518s vs gold=1262s (+20%, tolerance 10%)
  • Timing for eat_signal: gold=1262s, new=1518s

ecal_pn:

  • Text Differences Between Logs (60 lines differ)
  • Log character count differs by 0% (gold=26263133, new=26263151); within tolerance
  • Timing regression for ecal_pn: new=5779s vs gold=4486s (+28%, tolerance 10%)
  • Timing for ecal_pn: gold=4486s, new=5779s

hcal:

  • Text Differences Between Logs (4 lines differ)
  • Log character count differs by 0% (gold=1849368, new=1849386); within tolerance
  • Timing regression for hcal: new=898s vs gold=770s (+16%, tolerance 10%)
  • Timing for hcal: gold=770s, new=898s

inclusive:

  • Text Differences Between Logs (28 lines differ)
  • Log character count differs by 0% (gold=23452407, new=23452425); within tolerance
  • Timing for inclusive: gold=6822s, new=6773s
  • Timing within 0% of gold (tolerance 10%)

it_pileup:

  • Text Differences Between Logs (44 lines differ)
  • Log character count differs by 0% (gold=8841824, new=8841842); within tolerance
  • Timing for it_pileup: gold=569s, new=578s
  • Timing within 1% of gold (tolerance 10%)

kaon_enhanced:

  • Text Differences Between Logs (60 lines differ)
  • Log character count differs by 0% (gold=11121319, new=11121337); within tolerance
  • Timing regression for kaon_enhanced: new=3043s vs gold=2119s (+43%, tolerance 10%)
  • Timing for kaon_enhanced: gold=2119s, new=3043s

reduced_ldmx:

  • Text Differences Between Logs (28 lines differ)
  • Log character count differs by 0% (gold=33305648, new=33305666); within tolerance
  • Timing for reduced_ldmx: gold=143s, new=150s
  • Timing within 4% of gold (tolerance 10%)

signal:

  • Text Differences Between Logs (38 lines differ)
  • Log character count differs by 0% (gold=18410869, new=18410887); within tolerance
  • Timing for signal: gold=1209s, new=1208s
  • Timing within 0% of gold (tolerance 10%)

signal_target_al:

  • Text Differences Between Logs (46 lines differ)
  • Log character count differs by 0% (gold=16986051, new=16986069); within tolerance
  • Timing regression for signal_target_al: new=910s vs gold=719s (+26%, tolerance 10%)
  • Timing for signal_target_al: gold=719s, new=910s

target_genie:

  • Text Differences Between Logs (172 lines differ)
  • Timing for target_genie: gold=6848s, new=6948s
  • Timing within 1% of gold (tolerance 10%)

target_pn_lyso:

  • Text Differences Between Logs (56 lines differ)
  • Log character count differs by 0% (gold=20743639, new=20743657); within tolerance
  • Timing regression for target_pn_lyso: new=6343s vs gold=3752s (+69%, tolerance 10%)
  • Timing for target_pn_lyso: gold=3752s, new=6343s

target_ti_en:

  • Text Differences Between Logs (46 lines differ)
  • Log character count differs by 0% (gold=398437, new=398455); within tolerance
  • Timing for target_ti_en: gold=2448s, new=2480s
  • Timing within 1% of gold (tolerance 10%)

wab_lhe:

  • Text Differences Between Logs (18 lines differ)
  • Log character count differs by 0% (gold=12553398, new=12553416); within tolerance
  • Timing for wab_lhe: gold=6804s, new=6794s
  • Timing within 0% of gold (tolerance 10%)

@tomeichlersmith tomeichlersmith left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am very nervous about this "solution" because, while it allows you to specifically avoid the pre-emption issues, it will also enable other types of erroneous runs to be "valid" and readable. I don't really want to get into the situation where jobs that are killed produce files that are the same as valid files. This will just encourage users to ignore errors/warnings which I don't believe will be helpful.

Maybe a better solution is to have a "resume from file" run mode? so that a pre-empted run can be completed in a deterministic manner? As I mentioned in the issue, I'm not confident that handling pre-emption by a specific job control system at a specific cluster should be in the purview of ldmx-sw.

@tvami

tvami commented Sep 14, 2026

Copy link
Copy Markdown
Member Author

I don't really want to get into the situation where jobs that are killed produce files that are the same as valid files. This will just encourage users to ignore errors/warnings which I don't believe will be helpful.

I could have a follow-up PR that by default will not read file where isComplete is false. That should address your worry.

I'm not confident that handling pre-emption by a specific job control system at a specific cluster should be in the purview of ldmx-sw.

While I agree with this statement in general, this specific system is the one that we use for big productions so that does mean we need to deal with it somehow and getting out of the preemption would require money which we dont have.

@tomeichlersmith tomeichlersmith left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok, I understand the reasoning and while I'm not happy about it, I agree that it is a necessary feature to make do with our current resources. I like the idea of having a check to force the user to acknowledge that they are processing an incomplete file.

@tvami

tvami commented Sep 14, 2026

Copy link
Copy Markdown
Member Author

Thanks, please feel free to edit the issue I just made: #2135

@tvami tvami changed the title Iss2129 fix en production problems Fix production problems (coming from preemption at SLAC) Sep 14, 2026
@jmmans

jmmans commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

How long is the user to be allowed to add to the run tree?

@jmmans

jmmans commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

In general, is the pre-emption issue for both simulation and reco or is this primarily a simulation issue? I would assert that it is quite dangerous to allow uncontrolled splitting of reco outputs without a strong database backend tracking provenance.

@tvami

tvami commented Sep 14, 2026

Copy link
Copy Markdown
Member Author

I would assert that it is quite dangerous to allow uncontrolled splitting of reco outputs without a strong database backend tracking provenance.

I totally agree with this!

But for making the GEN events, having 70% of the events produced is better than 0% (I have had preemptions where 30 min more would have finished everything). But also 20% statistics is better than nothing too. We just need to count event after the fact.

But yes this is not to be used for RECO.

@tomeichlersmith

Copy link
Copy Markdown
Member

How long is the user to be allowed to add to the run tree?

In beforeNewRun, so for simulation/production runs, just at the beginning. If there is more than one run, then periodically throughout program execution.

@jmmans

jmmans commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

For GEN/SIM, I would be VERY comfortable with allowing this. This is effectively the "origin" step. Any assumption that a given file has exactly a certain number of events is foolish. If there is an input file, I would be uncomfortable with this, but with no input file... pre-empt away. :-)

@tvami

tvami commented Sep 14, 2026

Copy link
Copy Markdown
Member Author

Now that you brought this up, maybe it should be moved to Process::newRun?

@tvami

tvami commented Sep 14, 2026

Copy link
Copy Markdown
Member Author

@tomeichlersmith @jmmans is f04cb1a better?

@tvami

tvami commented Sep 14, 2026

Copy link
Copy Markdown
Member Author

/run-validation

@github-actions

Copy link
Copy Markdown
Contributor

The validation workflow is running here: https://github.com/LDMX-Software/ldmx-sw/actions/runs/34880100119.

@github-actions

Copy link
Copy Markdown
Contributor

Validation Results

Some validation samples failed! ❌

Sample Status
cascade_history ✅ PASS
deep_ecal_gun ✅ PASS
eat_signal ✅ PASS
hcal ✅ PASS
inclusive ✅ PASS
it_pileup ✅ PASS
kaon_enhanced ✅ PASS
reduced_ldmx ✅ PASS
signal ✅ PASS
signal_target_al ✅ PASS
target_genie ✅ PASS
target_pn_lyso ✅ PASS
target_ti_en ✅ PASS
wab_lhe ✅ PASS
ecal_pn ❌ FAIL (1110 histograms failed KS test) (artifact)

ecal_pn:

  • 1110 plots failed the KS test against gold.
  • Text Differences Between Logs (258277 lines differ)
  • Log character count differs by 0% (gold=26263133, new=26301217); within tolerance
  • Timing regression for ecal_pn: new=5110s vs gold=4486s (+13%, tolerance 10%)
  • Timing for ecal_pn: gold=4486s, new=5110s

cascade_history:

  • Text Differences Between Logs (4 lines differ)
  • Log character count differs by 0% (gold=4450280, new=4450298); within tolerance
  • Timing for cascade_history: gold=6031s, new=5996s
  • Timing within 0% of gold (tolerance 10%)

deep_ecal_gun:

  • Text Differences Between Logs (28 lines differ)
  • Log character count differs by 0% (gold=11769584, new=11769602); within tolerance
  • Timing anomaly for deep_ecal_gun: new=3000s vs gold=4108s (-26%, tolerance 10%)
  • Timing for deep_ecal_gun: gold=4108s, new=3000s

eat_signal:

  • Timing regression for eat_signal: new=1412s vs gold=1262s (+11%, tolerance 10%)
  • Timing for eat_signal: gold=1262s, new=1412s

hcal:

  • Text Differences Between Logs (4 lines differ)
  • Log character count differs by 0% (gold=1849368, new=1849386); within tolerance
  • Timing regression for hcal: new=859s vs gold=770s (+11%, tolerance 10%)
  • Timing for hcal: gold=770s, new=859s

inclusive:

  • Text Differences Between Logs (54 lines differ)
  • Log character count differs by 0% (gold=23452407, new=23452424); within tolerance
  • Timing anomaly for inclusive: new=5671s vs gold=6822s (-16%, tolerance 10%)
  • Timing for inclusive: gold=6822s, new=5671s

it_pileup:

  • Text Differences Between Logs (42 lines differ)
  • Log character count differs by 0% (gold=8841824, new=8841842); within tolerance
  • Timing for it_pileup: gold=569s, new=598s
  • Timing within 5% of gold (tolerance 10%)

kaon_enhanced:

  • Text Differences Between Logs (58 lines differ)
  • Log character count differs by 0% (gold=11121319, new=11121337); within tolerance
  • Timing regression for kaon_enhanced: new=3117s vs gold=2119s (+47%, tolerance 10%)
  • Timing for kaon_enhanced: gold=2119s, new=3117s

reduced_ldmx:

  • Text Differences Between Logs (24 lines differ)
  • Log character count differs by 0% (gold=33305648, new=33305666); within tolerance
  • Timing for reduced_ldmx: gold=143s, new=130s
  • Timing within 9% of gold (tolerance 10%)

signal:

  • Text Differences Between Logs (48 lines differ)
  • Log character count differs by 0% (gold=18410869, new=18410887); within tolerance
  • Timing for signal: gold=1209s, new=1202s
  • Timing within 0% of gold (tolerance 10%)

signal_target_al:

  • Text Differences Between Logs (52 lines differ)
  • Log character count differs by 0% (gold=16986051, new=16986069); within tolerance
  • Timing regression for signal_target_al: new=956s vs gold=719s (+32%, tolerance 10%)
  • Timing for signal_target_al: gold=719s, new=956s

target_genie:

  • Text Differences Between Logs (172 lines differ)
  • Timing for target_genie: gold=6848s, new=6918s
  • Timing within 1% of gold (tolerance 10%)

target_pn_lyso:

  • Text Differences Between Logs (52 lines differ)
  • Log character count differs by 0% (gold=20743639, new=20743657); within tolerance
  • Timing regression for target_pn_lyso: new=6164s vs gold=3752s (+64%, tolerance 10%)
  • Timing for target_pn_lyso: gold=3752s, new=6164s

target_ti_en:

  • Text Differences Between Logs (54 lines differ)
  • Log character count differs by 0% (gold=398437, new=398455); within tolerance
  • Timing regression for target_ti_en: new=3198s vs gold=2448s (+30%, tolerance 10%)
  • Timing for target_ti_en: gold=2448s, new=3198s

wab_lhe:

  • Text Differences Between Logs (34 lines differ)
  • Log character count differs by 0% (gold=12553398, new=12553416); within tolerance
  • Timing for wab_lhe: gold=6804s, new=6512s
  • Timing within 4% of gold (tolerance 10%)

@tvami
tvami merged commit 907300a into trunk Sep 14, 2026
2 checks passed
@tvami
tvami deleted the iss2129-fix-en-production-problems branch September 14, 2026 21:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Fix problems seen while EN production on SDF

4 participants