Fix production problems (coming from preemption at SLAC) - #2130
Conversation
Validation ResultsAll validation samples passed! ✅
|
tomeichlersmith
left a comment
There was a problem hiding this comment.
I am very nervous about this "solution" because, while it allows you to specifically avoid the pre-emption issues, it will also enable other types of erroneous runs to be "valid" and readable. I don't really want to get into the situation where jobs that are killed produce files that are the same as valid files. This will just encourage users to ignore errors/warnings which I don't believe will be helpful.
Maybe a better solution is to have a "resume from file" run mode? so that a pre-empted run can be completed in a deterministic manner? As I mentioned in the issue, I'm not confident that handling pre-emption by a specific job control system at a specific cluster should be in the purview of ldmx-sw.
I could have a follow-up PR that by default will not read file where
While I agree with this statement in general, this specific system is the one that we use for big productions so that does mean we need to deal with it somehow and getting out of the preemption would require money which we dont have. |
tomeichlersmith
left a comment
There was a problem hiding this comment.
Ok, I understand the reasoning and while I'm not happy about it, I agree that it is a necessary feature to make do with our current resources. I like the idea of having a check to force the user to acknowledge that they are processing an incomplete file.
|
Thanks, please feel free to edit the issue I just made: #2135 |
|
How long is the user to be allowed to add to the run tree? |
|
In general, is the pre-emption issue for both simulation and reco or is this primarily a simulation issue? I would assert that it is quite dangerous to allow uncontrolled splitting of reco outputs without a strong database backend tracking provenance. |
I totally agree with this! But for making the GEN events, having 70% of the events produced is better than 0% (I have had preemptions where 30 min more would have finished everything). But also 20% statistics is better than nothing too. We just need to count event after the fact. But yes this is not to be used for RECO. |
In |
|
For GEN/SIM, I would be VERY comfortable with allowing this. This is effectively the "origin" step. Any assumption that a given file has exactly a certain number of events is foolish. If there is an input file, I would be uncomfortable with this, but with no input file... pre-empt away. :-) |
|
Now that you brought this up, maybe it should be moved to |
|
@tomeichlersmith @jmmans is f04cb1a better? |
|
/run-validation |
|
The validation workflow is running here: https://github.com/LDMX-Software/ldmx-sw/actions/runs/34880100119. |
Validation ResultsSome validation samples failed! ❌
|
I am updating ldmx-sw, here are the details.
What are the issues that this addresses?
Resolves #2129
Check List