Repository navigation
LLM Post training - Question regarding seq-mask-tis and global_valid_toks denominator scaling. #905
Description
Activity
@mmarcinkiewicz can you please address this?
@mmarcinkiewicz checking in again.
Hi @entrpn,
If the filter drops 30% of the sequences, isn't the overall gradient silently scaled down to 70% of nominal, rather than acting as a true average over just the surviving tokens.
Is leaving the seq-mask-tis drops in the denominator intentional here to act as a dynamic learning rate penalty, or should this be following the same convention as the overlong filter where dropped sequences are removed from the normalization factor?Yes, that interpretation is correct and this behavior is intentional. I surveyed other RL framework implementations, and they are in agreement about the implementation of importance sampling corrections from research like seq-mask-tis: they all keep the normalization the same, see the below table.
I noticed that other sequence-dropping mechanisms (like the overlong filter) avoid this by zeroing out the sample_mask before global_valid_toks is calculated, which correctly removes them from the denominator.
One could make the distinction between filters (change normalization) and corrections (keep normalization) but I agree that it is somewhat arbitrary. Frameworks appear to implement various custom sample filters and there's little commonality there. We will want to review these in the next submission round after v6.1.
Related: what is_oob_ratio does the reference typically see at the [0.999, 1.002] bounds? It's computed in _is_filter_metrics but doesn't appear in the published RCP logs, so it's hard to tell from outside how often this path is active.
It depends on GBS but this is what it looks like for the 3 GBS's evaluated during RCPs. The interesting thing is that in the reference configuration roughly the same number of samples pass the IS mask for each GBS in the number of steps the training targets.

- added a commit that references this issue
on Sep 16, 2026 @jepio thanks for answer my questions and pulling all the references from different frameworks. The image was really helpful too.
Reference: https://github.com/mlcommons/training/tree/master/llm_post_training
While looking at the GRPO loss implementation, I ran into something with the seq-mask-tis filter that I wanted to verify.
It looks like when truncated_importance_sampling_type == "seq-mask-tis", sequences that fall outside the band have their importance weights zeroed out, but those dropped sequences are still kept in the global_valid_toks denominator during the final loss reduction
Here is the section in ClippedPGLossFn.call where the mask is applied to the weights:
And then immediately below, the loss is reduced using the upstream global_valid_toks:
Because global_valid_toks is calculated upstream (before this loss function zeroes out the actor_importance_weights_expanded), the dropped sequences are still acting as a denominator.
If the filter drops 30% of the sequences, isn't the overall gradient silently scaled down to 70% of nominal, rather than acting as a true average over just the surviving tokens.
I noticed that other sequence-dropping mechanisms (like the overlong filter) avoid this by zeroing out the sample_mask before global_valid_toks is calculated, which correctly removes them from the denominator.
Is leaving the seq-mask-tis drops in the denominator intentional here to act as a dynamic learning rate penalty, or should this be following the same convention as the overlong filter where dropped sequences are removed from the normalization factor?
Related: what is_oob_ratio does the reference typically see at the [0.999, 1.002] bounds? It's computed in _is_filter_metrics but doesn't appear in the published RCP logs, so it's hard to tell from outside how often this path is active.