Skip to content

LLM Post training - Question regarding seq-mask-tis and global_valid_toks denominator scaling. #905

Description

@entrpn

Reference: https://github.com/mlcommons/training/tree/master/llm_post_training

While looking at the GRPO loss implementation, I ran into something with the seq-mask-tis filter that I wanted to verify.

It looks like when truncated_importance_sampling_type == "seq-mask-tis", sequences that fall outside the band have their importance weights zeroed out, but those dropped sequences are still kept in the global_valid_toks denominator during the final loss reduction

Here is the section in ClippedPGLossFn.call where the mask is applied to the weights:

elif self.truncated_importance_sampling_type == "seq-mask-tis":
                # ... [snip] ...
                seq_kept_mask = (
                    (
                        seq_geomean_is_ratio
                        >= self.truncated_importance_sampling_ratio_min
                    )
                    & (seq_geomean_is_ratio <= self.truncated_importance_sampling_ratio)
                ).float()  # [B]
                # ...
                actor_importance_weights_expanded = (
                    actor_importance_weights_expanded * seq_kept_mask.unsqueeze(-1)
                )

And then immediately below, the loss is reduced using the upstream global_valid_toks:

if self.loss_type == LossType.TOKEN_LEVEL:
            actor_loss = masked_mean(
                importance_weights_to_use * clip_loss,
                mask,
                global_normalization_factor=global_valid_toks,
            )

Because global_valid_toks is calculated upstream (before this loss function zeroes out the actor_importance_weights_expanded), the dropped sequences are still acting as a denominator.

If the filter drops 30% of the sequences, isn't the overall gradient silently scaled down to 70% of nominal, rather than acting as a true average over just the surviving tokens.

I noticed that other sequence-dropping mechanisms (like the overlong filter) avoid this by zeroing out the sample_mask before global_valid_toks is calculated, which correctly removes them from the denominator.

Is leaving the seq-mask-tis drops in the denominator intentional here to act as a dynamic learning rate penalty, or should this be following the same convention as the overlong filter where dropped sequences are removed from the normalization factor?

Related: what is_oob_ratio does the reference typically see at the [0.999, 1.002] bounds? It's computed in _is_filter_metrics but doesn't appear in the published RCP logs, so it's hard to tell from outside how often this path is active.

Activity

  1. ShriyaRishab commented on Sep 4, 2026

    @ShriyaRishab
    Contributor

    @mmarcinkiewicz can you please address this?

  2. ShriyaRishab commented on Sep 10, 2026

    @ShriyaRishab
    Contributor

    @mmarcinkiewicz checking in again.

  3. jepio commented on Sep 11, 2026

    @jepio
    Contributor

    Hi @entrpn,

    If the filter drops 30% of the sequences, isn't the overall gradient silently scaled down to 70% of nominal, rather than acting as a true average over just the surviving tokens.
    Is leaving the seq-mask-tis drops in the denominator intentional here to act as a dynamic learning rate penalty, or should this be following the same convention as the overlong filter where dropped sequences are removed from the normalization factor?

    Yes, that interpretation is correct and this behavior is intentional. I surveyed other RL framework implementations, and they are in agreement about the implementation of importance sampling corrections from research like seq-mask-tis: they all keep the normalization the same, see the below table.

    Framework Support Sequence statistic / rule Denominator after rejection
    NeMo-RL Native seq-mask-tis Two-sided geometric-mean IS ratio Fixed: original valid-token or sequence count. [Code](https://github.com/NVIDIA-NeMo/RL/blob/5d49fbf4e74bdd3a6627c48759c4a66b8920743b/nemo_rl/algorithms/loss/loss_functions.py#L715-L770)
    OpenRLHF Native seq-mask-tis Two-sided geometric-mean IS ratio Fixed: original global token/sample count. [Mask](https://github.com/OpenRLHF/OpenRLHF/blob/0c550f39d215f6c859674d0d6d721e931397f6d5/openrlhf/models/loss.py#L195-L227), [reducer](https://github.com/OpenRLHF/OpenRLHF/blob/0c550f39d215f6c859674d0d6d721e931397f6d5/openrlhf/models/loss.py#L11-L39)
    verl Equivalent: Geo-RS-Token-TIS Configurable sequence mean, normally a two-sided geometric ratio Fixed in normal training: original global token/batch count. [Code](https://github.com/volcengine/verl/blob/10db40d0da4d59150bb389960b77585f81a89b8d/verl/trainer/ppo/core_algos.py#L2494-L2539)
    ms-swift Native rollout_importance_sampling_mode=sequence_mask One-sided geometric-mean ratio Fixed: sequence mask zeroes IS weights without changing completion_mask. [Code](https://github.com/modelscope/ms-swift/blob/0673cf75dca7d0b9b608b4a76632fb508ead5076/swift/rlhf_trainers/grpo_trainer.py#L2301-L2362)
    Hugging Face TRL Native vllm_importance_sampling_mode=sequence_mask Two-sided full-sequence product, not geometric mean Fixed: original completion mask and reducer denominator. [Mask](https://github.com/huggingface/trl/blob/f6229801adaa629816be2bc1aae51621d04d24a1/trl/trainer/grpo_trainer.py#L2681-L2725), [reducer](https://github.com/huggingface/trl/blob/f6229801adaa629816be2bc1aae51621d04d24a1/trl/trainer/grpo_trainer.py#L3233-L3266)
    Axolotl Via its Hugging Face TRL integration Same product-ratio sequence mask as TRL Fixed: same behavior as TRL. [Documentation](https://github.com/axolotl-ai-cloud/axolotl/blob/main/docs/rlhf.qmd)
    Miles Configurable equivalent through MIS/RS Geometric-level sequence rejection Fixed: original per-rollout rollout_mask_sums. [Mask](https://github.com/radixark/miles/blob/50ad28b47c8b7cd1f6bfed7767bb505a9fc8d9a9/examples/infra_features/train_infer_mismatch_helper/mis.py#L193-L271), [reducer](https://github.com/radixark/miles/blob/50ad28b47c8b7cd1f6bfed7767bb505a9fc8d9a9/miles/backends/training_utils/loss_hub/losses.py#L229-L284)
    SkyRL Configurable equivalent with sequence_mask_metric="geometric" Two-sided geometric-mean ratio Fixed: normalization uses the original loss mask before filtering. [Mask](https://github.com/NovaSky-AI/SkyRL/blob/52063b6567a350e02d326699cd7d0f380bbad466/skyrl/backends/skyrl_train/utils/off_policy_correction_utils.py#L191-L348), [normalization](https://github.com/NovaSky-AI/SkyRL/blob/52063b6567a350e02d326699cd7d0f380bbad466/skyrl/train/trainer.py#L1471-L1512)
    slime No exact seq-mask-TIS; related DeepSeek OPSM and configurable rejection One-sided mean log-ratio for negative-advantage sequences Fixed: original loss-mask denominator. [Code](https://github.com/THUDM/slime/blob/4c193f1f37509cca70f0e88807a9305b70f63f4e/slime/backends/megatron_utils/loss.py#L1009-L1104)
    PRIME-RL Not implemented — —
    TitanRL Not implemented — —

    I noticed that other sequence-dropping mechanisms (like the overlong filter) avoid this by zeroing out the sample_mask before global_valid_toks is calculated, which correctly removes them from the denominator.

    One could make the distinction between filters (change normalization) and corrections (keep normalization) but I agree that it is somewhat arbitrary. Frameworks appear to implement various custom sample filters and there's little commonality there. We will want to review these in the next submission round after v6.1.

    Related: what is_oob_ratio does the reference typically see at the [0.999, 1.002] bounds? It's computed in _is_filter_metrics but doesn't appear in the published RCP logs, so it's hard to tell from outside how often this path is active.

    It depends on GBS but this is what it looks like for the 3 GBS's evaluated during RCPs. The interesting thing is that in the reference configuration roughly the same number of samples pass the IS mask for each GBS in the number of steps the training targets.

    Image
  4. entrpn commented on Sep 29, 2026

    @entrpn
    Author

    @jepio thanks for answer my questions and pulling all the references from different frameworks. The image was really helpful too.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions