Skip to content

Fix model convert when use latest megatron - #2267

Open
alexqdh wants to merge 1 commit into
THUDM:mainfrom
alexqdh:fix_model_convert_use_latest_megatron
Open

Fix model convert when use latest megatron#2267
alexqdh wants to merge 1 commit into
THUDM:mainfrom
alexqdh:fix_model_convert_use_latest_megatron

Conversation

@alexqdh

@alexqdh alexqdh commented Aug 12, 2026

Copy link
Copy Markdown
Contributor
  • Add --use-gated-attention compatibility to the HF-to-torch_dist conversion tool while retaining --attention-output-gate.
  • Default enable_gloo_process_groups to True when it is unavailable in the Megatron argument namespace.
  • Support both old and new Megatron tokenizer utility import paths.
  • Allow the custom Qwen3.5 and Qwen3-Next attention modules to accept Megatron’s optional name argument.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant