Repository navigation
Conversation
…rk support) (#2430) * Make opencv-python an optional dependency Register a lightweight Pillow/NumPy-based cv2 stub via library/cv2_compat when opencv-python is not installed, so existing `import cv2` call sites continue to work without modification. The default requirements.txt keeps opencv-python; a new requirements-no-opencv.txt skips it for users who prefer the lighter install. Tools that genuinely need real OpenCV (tools/canny.py, tools/detect_face_rotate.py, and the ControlNet canny preprocessor) now exit early with a clear bilingual message when it is missing. * feat: make cv2 stub imshow/waitKey interactive instead of no-op Debug viewers (gen_img.py, --debug_dataset) are entered explicitly by the user, so surfacing the image is the expected behavior. Display via PIL.Image.show and block on input() for pacing; map 'q' to keycode 27 so callers that branch on ESC keep working. * feat: import cv2_compat before cv2 in dataset and mask_generator * fix: make the cv2 stub reproduce OpenCV results for the training pipeline The Pillow-based fallback in library/_cv2_stub diverged from real OpenCV in ways that would change training results on installs without opencv-python (e.g. Windows on ARM64 / RTX Spark PCs, where no wheel is available): - cvtColor BGR2HSV clipped hues that round to 180 instead of wrapping to 0. - resize of RGBA images went through Pillow's "RGBA" mode, which premultiplies by alpha and alters the colour channels wherever alpha < 255 (affects alpha_mask datasets). Channels are now resampled independently. - resize with INTER_AREA / INTER_LINEAR (the dataset pipeline defaults) used Pillow's antialiased filters. They are now re-implemented in NumPy with OpenCV's exact sampling: true area averaging when shrinking on both axes, OpenCV's two-tap "area_mode" otherwise, and OpenCV's bilinear positions. Floating-point ties at pixel boundaries follow OpenCV's arithmetic order, so results match up to +-1 (verified on 3000 random cases and 24 MP images). Area averaging uses a fixed-width gather + sum, which is several times faster than cumsum/reduceat in NumPy (~0.15 s for a 24 MP image). - Single-channel (H, W, 1) input now returns a 2-D array like OpenCV. Adds tests/test_cv2_stub.py (equivalence tests skip without real OpenCV) and documents the behaviour and the Windows on ARM64 use case in the READMEs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Use tensorboardX in requirements-no-opencv.txt for Windows on ARM64 tensorboard 2.x depends on grpcio, which has no Windows ARM64 wheel, so on such platforms (e.g. RTX Spark PCs) pip silently resolves tensorboard to the ancient 1.10.0 release, which does not work with current protobuf. sd-scripts only touches TensorBoard through accelerate, and accelerate's tracker falls back to tensorboardX (pure Python) when torch.utils.tensorboard cannot be imported, so `--log_with tensorboard` keeps working; the event files can be viewed with TensorBoard on another machine. Documented in the READMEs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…s, make dadaptation optional in tests (#2431) * deps: bump schedulefree to 1.4.1 and safetensors to 0.8.0 for Windows ARM64 wheels Part of RTX Spark (Windows on ARM64) support. On win_arm64 pip can only install prebuilt wheels for these two packages from these versions on: - schedulefree 1.4 is published as an sdist only; 1.4.1 ships a universal wheel. Its Python code is identical to 1.4 apart from an added `algoperf` subpackage. - safetensors 0.4.5 has no win_arm64 wheel; 0.8.0 is the first release that ships one. The high-level `safetensors.torch` API used here is unchanged (verified with the optimizer / model-spec / inpainting tests and a save/load/safe_open smoke test including library.safetensors_utils). 0.8.0 requires Python >= 3.10, which is already the documented minimum. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * deps: bump transformers to 4.57.6 (tokenizers 0.22) for Windows ARM64 wheels transformers 4.54.1 requires tokenizers<0.22, and tokenizers only ships win_arm64 wheels from 0.22.2 on, so `pip install -r requirements.txt` fails on Windows on ARM64 (RTX Spark PCs) at dependency resolution. 4.57.6 is the latest 4.x release and accepts tokenizers 0.22.x. Verified with the full test suite and by comparing T5 / Qwen3 / CLIP tokenizer outputs with 4.54.1 + tokenizers 0.21.4 (identical). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * tests: skip the D-Adaptation optimizer cases when dadaptation is not installed dadaptation is not part of requirements.txt (CI installs it separately), so tests/test_optimizer.py failed at import on any plain install, e.g. on Windows on ARM64 (RTX Spark). The six D-Adaptation cases are now added only when the package is importable; the rest of the test still runs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
#2430 and #2431 were developed in parallel, so requirements-no-opencv.txt still had the old transformers / schedulefree / safetensors pins, which do not resolve on Windows on ARM64. Now identical to requirements.txt except for the opencv-python and tensorboard lines. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sync requirements-no-opencv.txt with the dependency bumps from #2431
…nt markers (Windows on ARM64 / RTX Spark support) Windows on ARM64 (e.g. NVIDIA RTX Spark PCs) has no wheels for opencv-python and for grpcio (a dependency of tensorboard 2.x). Instead of maintaining a separate requirements-no-opencv.txt, requirements.txt now uses environment markers so that a single `pip install -r requirements.txt` works everywhere: - opencv-python is skipped when platform_machine == "ARM64" (the value reported by Python on Windows on ARM64; Linux aarch64 and macOS arm64 are not affected and keep OpenCV) - tensorboard is replaced by tensorboardX on that platform requirements-no-opencv.txt is removed. Users on other platforms who want to avoid OpenCV can uninstall it after installing the requirements; the Pillow/NumPy fallback from #2430 takes over automatically. README / README-ja: rewrote the "Installing without OpenCV" section accordingly and added a change-history entry for the Windows on ARM64 support (#2430, #2431 and this PR). Verified by evaluating the markers with `packaging` for win_arm64 / win_amd64 / linux_aarch64 / macos_arm64 environments, and by resolving the resulting package sets with `pip install --dry-run --platform win_arm64` and `--platform win_amd64` (Python 3.13): both resolve fully and the versions of the shared packages are identical. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Merge requirements-no-opencv.txt into requirements.txt with environment markers (Windows on ARM64 / RTX Spark support)
Preparation for the transformers 5.x / diffusers 0.40 upgrade. The harness records the outputs of the code paths that depend on those libraries with the currently pinned versions and compares them after the upgrade: - text encoders of every model family (SD1/2, SDXL, SD3, FLUX.1, Lumina, HunyuanImage, Anima) through the trainers' own TokenizeStrategy / TextEncodingStrategy and family loaders, for a fixed prompt set (empty, short, tag list, Japanese + emoji, >77 tokens, >225 tokens), per prompt and batched; token ids / masks must match exactly, float outputs are compared with per-model tolerances and max abs / rel diff and cosine similarity are printed for every array - diffusers AutoencoderKL encode / decode of a fixed image for SD / SDXL - the diffusers noise schedulers used for sample generation (no weights needed): timesteps, alphas_cumprod, sigmas and a few step() calls Model paths live in tests/local/models.toml (git-ignored; template in models.example.toml), references in tests/local/references/ (git-ignored). `pytest` from the repository root does not descend into tests/local (norecursedirs); run `python tests/local/regression_te.py record|compare` or `pytest tests/local` explicitly. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ce Hub (SD2, Lumina) - stabilityai/stable-diffusion-2 (source of the SD2.x tokenizer) returns 401 on the Hub now, so v2 training failed to load the tokenizer without a warm cache. The SD2 tokenizer is the v1 (OpenAI CLIP) tokenizer with "!" (id 0) as the pad token: same vocabulary, same merges, same token ids. It is now built from openai/clip-vit-large-patch14 with pad_token="!" (verified: identical vocab, identical ids / attention masks incl. padding and truncation). The cache directory name under --tokenizer_cache_dir is kept so existing caches are still used. Same change in networks/lora_interrogator.py. - google/gemma-2-2b (Lumina tokenizer) is gated. The tokenizer files of the official Alpha-VLLM/Lumina-Image-2.0 repository are byte-identical (sha256) to those of google/gemma-2-2b, so the tokenizer is loaded from its tokenizer/ subfolder instead. - DIFFUSERS_REF_MODEL_ID_V2 (scheduler / tokenizer / vae reference when saving SD2 models in Diffusers format) now points to the sd2-community mirror of stabilityai/stable-diffusion-2-1. - strategy_base._load_tokenizer accepts a cache directory name override and extra from_pretrained kwargs for the above. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fix tokenizer sources that are no longer accessible on the Hugging Face Hub (SD2, Lumina)
…age glyph path load_byt5 returns (tokenizer, model). The new prompt makes the byT5 glyph encoder run on real text instead of the empty-token shortcut. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Add local-only regression harness for transformers / diffusers upgrades (tests/local)
….0, huggingface-hub to 1.32.0 Security maintenance update: the transformers 4.x line and diffusers < 0.38 no longer receive fixes (Dependabot: transformers < 5.3 RCE via config.json, < 5.5 LightGlue, diffusers < 0.38 trust_remote_code bypass). - requirements.txt: transformers==5.5.4, diffusers[torch]==0.40.0, accelerate==1.15.0, huggingface-hub==1.32.0. Resolves for win_amd64 and win_arm64 (RTX Spark). - transformers 5.6+ flattened CLIPTextModel (no `text_model` submodule), which would change the text encoder checkpoint keys and the LoRA weight names (lora_te_text_model_encoder_layers_*). 5.5.4 is the last release with the previous structure, so the pin stays there for now. - In transformers 5.x CLIPTokenizer is the fast tokenizer and no longer runs ftfy.fix_text like the slow tokenizer (and the original CLIP tokenizer) did, so curly quotes, full-width characters, HTML entities and mojibake were tokenized differently. library/clip_tokenizer.py provides a CLIPTokenizer subclass that applies the legacy normalization for the fast tokenizer (no-op for the slow one); all CLIP tokenizer users switched to it. tests/test_clip_tokenizer.py covers the normalization. - diffusers 0.40 removed the old module paths used by library/slicing_vae.py and gen_img_diffusers.py (diffusers.models.vae / unet_2d_blocks / unet_2d_condition / autoencoder_kl); updated to the current paths. - tests/local/regression_te.py: zero the padded positions of the HunyuanImage VLM embeddings before comparing (they are masked out downstream and their garbage values depend on the transformers version). - README / README-ja: change history entry. Verified with the local regression harness (tests/local, references recorded with transformers 4.57.6 / diffusers 0.32.1): text encoder outputs of SD1.5, SD2.1, SDXL, SD3, FLUX.1, Lumina, HunyuanImage and Anima, the SD/SDXL VAE and the noise schedulers are bit-identical with the new versions. pytest: 248 passed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…users 0.40 diffusers >= 0.40 warns on every `ModelMixin.to(dtype)` even when the model has no module to keep in float32 (`fp32_modules = self._keep_in_fp32_modules or []` followed by `if ... and fp32_modules is not None`, which is always true; still present on diffusers main). sd-scripts casts the VAE / U-Net with `.to(dtype)` on purpose, so every training and generation script printed the warning. library/utils.py now installs a logging filter on the diffusers logger that drops the warning only when the list is empty; a real "keep in float32" warning still shows. tests/test_diffusers_warning_filter.py covers both cases. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…gged "caf") Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…l if torch gets upgraded diffusers[torch]==0.40.0 requires torch >= 2.6. On the PyTorch 2.4.0 job pip silently upgraded torch to the latest release while installing requirements.txt and left torchvision 0.19 behind, so every import failed with "operator torchvision::nms does not exist". PyTorch 2.6.0 or later has been the documented requirement of sd-scripts already, so the 2.4.0 job is replaced by 2.8.0 (the version recommended for RTX 50 series GPUs). A new step asserts that the pinned torch version is still installed after requirements.txt, so such an upgrade fails with a clear message next time. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ers040 Update transformers to 5.5.4, diffusers to 0.40.0, accelerate to 1.15.0, huggingface-hub to 1.32.0
…r T5 attention transformers 5.6 flattened CLIPTextModel (the `text_model` submodule was removed), which changes the state dict keys, the LoRA/OFT module names of the text encoders (`lora_te_text_model_encoder_layers_*`) and breaks `text_encoder.text_model.*` access. Add library/clip_text_model.py with `CLIPTextModelWrapper`, which holds the flattened model as `text_model` and delegates forward / config / dtype / device / input embeddings / gradient checkpointing / save_pretrained to it. The wrapper is applied only when the model is flattened (transformers >= 5.6), so nothing changes for older versions. Construction sites (SD1/2, SDXL text encoder 1, FLUX CLIP-L and the Diffusers pipeline loaders) wrap the model, and the Diffusers-format savers unwrap it so model_index.json keeps `transformers.CLIPTextModel`. transformers 5.6 also switched T5 attention to SDPA by default, which changes the bf16 outputs of T5-XXL (FLUX.1 / SD3) and byT5 (HunyuanImage) slightly. Pin the eager implementation in the loaders so the (cached) text encoder outputs stay identical across versions. Verified with tests/local: all 9 entries are bit-identical to the transformers 4.57.6 references with both 5.5.4 and 5.17.0. LoRA module names are unchanged for SD1.5 and SDXL, and save_pretrained/from_pretrained round-trips with 5.17.0. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Drop the attn_implementation="eager" pin for T5-XXL / byT5: the SDPA path that transformers 5.6+ selects by default is as accurate as eager against fp32 and should be faster, so accept the slight bf16 output difference and document it in the README instead. Update the transformers pin to 5.17.0. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Support transformers 5.6+ (CLIPTextModel layout wrapper, eager T5 attention)
… without scipy `get_my_scheduler` (sample image generation during training) failed for two of the `--sample_sampler` choices with recent versions of diffusers: - `dpmsolver`: diffusers rejects its default `final_sigmas_type="zero"` for the non-++ algorithm. Pass `final_sigmas_type="sigma_min"` for this case (`dpmsolver++` keeps the default, so its output is unchanged). The same fix is applied to `--sampler dpmsolver` of gen_img.py / sdxl_gen_img.py. - `dpmsingle`: `DPMSolverSinglestepScheduler` has no `steps_offset` argument. Do not pass it, as gen_img.py already does. `lms` / `k_lms` need scipy, which is not in requirements.txt. Instead of failing at the first sample generation after the models are loaded, a new `check_sampler_requirements` raises a clear ImportError (with the pip command) at startup: in `verify_training_args` when `--sample_prompts` is set, and at the top of `main` of gen_img.py / sdxl_gen_img.py. tests/test_sampling_scheduler.py builds every sampler choice and checks the new behavior. The other samplers were verified to be unchanged with the local regression harness (tests/local, schedulers entry). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fix the dpmsolver / dpmsingle samplers and give a clear error for lms without scipy
…cript It has not worked since the refactoring removed `train_util.load_tokenizer` (the only thing it needed to run again), and gen_img.py supports everything it did except the experimental CLIP / VGG16 guidance. No user has reported the breakage, so rather than keeping 4,000 lines of duplicated code alive, remove it. The file remains available in the previous releases. The remaining references in the docs (train_network_README-ja / -zh and train_ti_README-ja) now point to gen_img.py. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Remove gen_img_diffusers.py, the old SD1.x / SD2.x image generation script
Turn the "changes planned for the next release" section into the 0.12.0 entry. Also add the missing Japanese entry for the transformers 5.6+ support (#2437) and drop the stale note that 5.6+ was not supported yet. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ll note Windows on ARM64 support first, then the dependency update, then the other changes. The previous library versions were verified to still work with this release (local regression harness and gen_img.py with transformers 4.57.6 / diffusers 0.32.1 / accelerate 1.6.0 / huggingface-hub 0.34.3), so the note now recommends updating soon rather than requiring it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
docs: Add version 0.12.0 to the change history
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release 0.12.0: merge dev into main.
Main changes since 0.11.1 (see the change history in README for details):
opencv-pythonis optional,requirements.txtselects ARM64-compatible packages with environment markers. Make opencv-python an optional dependency (Windows on ARM64 / RTX Spark support) #2430, Bump transformers/schedulefree/safetensors for Windows on ARM64 wheels, make dadaptation optional in tests #2431, Merge requirements-no-opencv.txt into requirements.txt with environment markers (Windows on ARM64 / RTX Spark support) #2433transformers4.57.6 → 5.17.0,diffusers0.32.1 → 0.40.0,accelerate1.6.0 → 1.15.0,huggingface-hub0.34.3 → 1.32.0. Requirespip install --upgrade -r requirements.txt. Update transformers to 5.5.4, diffusers to 0.40.0, accelerate to 1.15.0, huggingface-hub to 1.32.0 #2436, Support transformers 5.6+ (CLIPTextModel layout wrapper, eager T5 attention) #2437tests/local). Add local-only regression harness for transformers / diffusers upgrades (tests/local) #2434dpmsolver/dpmsinglesamplers, clear error forlmswithout scipy. Fix the dpmsolver / dpmsingle samplers and give a clear error for lms without scipy #2438gen_img_diffusers.py. Remove gen_img_diffusers.py, the old SD1.x / SD2.x image generation script #2439--show_timesteps_offset. feat: add per-subset timestep_bias for semantic-aware timestep sampling #2401, Add --show_timesteps_offset and document offset behavior under shift/flux_shift #2410Note: the README change-history PR (
doc-update-readme-for-release) should be merged into dev before this PR.🤖 Generated with Claude Code