Skip to content

Version 0.12.0 - #2441

Merged
kohya-ss merged 33 commits into
mainfrom
dev
Sep 24, 2026
Merged

kohya-ss merged 33 commits into
mainfrom
dev

Conversation

@kohya-ss

Copy link
Copy Markdown
Owner

Release 0.12.0: merge dev into main.

Main changes since 0.11.1 (see the change history in README for details):

Note: the README change-history PR (doc-update-readme-for-release) should be merged into dev before this PR.

🤖 Generated with Claude Code

kohya-ss and others added 30 commits September 22, 2026 14:50
…rk support) (#2430)

* Make opencv-python an optional dependency

Register a lightweight Pillow/NumPy-based cv2 stub via library/cv2_compat
when opencv-python is not installed, so existing `import cv2` call sites
continue to work without modification. The default requirements.txt keeps
opencv-python; a new requirements-no-opencv.txt skips it for users who
prefer the lighter install.

Tools that genuinely need real OpenCV (tools/canny.py,
tools/detect_face_rotate.py, and the ControlNet canny preprocessor)
now exit early with a clear bilingual message when it is missing.

* feat: make cv2 stub imshow/waitKey interactive instead of no-op

Debug viewers (gen_img.py, --debug_dataset) are entered explicitly by the
user, so surfacing the image is the expected behavior. Display via
PIL.Image.show and block on input() for pacing; map 'q' to keycode 27 so
callers that branch on ESC keep working.

* feat: import cv2_compat before cv2 in dataset and mask_generator

* fix: make the cv2 stub reproduce OpenCV results for the training pipeline

The Pillow-based fallback in library/_cv2_stub diverged from real OpenCV in
ways that would change training results on installs without opencv-python
(e.g. Windows on ARM64 / RTX Spark PCs, where no wheel is available):

- cvtColor BGR2HSV clipped hues that round to 180 instead of wrapping to 0.
- resize of RGBA images went through Pillow's "RGBA" mode, which
  premultiplies by alpha and alters the colour channels wherever alpha < 255
  (affects alpha_mask datasets). Channels are now resampled independently.
- resize with INTER_AREA / INTER_LINEAR (the dataset pipeline defaults) used
  Pillow's antialiased filters. They are now re-implemented in NumPy with
  OpenCV's exact sampling: true area averaging when shrinking on both axes,
  OpenCV's two-tap "area_mode" otherwise, and OpenCV's bilinear positions.
  Floating-point ties at pixel boundaries follow OpenCV's arithmetic order,
  so results match up to +-1 (verified on 3000 random cases and 24 MP images).
  Area averaging uses a fixed-width gather + sum, which is several times
  faster than cumsum/reduceat in NumPy (~0.15 s for a 24 MP image).
- Single-channel (H, W, 1) input now returns a 2-D array like OpenCV.

Adds tests/test_cv2_stub.py (equivalence tests skip without real OpenCV) and
documents the behaviour and the Windows on ARM64 use case in the READMEs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Use tensorboardX in requirements-no-opencv.txt for Windows on ARM64

tensorboard 2.x depends on grpcio, which has no Windows ARM64 wheel, so on
such platforms (e.g. RTX Spark PCs) pip silently resolves tensorboard to the
ancient 1.10.0 release, which does not work with current protobuf. sd-scripts
only touches TensorBoard through accelerate, and accelerate's tracker falls
back to tensorboardX (pure Python) when torch.utils.tensorboard cannot be
imported, so `--log_with tensorboard` keeps working; the event files can be
viewed with TensorBoard on another machine. Documented in the READMEs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…s, make dadaptation optional in tests (#2431)

* deps: bump schedulefree to 1.4.1 and safetensors to 0.8.0 for Windows ARM64 wheels

Part of RTX Spark (Windows on ARM64) support. On win_arm64 pip can only
install prebuilt wheels for these two packages from these versions on:

- schedulefree 1.4 is published as an sdist only; 1.4.1 ships a universal
  wheel. Its Python code is identical to 1.4 apart from an added
  `algoperf` subpackage.
- safetensors 0.4.5 has no win_arm64 wheel; 0.8.0 is the first release
  that ships one. The high-level `safetensors.torch` API used here is
  unchanged (verified with the optimizer / model-spec / inpainting tests
  and a save/load/safe_open smoke test including library.safetensors_utils).
  0.8.0 requires Python >= 3.10, which is already the documented minimum.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* deps: bump transformers to 4.57.6 (tokenizers 0.22) for Windows ARM64 wheels

transformers 4.54.1 requires tokenizers<0.22, and tokenizers only ships
win_arm64 wheels from 0.22.2 on, so `pip install -r requirements.txt` fails
on Windows on ARM64 (RTX Spark PCs) at dependency resolution. 4.57.6 is the
latest 4.x release and accepts tokenizers 0.22.x. Verified with the full
test suite and by comparing T5 / Qwen3 / CLIP tokenizer outputs with
4.54.1 + tokenizers 0.21.4 (identical).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* tests: skip the D-Adaptation optimizer cases when dadaptation is not installed

dadaptation is not part of requirements.txt (CI installs it separately), so
tests/test_optimizer.py failed at import on any plain install, e.g. on
Windows on ARM64 (RTX Spark). The six D-Adaptation cases are now added only
when the package is importable; the rest of the test still runs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
#2430 and #2431 were developed in parallel, so requirements-no-opencv.txt
still had the old transformers / schedulefree / safetensors pins, which do
not resolve on Windows on ARM64. Now identical to requirements.txt except
for the opencv-python and tensorboard lines.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sync requirements-no-opencv.txt with the dependency bumps from #2431
…nt markers (Windows on ARM64 / RTX Spark support)

Windows on ARM64 (e.g. NVIDIA RTX Spark PCs) has no wheels for opencv-python
and for grpcio (a dependency of tensorboard 2.x). Instead of maintaining a
separate requirements-no-opencv.txt, requirements.txt now uses environment
markers so that a single `pip install -r requirements.txt` works everywhere:

- opencv-python is skipped when platform_machine == "ARM64" (the value
  reported by Python on Windows on ARM64; Linux aarch64 and macOS arm64 are
  not affected and keep OpenCV)
- tensorboard is replaced by tensorboardX on that platform

requirements-no-opencv.txt is removed. Users on other platforms who want to
avoid OpenCV can uninstall it after installing the requirements; the
Pillow/NumPy fallback from #2430 takes over automatically.

README / README-ja: rewrote the "Installing without OpenCV" section
accordingly and added a change-history entry for the Windows on ARM64 support
(#2430, #2431 and this PR).

Verified by evaluating the markers with `packaging` for win_arm64 / win_amd64 /
linux_aarch64 / macos_arm64 environments, and by resolving the resulting
package sets with `pip install --dry-run --platform win_arm64` and
`--platform win_amd64` (Python 3.13): both resolve fully and the versions of
the shared packages are identical.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Merge requirements-no-opencv.txt into requirements.txt with environment markers (Windows on ARM64 / RTX Spark support)
Preparation for the transformers 5.x / diffusers 0.40 upgrade. The harness
records the outputs of the code paths that depend on those libraries with the
currently pinned versions and compares them after the upgrade:

- text encoders of every model family (SD1/2, SDXL, SD3, FLUX.1, Lumina,
  HunyuanImage, Anima) through the trainers' own TokenizeStrategy /
  TextEncodingStrategy and family loaders, for a fixed prompt set (empty,
  short, tag list, Japanese + emoji, >77 tokens, >225 tokens), per prompt and
  batched; token ids / masks must match exactly, float outputs are compared
  with per-model tolerances and max abs / rel diff and cosine similarity are
  printed for every array
- diffusers AutoencoderKL encode / decode of a fixed image for SD / SDXL
- the diffusers noise schedulers used for sample generation (no weights
  needed): timesteps, alphas_cumprod, sigmas and a few step() calls

Model paths live in tests/local/models.toml (git-ignored; template in
models.example.toml), references in tests/local/references/ (git-ignored).
`pytest` from the repository root does not descend into tests/local
(norecursedirs); run `python tests/local/regression_te.py record|compare` or
`pytest tests/local` explicitly.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ce Hub (SD2, Lumina)

- stabilityai/stable-diffusion-2 (source of the SD2.x tokenizer) returns 401 on
  the Hub now, so v2 training failed to load the tokenizer without a warm cache.
  The SD2 tokenizer is the v1 (OpenAI CLIP) tokenizer with "!" (id 0) as the
  pad token: same vocabulary, same merges, same token ids. It is now built from
  openai/clip-vit-large-patch14 with pad_token="!" (verified: identical vocab,
  identical ids / attention masks incl. padding and truncation). The cache
  directory name under --tokenizer_cache_dir is kept so existing caches are
  still used. Same change in networks/lora_interrogator.py.
- google/gemma-2-2b (Lumina tokenizer) is gated. The tokenizer files of the
  official Alpha-VLLM/Lumina-Image-2.0 repository are byte-identical (sha256)
  to those of google/gemma-2-2b, so the tokenizer is loaded from its tokenizer/
  subfolder instead.
- DIFFUSERS_REF_MODEL_ID_V2 (scheduler / tokenizer / vae reference when saving
  SD2 models in Diffusers format) now points to the sd2-community mirror of
  stabilityai/stable-diffusion-2-1.
- strategy_base._load_tokenizer accepts a cache directory name override and
  extra from_pretrained kwargs for the above.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fix tokenizer sources that are no longer accessible on the Hugging Face Hub (SD2, Lumina)
…age glyph path

load_byt5 returns (tokenizer, model). The new prompt makes the byT5 glyph
encoder run on real text instead of the empty-token shortcut.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Add local-only regression harness for transformers / diffusers upgrades (tests/local)
….0, huggingface-hub to 1.32.0

Security maintenance update: the transformers 4.x line and diffusers < 0.38
no longer receive fixes (Dependabot: transformers < 5.3 RCE via config.json,
< 5.5 LightGlue, diffusers < 0.38 trust_remote_code bypass).

- requirements.txt: transformers==5.5.4, diffusers[torch]==0.40.0,
  accelerate==1.15.0, huggingface-hub==1.32.0. Resolves for win_amd64 and
  win_arm64 (RTX Spark).
- transformers 5.6+ flattened CLIPTextModel (no `text_model` submodule),
  which would change the text encoder checkpoint keys and the LoRA weight
  names (lora_te_text_model_encoder_layers_*). 5.5.4 is the last release
  with the previous structure, so the pin stays there for now.
- In transformers 5.x CLIPTokenizer is the fast tokenizer and no longer runs
  ftfy.fix_text like the slow tokenizer (and the original CLIP tokenizer)
  did, so curly quotes, full-width characters, HTML entities and mojibake
  were tokenized differently. library/clip_tokenizer.py provides a
  CLIPTokenizer subclass that applies the legacy normalization for the fast
  tokenizer (no-op for the slow one); all CLIP tokenizer users switched to
  it. tests/test_clip_tokenizer.py covers the normalization.
- diffusers 0.40 removed the old module paths used by library/slicing_vae.py
  and gen_img_diffusers.py (diffusers.models.vae / unet_2d_blocks /
  unet_2d_condition / autoencoder_kl); updated to the current paths.
- tests/local/regression_te.py: zero the padded positions of the
  HunyuanImage VLM embeddings before comparing (they are masked out
  downstream and their garbage values depend on the transformers version).
- README / README-ja: change history entry.

Verified with the local regression harness (tests/local, references recorded
with transformers 4.57.6 / diffusers 0.32.1): text encoder outputs of SD1.5,
SD2.1, SDXL, SD3, FLUX.1, Lumina, HunyuanImage and Anima, the SD/SDXL VAE and
the noise schedulers are bit-identical with the new versions. pytest: 248
passed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…users 0.40

diffusers >= 0.40 warns on every `ModelMixin.to(dtype)` even when the model
has no module to keep in float32 (`fp32_modules = self._keep_in_fp32_modules
or []` followed by `if ... and fp32_modules is not None`, which is always
true; still present on diffusers main). sd-scripts casts the VAE / U-Net with
`.to(dtype)` on purpose, so every training and generation script printed the
warning. library/utils.py now installs a logging filter on the diffusers
logger that drops the warning only when the list is empty; a real "keep in
float32" warning still shows. tests/test_diffusers_warning_filter.py covers
both cases.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…gged "caf")

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…l if torch gets upgraded

diffusers[torch]==0.40.0 requires torch >= 2.6. On the PyTorch 2.4.0 job pip
silently upgraded torch to the latest release while installing
requirements.txt and left torchvision 0.19 behind, so every import failed
with "operator torchvision::nms does not exist". PyTorch 2.6.0 or later has
been the documented requirement of sd-scripts already, so the 2.4.0 job is
replaced by 2.8.0 (the version recommended for RTX 50 series GPUs). A new
step asserts that the pinned torch version is still installed after
requirements.txt, so such an upgrade fails with a clear message next time.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ers040

Update transformers to 5.5.4, diffusers to 0.40.0, accelerate to 1.15.0, huggingface-hub to 1.32.0
…r T5 attention

transformers 5.6 flattened CLIPTextModel (the `text_model` submodule was removed),
which changes the state dict keys, the LoRA/OFT module names of the text encoders
(`lora_te_text_model_encoder_layers_*`) and breaks `text_encoder.text_model.*` access.

Add library/clip_text_model.py with `CLIPTextModelWrapper`, which holds the
flattened model as `text_model` and delegates forward / config / dtype / device /
input embeddings / gradient checkpointing / save_pretrained to it. The wrapper is
applied only when the model is flattened (transformers >= 5.6), so nothing changes
for older versions. Construction sites (SD1/2, SDXL text encoder 1, FLUX CLIP-L and
the Diffusers pipeline loaders) wrap the model, and the Diffusers-format savers
unwrap it so model_index.json keeps `transformers.CLIPTextModel`.

transformers 5.6 also switched T5 attention to SDPA by default, which changes the
bf16 outputs of T5-XXL (FLUX.1 / SD3) and byT5 (HunyuanImage) slightly. Pin the
eager implementation in the loaders so the (cached) text encoder outputs stay
identical across versions.

Verified with tests/local: all 9 entries are bit-identical to the transformers
4.57.6 references with both 5.5.4 and 5.17.0. LoRA module names are unchanged for
SD1.5 and SDXL, and save_pretrained/from_pretrained round-trips with 5.17.0.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Drop the attn_implementation="eager" pin for T5-XXL / byT5: the SDPA path that
transformers 5.6+ selects by default is as accurate as eager against fp32 and
should be faster, so accept the slight bf16 output difference and document it
in the README instead. Update the transformers pin to 5.17.0.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Support transformers 5.6+ (CLIPTextModel layout wrapper, eager T5 attention)
… without scipy

`get_my_scheduler` (sample image generation during training) failed for two
of the `--sample_sampler` choices with recent versions of diffusers:

- `dpmsolver`: diffusers rejects its default `final_sigmas_type="zero"` for
  the non-++ algorithm. Pass `final_sigmas_type="sigma_min"` for this case
  (`dpmsolver++` keeps the default, so its output is unchanged). The same
  fix is applied to `--sampler dpmsolver` of gen_img.py / sdxl_gen_img.py.
- `dpmsingle`: `DPMSolverSinglestepScheduler` has no `steps_offset`
  argument. Do not pass it, as gen_img.py already does.

`lms` / `k_lms` need scipy, which is not in requirements.txt. Instead of
failing at the first sample generation after the models are loaded, a new
`check_sampler_requirements` raises a clear ImportError (with the pip
command) at startup: in `verify_training_args` when `--sample_prompts` is
set, and at the top of `main` of gen_img.py / sdxl_gen_img.py.

tests/test_sampling_scheduler.py builds every sampler choice and checks the
new behavior. The other samplers were verified to be unchanged with the
local regression harness (tests/local, schedulers entry).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fix the dpmsolver / dpmsingle samplers and give a clear error for lms without scipy
…cript

It has not worked since the refactoring removed `train_util.load_tokenizer`
(the only thing it needed to run again), and gen_img.py supports everything
it did except the experimental CLIP / VGG16 guidance. No user has reported
the breakage, so rather than keeping 4,000 lines of duplicated code alive,
remove it. The file remains available in the previous releases.

The remaining references in the docs (train_network_README-ja / -zh and
train_ti_README-ja) now point to gen_img.py.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Remove gen_img_diffusers.py, the old SD1.x / SD2.x image generation script
kohya-ss and others added 3 commits September 24, 2026 19:08
Turn the "changes planned for the next release" section into the 0.12.0
entry. Also add the missing Japanese entry for the transformers 5.6+ support
(#2437) and drop the stale note that 5.6+ was not supported yet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ll note

Windows on ARM64 support first, then the dependency update, then the other
changes. The previous library versions were verified to still work with this
release (local regression harness and gen_img.py with transformers 4.57.6 /
diffusers 0.32.1 / accelerate 1.6.0 / huggingface-hub 0.34.3), so the note
now recommends updating soon rather than requiring it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
docs: Add version 0.12.0 to the change history
@kohya-ss
kohya-ss merged commit 690ea7f into main Sep 24, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant