Skip to content

Release 0.11.1 - #820

Merged
dkackman merged 18 commits into
masterfrom
develop
Oct 11, 2026
Merged

dkackman merged 18 commits into
masterfrom
develop

Conversation

@dkackman

Copy link
Copy Markdown
Owner

Release 0.11.1: the workflow speed range (v0.11.0..develop). Notes are drafted in docs/RELEASING.md under 0.11.1.

  • FLUX templates (14 plus flux-krea) stream the transformer from pinned host memory instead of sequential offload: 78 s against 273 s for a 25-step 1024x1024 image on a 3090. Z-Image Turbo takes model offload: 41 s against 107 s.
  • A LoRA scale/alpha or scheduler shift change re-applies on the warm pipeline instead of reloading. A workflow switch keeps the warm pipelines the next workflow loads.
  • restore_faces builds its face helper once per run, and picture frames decode with libav threading.
  • The four MiniMax-H3 baselines expose attention_backend (null by default).
  • Stale catalog step counts and caches fixed.

Verification: CI green on e351795. GPU-verified on lem (CUDA) and mini-ai (MPS). Preflight on the Mac passes except one UI e2e test, editor validates, saves into a new folder. It fails only on a box with z-image run history: the inherited-cost warning from #797 replaces the "schema-valid" line. That's a fixture-isolation gap from 0.11.0, not this range.

🤖 Generated with Claude Code

dkackman and others added 18 commits October 10, 2026 16:43
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…elease

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…were removed

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
per_frame calls restore_faces once per frame and the FaceRestoreHelper
constructor loads RetinaFace (and ParseNet with use_parse). Cache it through
cached_model and clean_all() before each read.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sequential decoders (video_shape count, count_video_frames,
decode_audio_video, decode_rgb_frames, read_thumbnails_and_track) set
thread_type AUTO. read_frames and read_frame_range stay single-threaded:
with frame threads a backward seek after a decoded frame yields stale
frames.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…pipeline instead of reloading

pipeline_cache_key now hashes weights_identity(), which drops a LoRA's
scale and alpha and a scheduler's shift; wrap_resident re-applies them
through Pipeline.apply_runtime_settings(). load_loras derives its names,
weights and alphas from adapter_settings() so a load and a hit cannot
disagree. 38 recorded template key rows regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…its source's runtime settings

weights_identity keeps an alpha's or a shift's presence (True, never the
value): a hit cannot restore a checkpoint default, so dropping one must
reload, while 16 -> 8 still re-applies in place.

The step cache folds a borrowed pipeline in by step_definition_keys - the
full definition hash (pipeline_definition_key), runtime settings included -
so a pipeline_reference or reused_components step misses when only its
source's LoRA scale, alpha or shift moved. Pipeline ownership stays on
step_pipeline_keys. 30 recorded template key rows regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
On a workflow switch the worker now asks workflow_run.pipeline_keys for
the pipelines the next workflow loads and keeps those resident; the rest,
the task model cache and the step cache are released as before. A failure
to prepare the next definition falls back to the full release.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ale, alpha and shift

Steps whose pipelines differ only in runtime settings now share one warm
model, and the last of them to run left its values on it. PipelineOwnership
records each step's realized pipeline definition beside its key, and
_reference_action re-applies the referenced step's values through a
module-level apply_runtime_settings(pipeline_definition, model) before
wrapping - from the workflow's definition, never the loading copy whose
LoRA entries lost model_name to load_loras.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The probe keyed the unrealized definition, so its keys never equalled the
resident pipelines' keys and the worker's keep path always kept nothing.
It now realizes the steps as cache_hits does. Also: stray blank line in
media.py, worker comment wording, release-note and guide clarifications.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…alog speed defaults

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…h set_adapters

diffusers' set_adapters() re-activates the adapters through peft's
set_adapter(), which sets requires_grad on the adapter weights. After a
run under model offload those are inference tensors and torch refuses
("Setting requires_grad=True on inference tensor outside InferenceMode"),
seen on lem the first time a scale changed on a warm Flux pipeline. A
warm pipeline holds the adapters active already, so a cache hit now
walks the loaded peft layers and calls set_scale() only; the cold path
keeps set_adapters().

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…Z-Image takes model offload, H3 baselines expose attention_backend

Measured on an RTX 3090 (docs/RECIPES_24GB.md): sequential offload costs
10.8 s/step on FLUX dev and 3.4x on Z-Image; model offload on unquantized
FLUX is the marginal fit #580 retired (491 MB free at 1024, OOM at four
images). A block-level, streamed group offload of the transformer from
pinned host memory runs at the resident per-step cost with 6 GB to spare,
so every unquantized FLUX template takes it, with the encoders and VAE
placed on the card explicitly. Z-Image fits model offload at 13 GB.

The four MiniMax-H3 family baselines expose attention_backend (null by
default; sage_hub measured ~30% per step). LTX-2.5 refuses sage attention
(attn_mask), so its templates expose nothing. Compact listing budget
10_650 -> 10_700 for the four variables.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The workflow-speed tests added eight string patch("dw...") targets
(325 -> 333) and develop CI has failed the architecture ratchet since the
merge. The runtime-settings tests now run the real alpha/scale path against
fake peft layers instead of stubbing it, emit_phase needs no patch outside a
run, and the two remaining stubs use patch.object on the imported module,
which fails loudly on a module move rather than silently.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The e2e server read the developer's real ~/.diffusers_helper, including the
job history plans price from. On a box that had run models/z-image, the
editor spec's validation carried the #797 inherited-cost warning instead of
"schema-valid" and failed, while CI (no history) passed. The settings root
is now scratch like the workflows, prompts and outputs already were.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dkackman
dkackman merged commit ca0bef9 into master Oct 11, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant