Repository navigation
Conversation
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…elease Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…were removed Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
per_frame calls restore_faces once per frame and the FaceRestoreHelper constructor loads RetinaFace (and ParseNet with use_parse). Cache it through cached_model and clean_all() before each read. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sequential decoders (video_shape count, count_video_frames, decode_audio_video, decode_rgb_frames, read_thumbnails_and_track) set thread_type AUTO. read_frames and read_frame_range stay single-threaded: with frame threads a backward seek after a decoded frame yields stale frames. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…pipeline instead of reloading pipeline_cache_key now hashes weights_identity(), which drops a LoRA's scale and alpha and a scheduler's shift; wrap_resident re-applies them through Pipeline.apply_runtime_settings(). load_loras derives its names, weights and alphas from adapter_settings() so a load and a hit cannot disagree. 38 recorded template key rows regenerated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…its source's runtime settings weights_identity keeps an alpha's or a shift's presence (True, never the value): a hit cannot restore a checkpoint default, so dropping one must reload, while 16 -> 8 still re-applies in place. The step cache folds a borrowed pipeline in by step_definition_keys - the full definition hash (pipeline_definition_key), runtime settings included - so a pipeline_reference or reused_components step misses when only its source's LoRA scale, alpha or shift moved. Pipeline ownership stays on step_pipeline_keys. 30 recorded template key rows regenerated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
On a workflow switch the worker now asks workflow_run.pipeline_keys for the pipelines the next workflow loads and keeps those resident; the rest, the task model cache and the step cache are released as before. A failure to prepare the next definition falls back to the full release. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ale, alpha and shift Steps whose pipelines differ only in runtime settings now share one warm model, and the last of them to run left its values on it. PipelineOwnership records each step's realized pipeline definition beside its key, and _reference_action re-applies the referenced step's values through a module-level apply_runtime_settings(pipeline_definition, model) before wrapping - from the workflow's definition, never the loading copy whose LoRA entries lost model_name to load_loras. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The probe keyed the unrealized definition, so its keys never equalled the resident pipelines' keys and the worker's keep path always kept nothing. It now realizes the steps as cache_hits does. Also: stray blank line in media.py, worker comment wording, release-note and guide clarifications. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…alog speed defaults Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…h set_adapters
diffusers' set_adapters() re-activates the adapters through peft's
set_adapter(), which sets requires_grad on the adapter weights. After a
run under model offload those are inference tensors and torch refuses
("Setting requires_grad=True on inference tensor outside InferenceMode"),
seen on lem the first time a scale changed on a warm Flux pipeline. A
warm pipeline holds the adapters active already, so a cache hit now
walks the loaded peft layers and calls set_scale() only; the cold path
keeps set_adapters().
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…Z-Image takes model offload, H3 baselines expose attention_backend Measured on an RTX 3090 (docs/RECIPES_24GB.md): sequential offload costs 10.8 s/step on FLUX dev and 3.4x on Z-Image; model offload on unquantized FLUX is the marginal fit #580 retired (491 MB free at 1024, OOM at four images). A block-level, streamed group offload of the transformer from pinned host memory runs at the resident per-step cost with 6 GB to spare, so every unquantized FLUX template takes it, with the encoders and VAE placed on the card explicitly. Z-Image fits model offload at 13 GB. The four MiniMax-H3 family baselines expose attention_backend (null by default; sage_hub measured ~30% per step). LTX-2.5 refuses sage attention (attn_mask), so its templates expose nothing. Compact listing budget 10_650 -> 10_700 for the four variables. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The workflow-speed tests added eight string patch("dw...") targets
(325 -> 333) and develop CI has failed the architecture ratchet since the
merge. The runtime-settings tests now run the real alpha/scale path against
fake peft layers instead of stubbing it, emit_phase needs no patch outside a
run, and the two remaining stubs use patch.object on the imported module,
which fails loudly on a module move rather than silently.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The e2e server read the developer's real ~/.diffusers_helper, including the job history plans price from. On a box that had run models/z-image, the editor spec's validation carried the #797 inherited-cost warning instead of "schema-valid" and failed, while CI (no history) passed. The settings root is now scratch like the workflows, prompts and outputs already were. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release 0.11.1: the workflow speed range (v0.11.0..develop). Notes are drafted in docs/RELEASING.md under 0.11.1.
attention_backend(null by default).Verification: CI green on e351795. GPU-verified on lem (CUDA) and mini-ai (MPS). Preflight on the Mac passes except one UI e2e test,
editor validates, saves into a new folder. It fails only on a box with z-image run history: the inherited-cost warning from #797 replaces the "schema-valid" line. That's a fixture-isolation gap from 0.11.0, not this range.🤖 Generated with Claude Code