Custom Node Testing
Expected Behavior
Repeated executions of the same workflow should have approximately stable sampling performance.
Actual Behavior
Without explicitly freeing/unloading models, the first run is fast but subsequent runs become roughly 3–4x slower per sampling step.
Calling /free between generations restores normal performance and keeps subsequent runs stable.
Steps to Reproduce
- Start ComfyUI with a clean process.
- Run the same 1024x1024 / 8-step Krea2 workflow repeatedly.
- Do not change prompt, model, resolution, sampler settings, or LoRAs.
- Compare sampling speed between the first and following generations.
Debug Logs
ComfyUI AMD/ROCm DynamicVRAM slowdown - relevant logs
========================================================================
[Environment / startup]
[INFO] Total VRAM 16304 MB, total RAM 65461 MB
[INFO] pytorch version: 2.12.0+rocm7.14.0
[INFO] Set: torch.backends.cudnn.enabled = False for better AMD performance.
[INFO] AMD arch: gfx1200
[INFO] ROCm version: (7, 14)
[INFO] Set vram state to: NORMAL_VRAM
[INFO] Device: cuda:0 AMD Radeon RX 9060 XT : native
[INFO] Using async weight offloading with 2 streams
[INFO] Enabled pinned memory 26184.0
[INFO] Using pytorch attention
[INFO] aimdo: src-win/cuda-detour.c:38:INFO:aimdo_setup_hooks: installing 6 hooks
[INFO] aimdo: src-win/shmem-detect.c:80:INFO:comfy-aimdo WDDM adapter match: AMD Radeon RX 9060 XT runtime_luid=00000000:0000f336 dxgi_luid=00000000:0000f336
[INFO] aimdo: src/control.c:284:INFO:comfy-aimdo inited for GPU: AMD Radeon RX 9060 XT (VRAM: 16304 MB)
[INFO] DynamicVRAM support detected and enabled
[INFO] Python version: 3.13.12 (main, Feb 12 2026, 00:38:53) [MSC v.1944 64 bit (AMD64)]
[INFO] ComfyUI version: 0.37.0
[INFO] comfy-aimdo version: 0.5.5
[Bad state: first run fast]
[INFO] got prompt
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\vae\\qwen_image_vae.safetensors']
[INFO] Using split attention in VAE
[INFO] Using split attention in VAE
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
[INFO] Found quantization metadata version 1
[INFO] Using MixedPrecisionOps for text encoder
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\text_encoders\\qwen3vl_4b_fp8_scaled.safetensors']
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
[INFO] Requested to load Krea2TEModel_
[INFO] Model Krea2TEModel_ prepared for dynamic VRAM loading. 4999MB Staged. 0 patches attached. Force pre-loaded 249 weights: 627 KB.
[INFO] Found quantization metadata version 1
[INFO] Detected mixed precision quantization
[INFO] Using mixed precision operations
[INFO] Native ops: asym_w4a8_int8, convrot_w4a4, float8_e5m2, int8_tensorwise, float8_e4m3fn , emulated ops: nvfp4, mxfp8
[INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16
[INFO] model_type FLUX
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\diffusion_models\\krea2_turbo_fp8_scaled.safetensors']
[INFO] Requested to load Krea2
[INFO] Model Krea2 prepared for dynamic VRAM loading. 12530MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB.
0%| | 0/8 [00:00<?, ?it/s]
0%| | 0/8 [00:00<?, ?it/s, Model Initializing ... ]
0%| | 0/8 [00:04<?, ?it/s, Model Initialization complete! ]
12%|█▎ | 1/8 [00:04<00:33, 4.74s/it, Model Initialization complete! ]
12%|█▎ | 1/8 [00:00<00:33, 4.74s/it]
25%|██▌ | 2/8 [00:00<00:00, 7.59it/s]
38%|███▊ | 3/8 [00:02<00:11, 2.27s/it]
50%|█████ | 4/8 [00:04<00:09, 2.27s/it]
62%|██████▎ | 5/8 [00:07<00:06, 2.26s/it]
75%|███████▌ | 6/8 [00:09<00:04, 2.27s/it]
88%|████████▊ | 7/8 [00:11<00:02, 2.27s/it]
100%|██████████| 8/8 [00:13<00:00, 2.27s/it]
100%|██████████| 8/8 [00:13<00:00, 1.73s/it]
[INFO] Requested to load WanVAE
[INFO] 0 models unloaded.
[INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
[INFO] Prompt executed in 33.52 seconds
[Bad state: second run slow]
[INFO] got prompt
[INFO] Model Krea2TEModel_ prepared for dynamic VRAM loading. 4999MB Staged. 0 patches attached. Force pre-loaded 249 weights: 627 KB.
[INFO] Model Krea2 prepared for dynamic VRAM loading. 12530MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB.
0%| | 0/8 [00:00<?, ?it/s]
0%| | 0/8 [00:00<?, ?it/s, Model Initializing ... ]
0%| | 0/8 [00:08<?, ?it/s, Model Initialization complete! ]
12%|█▎ | 1/8 [00:08<01:00, 8.65s/it, Model Initialization complete! ]
12%|█▎ | 1/8 [00:00<01:00, 8.65s/it]
25%|██▌ | 2/8 [00:00<00:02, 2.63it/s]
38%|███▊ | 3/8 [00:08<00:40, 8.16s/it]
50%|█████ | 4/8 [00:17<00:32, 8.16s/it]
62%|██████▎ | 5/8 [00:25<00:24, 8.15s/it]
75%|███████▌ | 6/8 [00:33<00:16, 8.14s/it]
88%|████████▊ | 7/8 [00:41<00:08, 8.15s/it]
100%|██████████| 8/8 [00:49<00:00, 8.13s/it]
100%|██████████| 8/8 [00:49<00:00, 6.21s/it]
[INFO] 0 models unloaded.
[INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
[INFO] Prompt executed in 70.31 seconds
[Bad state: additional slow runs]
[INFO] got prompt
[INFO] Model Krea2TEModel_ prepared for dynamic VRAM loading. 4999MB Staged. 0 patches attached. Force pre-loaded 249 weights: 627 KB.
[INFO] Model Krea2 prepared for dynamic VRAM loading. 12530MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB.
0%| | 0/8 [00:00<?, ?it/s]
0%| | 0/8 [00:00<?, ?it/s, Model Initializing ... ]
0%| | 0/8 [00:08<?, ?it/s, Model Initialization complete! ]
12%|█▎ | 1/8 [00:08<00:57, 8.27s/it, Model Initialization complete! ]
12%|█▎ | 1/8 [00:00<00:57, 8.27s/it]
25%|██▌ | 2/8 [00:00<00:02, 2.62it/s][INFO] FETCH ComfyRegistry Data [DONE]
[INFO] [ComfyUI-Manager] default cache updated: https://api.comfy.org/nodes
FETCH DATA from: C:\Users\Michael\AppData\Local\Comfy-Desktop\ComfyUI-Installs\ComfyUI\ComfyUI\user\__manager\cache\1514988643_custom-node-list.json [DONE]
[INFO] [ComfyUI-Manager] All startup tasks have been completed.
38%|███▊ | 3/8 [00:08<00:39, 7.94s/it]
50%|█████ | 4/8 [00:16<00:32, 8.04s/it]
62%|██████▎ | 5/8 [00:24<00:24, 8.02s/it]
75%|███████▌ | 6/8 [00:32<00:15, 7.96s/it]
88%|████████▊ | 7/8 [00:40<00:07, 7.97s/it]
100%|██████████| 8/8 [00:48<00:00, 7.98s/it]
100%|██████████| 8/8 [00:48<00:00, 6.08s/it]
[INFO] 0 models unloaded.
[INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
[INFO] Prompt executed in 69.28 seconds
[INFO] got prompt
[INFO] Model Krea2TEModel_ prepared for dynamic VRAM loading. 4999MB Staged. 0 patches attached. Force pre-loaded 249 weights: 627 KB.
[INFO] Model Krea2 prepared for dynamic VRAM loading. 12530MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB.
0%| | 0/8 [00:00<?, ?it/s]
0%| | 0/8 [00:00<?, ?it/s, Model Initializing ... ]
0%| | 0/8 [00:08<?, ?it/s, Model Initialization complete! ]
12%|█▎ | 1/8 [00:08<00:58, 8.32s/it, Model Initialization complete! ]
12%|█▎ | 1/8 [00:00<00:58, 8.32s/it]
25%|██▌ | 2/8 [00:00<00:02, 2.68it/s]
38%|███▊ | 3/8 [00:08<00:39, 7.98s/it]
50%|█████ | 4/8 [00:16<00:31, 7.95s/it]
62%|██████▎ | 5/8 [00:24<00:23, 7.99s/it]
75%|███████▌ | 6/8 [00:32<00:15, 7.98s/it]
88%|████████▊ | 7/8 [00:40<00:08, 8.01s/it]
100%|██████████| 8/8 [00:48<00:00, 7.94s/it]
100%|██████████| 8/8 [00:48<00:00, 6.07s/it]
[INFO] 0 models unloaded.
[INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
[INFO] Prompt executed in 69.26 seconds
[After /free between jobs: stable fast run, no LoRA]
[INFO] got prompt
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\vae\\qwen_image_vae.safetensors']
[INFO] Using split attention in VAE
[INFO] Using split attention in VAE
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
[INFO] Found quantization metadata version 1
[INFO] Using MixedPrecisionOps for text encoder
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\text_encoders\\qwen3vl_4b_fp8_scaled.safetensors']
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
[INFO] Requested to load Krea2TEModel_
[INFO] Model Krea2TEModel_ prepared for dynamic VRAM loading. 4999MB Staged. 0 patches attached. Force pre-loaded 249 weights: 627 KB.
[INFO] Found quantization metadata version 1
[INFO] Detected mixed precision quantization
[INFO] Using mixed precision operations
[INFO] Native ops: float8_e4m3fn, convrot_w4a4, float8_e5m2, int8_tensorwise, asym_w4a8_int8 , emulated ops: mxfp8, nvfp4
[INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16
[INFO] model_type FLUX
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\diffusion_models\\krea2_turbo_fp8_scaled.safetensors']
[INFO] Requested to load Krea2
[INFO] Model Krea2 prepared for dynamic VRAM loading. 12530MB Staged. 0 patches attached. Force pre-loaded 160 weights: 2824 KB.
0%| | 0/8 [00:00<?, ?it/s]
0%| | 0/8 [00:00<?, ?it/s, Model Initializing ... ]
0%| | 0/8 [00:03<?, ?it/s, Model Initialization complete! ]
12%|█▎ | 1/8 [00:03<00:24, 3.53s/it, Model Initialization complete! ]
12%|█▎ | 1/8 [00:00<00:24, 3.53s/it]
25%|██▌ | 2/8 [00:00<00:00, 6.76it/s]
38%|███▊ | 3/8 [00:02<00:12, 2.50s/it]
50%|█████ | 4/8 [00:05<00:09, 2.49s/it]
62%|██████▎ | 5/8 [00:07<00:07, 2.49s/it]
75%|███████▌ | 6/8 [00:10<00:04, 2.49s/it]
88%|████████▊ | 7/8 [00:12<00:02, 2.53s/it]
100%|██████████| 8/8 [00:15<00:00, 2.51s/it]
100%|██████████| 8/8 [00:15<00:00, 1.91s/it]
[INFO] Requested to load WanVAE
[INFO] 0 models unloaded.
[INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
[INFO] Prompt executed in 27.52 seconds
[After /free between jobs: stable fast run, LoRA active (256 patches)]
[INFO] got prompt
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\vae\\qwen_image_vae.safetensors']
[INFO] Using split attention in VAE
[INFO] Using split attention in VAE
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
[INFO] Found quantization metadata version 1
[INFO] Using MixedPrecisionOps for text encoder
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\text_encoders\\qwen3vl_4b_fp8_scaled.safetensors']
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
[INFO] Requested to load Krea2TEModel_
[INFO] Model Krea2TEModel_ prepared for dynamic VRAM loading. 4999MB Staged. 0 patches attached. Force pre-loaded 249 weights: 627 KB.
[INFO] Found quantization metadata version 1
[INFO] Detected mixed precision quantization
[INFO] Using mixed precision operations
[INFO] Native ops: float8_e4m3fn, convrot_w4a4, float8_e5m2, int8_tensorwise, asym_w4a8_int8 , emulated ops: mxfp8, nvfp4
[INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16
[INFO] model_type FLUX
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\diffusion_models\\krea2_turbo_fp8_scaled.safetensors']
[INFO] Requested to load Krea2
[INFO] Model Krea2 prepared for dynamic VRAM loading. 12530MB Staged. 256 patches attached. Force pre-loaded 160 weights: 2824 KB.
0%| | 0/8 [00:00<?, ?it/s]
0%| | 0/8 [00:00<?, ?it/s, Model Initializing ... ]
0%| | 0/8 [00:05<?, ?it/s, Model Initialization complete! ]
12%|█▎ | 1/8 [00:05<00:38, 5.44s/it, Model Initialization complete! ]
12%|█▎ | 1/8 [00:00<00:38, 5.44s/it]
25%|██▌ | 2/8 [00:00<00:02, 2.11it/s]
38%|███▊ | 3/8 [00:03<00:12, 2.45s/it]
50%|█████ | 4/8 [00:05<00:10, 2.60s/it]
62%|██████▎ | 5/8 [00:08<00:07, 2.60s/it]
75%|███████▌ | 6/8 [00:11<00:05, 2.61s/it]
88%|████████▊ | 7/8 [00:13<00:02, 2.61s/it]
100%|██████████| 8/8 [00:16<00:00, 2.61s/it]
100%|██████████| 8/8 [00:16<00:00, 2.05s/it]
[INFO] Requested to load WanVAE
[INFO] 0 models unloaded.
[INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
[INFO] Prompt executed in 29.54 seconds
[Additional stable LoRA runs]
[INFO] got prompt
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\vae\\qwen_image_vae.safetensors']
[INFO] Using split attention in VAE
[INFO] Using split attention in VAE
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
[INFO] Found quantization metadata version 1
[INFO] Using MixedPrecisionOps for text encoder
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\text_encoders\\qwen3vl_4b_fp8_scaled.safetensors']
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
[INFO] Requested to load Krea2TEModel_
[INFO] Model Krea2TEModel_ prepared for dynamic VRAM loading. 4999MB Staged. 0 patches attached. Force pre-loaded 249 weights: 627 KB.
[INFO] Found quantization metadata version 1
[INFO] Detected mixed precision quantization
[INFO] Using mixed precision operations
[INFO] Native ops: float8_e4m3fn, convrot_w4a4, float8_e5m2, int8_tensorwise, asym_w4a8_int8 , emulated ops: mxfp8, nvfp4
[INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16
[INFO] model_type FLUX
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\diffusion_models\\krea2_turbo_fp8_scaled.safetensors']
[INFO] Requested to load Krea2
[INFO] Model Krea2 prepared for dynamic VRAM loading. 12530MB Staged. 256 patches attached. Force pre-loaded 160 weights: 2824 KB.
0%| | 0/8 [00:00<?, ?it/s]
0%| | 0/8 [00:00<?, ?it/s, Model Initializing ... ]
0%| | 0/8 [00:05<?, ?it/s, Model Initialization complete! ]
12%|█▎ | 1/8 [00:05<00:38, 5.47s/it, Model Initialization complete! ]
12%|█▎ | 1/8 [00:00<00:38, 5.47s/it]
25%|██▌ | 2/8 [00:00<00:02, 2.15it/s]
38%|███▊ | 3/8 [00:03<00:12, 2.43s/it]
50%|█████ | 4/8 [00:05<00:10, 2.60s/it]
62%|██████▎ | 5/8 [00:08<00:07, 2.59s/it]
75%|███████▌ | 6/8 [00:11<00:05, 2.59s/it]
88%|████████▊ | 7/8 [00:13<00:02, 2.60s/it]
100%|██████████| 8/8 [00:16<00:00, 2.60s/it]
100%|██████████| 8/8 [00:16<00:00, 2.04s/it]
[INFO] Requested to load WanVAE
[INFO] 0 models unloaded.
[INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
[INFO] Prompt executed in 29.14 seconds
[INFO] Using RAM pressure cache.
[INFO] got prompt
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\vae\\qwen_image_vae.safetensors']
[INFO] Using split attention in VAE
[INFO] Using split attention in VAE
[INFO] VAE load device: cuda:0, offload device: cpu, dtype: torch.bfloat16
[INFO] Found quantization metadata version 1
[INFO] Using MixedPrecisionOps for text encoder
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\text_encoders\\qwen3vl_4b_fp8_scaled.safetensors']
[INFO] CLIP/text encoder model load device: cuda:0, offload device: cpu, current: cpu, dtype: torch.float16
[INFO] Requested to load Krea2TEModel_
[INFO] Model Krea2TEModel_ prepared for dynamic VRAM loading. 4999MB Staged. 0 patches attached. Force pre-loaded 249 weights: 627 KB.
[INFO] Found quantization metadata version 1
[INFO] Detected mixed precision quantization
[INFO] Using mixed precision operations
[INFO] Native ops: float8_e4m3fn, convrot_w4a4, float8_e5m2, int8_tensorwise, asym_w4a8_int8 , emulated ops: mxfp8, nvfp4
[INFO] model weight dtype torch.bfloat16, manual cast: torch.bfloat16
[INFO] model_type FLUX
[INFO] Model storage policy: fast_disk=True paths=['C:\\Users\\Michael\\AppData\\Local\\Comfy-Desktop\\ComfyUI-Shared\\models\\diffusion_models\\krea2_turbo_fp8_scaled.safetensors']
[INFO] Requested to load Krea2
[INFO] Model Krea2 prepared for dynamic VRAM loading. 12530MB Staged. 256 patches attached. Force pre-loaded 160 weights: 2824 KB.
0%| | 0/8 [00:00<?, ?it/s]
0%| | 0/8 [00:00<?, ?it/s, Model Initializing ... ]
0%| | 0/8 [00:05<?, ?it/s, Model Initialization complete! ]
12%|█▎ | 1/8 [00:05<00:38, 5.45s/it, Model Initialization complete! ]
12%|█▎ | 1/8 [00:00<00:38, 5.45s/it]
25%|██▌ | 2/8 [00:00<00:02, 2.14it/s]
38%|███▊ | 3/8 [00:03<00:12, 2.43s/it]
50%|█████ | 4/8 [00:05<00:10, 2.60s/it]
62%|██████▎ | 5/8 [00:08<00:07, 2.60s/it]
75%|███████▌ | 6/8 [00:11<00:05, 2.60s/it]
88%|████████▊ | 7/8 [00:13<00:02, 2.60s/it]
100%|██████████| 8/8 [00:16<00:00, 2.60s/it]
100%|██████████| 8/8 [00:16<00:00, 2.05s/it]
[INFO] Requested to load WanVAE
[INFO] 0 models unloaded.
[INFO] Model WanVAE prepared for dynamic VRAM loading. 241MB Staged. 0 patches attached. Force pre-loaded 60 weights: 61 KB.
[INFO] Prompt executed in 29.40 seconds
Other
Description
On AMD ROCm, image generation performance degrades heavily after the first generation in the same ComfyUI process.
The first generation runs normally, but subsequent generations become roughly 3–4x slower per sampling step, even with the exact same workflow and with no LoRAs applied.
Calling the /free endpoint with both unload_models and free_memory before the next generation restores normal performance.
Interestingly, when /free is used between every generation, performance remains stable and is even slightly faster overall than before the issue appeared.
This strongly suggests a problem related to memory residency / offloading / DynamicVRAM state rather than the sampler, prompt, or LoRA processing itself.
Environment
- GPU: AMD Radeon RX 9060 XT, 16 GB VRAM
- Windows
- ROCm 7.14
- PyTorch 2.12.0+rocm7.14.0
- ComfyUI 0.37.0
- comfy-aimdo 0.5.5
- DynamicVRAM enabled
- NORMAL_VRAM
- async weight offloading with 2 streams
- fast_disk enabled
- Model: Krea2 Turbo FP8
- Resolution: 1024x1024
- Sampler: 8 steps
ComfyUI reports DynamicVRAM as enabled and uses async weight offloading. :contentReference[oaicite:0]{index=0}
Reproduction
- Start ComfyUI with a clean process.
- Run the same 1024x1024 / 8-step Krea2 workflow repeatedly.
- Do not change prompt, model, resolution, sampler settings, or LoRAs.
- Compare sampling speed between the first and following generations.
Without /free
First generation:
- ~2.27 s/it
- 33.52 seconds total
Second generation:
- ~8.13–8.16 s/it
- 70.31 seconds total
Further generations remain around ~8 s/it and ~69 seconds total.
The model reports 0 patches attached, so LoRA patching is not involved in these slow runs. :contentReference[oaicite:1]{index=1}
Additional runs remain slow at approximately 7.9–8.0 s/it and ~69 seconds total. :contentReference[oaicite:2]{index=2}
Workaround
Calling:
POST /free
with:
{
"unload_models": true,
"free_memory": true
}
between every generation prevents the slowdown.
With this workaround enabled, repeated runs are stable:
- 27.52 seconds
- 29.54 seconds
- 29.54 seconds
- 29.14 seconds
- 29.40 seconds
- 29.70 seconds
Sampling remains around ~2.5–2.6 s/it. :contentReference[oaicite:3]{index=3} :contentReference[oaicite:4]{index=4} :contentReference[oaicite:5]{index=5}
LoRA comparison
The issue does not appear to be caused by LoRAs.
Without LoRA patches:
- Krea2 reports
0 patches attached
- sampling is ~2.5–2.6 s/it
- total time is ~27–29 seconds
With LoRA enabled:
- Krea2 reports
256 patches attached
- sampling remains ~2.6 s/it
- total time remains ~29–30 seconds
Example with 256 patches:
- initialization: ~5.44 s
- sampling: ~2.60–2.61 s/it
- total: 29.54 seconds
:contentReference[oaicite:6]{index=6}
Further LoRA runs show the same behavior:
- 29.14 seconds
- 29.40 seconds
- 29.70 seconds
:contentReference[oaicite:7]{index=7} :contentReference[oaicite:8]{index=8} :contentReference[oaicite:9]{index=9}
Observation
The important difference appears to be the runtime state between generations.
The workflow loads/stages approximately:
- Krea2TEModel_: 4999 MB
- Krea2: 12530 MB
- WanVAE: 241 MB
and all are reported as being prepared for dynamic VRAM loading. :contentReference[oaicite:10]{index=10}
Because /free consistently restores the fast state, my current suspicion is that some memory residency, offloading, allocator, or DynamicVRAM state persists after a generation and causes later generations to use a much slower execution path.
I cannot determine from the logs whether the underlying problem is in ComfyUI DynamicVRAM, comfy-aimdo, PyTorch/ROCm memory management, or WDDM interaction.
Expected behavior
Repeated executions of the same workflow should have approximately stable sampling performance.
Actual behavior
Without explicitly freeing/unloading models, the first run is fast but subsequent runs become roughly 3–4x slower per sampling step.
Calling /free between generations restores normal performance and keeps subsequent runs stable.
Custom Node Testing
Expected Behavior
Repeated executions of the same workflow should have approximately stable sampling performance.
Actual Behavior
Without explicitly freeing/unloading models, the first run is fast but subsequent runs become roughly 3–4x slower per sampling step.
Calling
/freebetween generations restores normal performance and keeps subsequent runs stable.Steps to Reproduce
Debug Logs
Other
Description
On AMD ROCm, image generation performance degrades heavily after the first generation in the same ComfyUI process.
The first generation runs normally, but subsequent generations become roughly 3–4x slower per sampling step, even with the exact same workflow and with no LoRAs applied.
Calling the
/freeendpoint with bothunload_modelsandfree_memorybefore the next generation restores normal performance.Interestingly, when
/freeis used between every generation, performance remains stable and is even slightly faster overall than before the issue appeared.This strongly suggests a problem related to memory residency / offloading / DynamicVRAM state rather than the sampler, prompt, or LoRA processing itself.
Environment
ComfyUI reports DynamicVRAM as enabled and uses async weight offloading. :contentReference[oaicite:0]{index=0}
Reproduction
Without
/freeFirst generation:
Second generation:
Further generations remain around ~8 s/it and ~69 seconds total.
The model reports
0 patches attached, so LoRA patching is not involved in these slow runs. :contentReference[oaicite:1]{index=1}Additional runs remain slow at approximately 7.9–8.0 s/it and ~69 seconds total. :contentReference[oaicite:2]{index=2}
Workaround
Calling:
POST /free
with:
{
"unload_models": true,
"free_memory": true
}
between every generation prevents the slowdown.
With this workaround enabled, repeated runs are stable:
Sampling remains around ~2.5–2.6 s/it. :contentReference[oaicite:3]{index=3} :contentReference[oaicite:4]{index=4} :contentReference[oaicite:5]{index=5}
LoRA comparison
The issue does not appear to be caused by LoRAs.
Without LoRA patches:
0 patches attachedWith LoRA enabled:
256 patches attachedExample with 256 patches:
:contentReference[oaicite:6]{index=6}
Further LoRA runs show the same behavior:
:contentReference[oaicite:7]{index=7} :contentReference[oaicite:8]{index=8} :contentReference[oaicite:9]{index=9}
Observation
The important difference appears to be the runtime state between generations.
The workflow loads/stages approximately:
and all are reported as being prepared for dynamic VRAM loading. :contentReference[oaicite:10]{index=10}
Because
/freeconsistently restores the fast state, my current suspicion is that some memory residency, offloading, allocator, or DynamicVRAM state persists after a generation and causes later generations to use a much slower execution path.I cannot determine from the logs whether the underlying problem is in ComfyUI DynamicVRAM, comfy-aimdo, PyTorch/ROCm memory management, or WDDM interaction.
Expected behavior
Repeated executions of the same workflow should have approximately stable sampling performance.
Actual behavior
Without explicitly freeing/unloading models, the first run is fast but subsequent runs become roughly 3–4x slower per sampling step.
Calling
/freebetween generations restores normal performance and keeps subsequent runs stable.