Conversation
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
| return | ||
| self.stop.wait(0.1) |
There was a problem hiding this comment.
Failed samples report zero memory If the cgroup memory files are unavailable or a read fails, the sampling thread stops, but the benchmark still records host-memory peaks as 0.0 GiB. That can make an unmeasured run look like a valid low-memory result. Mark the measurement unavailable or fail the run instead of publishing zero peaks.
Prompt To Fix With AI
This is a comment left during a code review.
Path: scripts/benchmarks/minimax_h3_4090/bench_pod.py
Line: 48-49
Comment:
**Failed samples report zero memory** If the cgroup memory files are unavailable or a read fails, the sampling thread stops, but the benchmark still records host-memory peaks as 0.0 GiB. That can make an unmeasured run look like a valid low-memory result. Mark the measurement unavailable or fail the run instead of publishing zero peaks.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
FastH3's FP8 consumer path previously spent substantial time copying attention layouts, decoding INT8 VAE projections, and moving the text encoder. This adds opt-in RTX 4090 kernels and memory controls that reduce warmed end-to-end generation to 41.75 s for 124 frames at 832×480 and 79.67 s for 243 frames at 832×480, including audio and MP4 export.
Stacked on the shared release core in #45 (
h3-release-core,a97d23f09). This is a draft staging PR in the fork; upstream submission follows the shared-core split.Changes:
Measured on one RTX 4090, checkpoint
FastH3-Pruned-8Step-FP8-ckpt300, sourcefb92af176, eight DMD forwards, sparsity 0.8, tile 64, light H3 VAE. Each configuration uses one warmup and two timed prompts; construction and cold compilation are outside the headline median.Earlier 16 and 12 GiB PyTorch allocator-cap trials completed at 104.03 and 107.27 s for the 243-frame 480p clip. These emulate capacity, not the throughput of smaller GPUs. The latest 7.25 GiB allocator-cap trial exceeded the strict 8 GiB total-device target (sampled NVML ~8.28 GiB); further tightening and the updated 768p benchmark are pending. The pod stopped accepting SSH connections before those results could be collected; the last verified 768p median remains 279.94 s. Current uncapped recipes use essentially all 24 GB of the board. A 32 GB system-RAM configuration has not been established.
Validation:
Exact commands, environment flags, checkpoint identity, historical measurements, limitations, and reproduction details are in
scripts/benchmarks/minimax_h3_4090/README.md.The PR appears safe to merge, with a non-blocking benchmark reporting issue to address.
Fix with agent prompt
Summary
This PR adds opt-in consumer-GPU attention, encoder-streaming, pinned-memory, and VAE INT8 optimizations for FastH3, alongside tests and RTX 4090 benchmark recipes. The default attention route remains unchanged.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR A[Text conditioning] --> B[Stream encoder layers] B --> C[DiT denoising] C --> D[Optional tile-first VSA] D --> E[Original or sm89 fine attention] E --> F[On-demand VAE decode] F --> G[Clip and benchmark measurements]Reviews (1) · Last reviewed commit: "[docs]: record Track B PR and strict 8 G..."