I’m getting much better results with Minimax, and little quicker, since I moved to the latest 1.1 version of the most popular 4-step turbo LoRA.
I’m now using 32Gb of files in total on a 12Gb VRAM card, an incredible feat made possible by the behind-the-scenes magic of Minimax, ComfyUI and the .W4A8 format (successor to GGUFs)…
Model: minimax_h3_ref2va_hybrid_b20-49_pruned_w4a8_mixed.safetensors (11.6Gb)
Text encoder: qwen3vl_32b_minimax_h3-w4a8_convrot.safetensors (14.6Gb)
Video VAE: minimax_h3_video_vae_int8_convrot.safetensors (2.9Gb) (still the good old fast one)
Audio VAE: still the official one (577Mb)
Turbo LoRA: minimax_h3_fl2v_turbo_4step_v1.1_768p_comfyui_bf16.safetensors (1.8Gb)
Generating in 4 steps, er_sde with sgm_uniform for visual quality.
The audio is bad, due to the turbo 4-steps and the er-sde sampler, but audio would be replaced manually. Tests initially suggest the turbo LoRA should be left at 1.0, regardless of video size.
Tests, 10 seconds:
0.3 megapixels.
0.5 megapixels.
1.0 megapixels.
As you can see, much of the graininess on the complex fast-moving net and dragonfly wings goes away at the native 1.0. Not completely, but apparently such problems are vanquished as one nears 2.0.
I’m still testing and building the text-to-video workflow, and have not yet started using reference images which would fix the owl’s visual character etc. Also, the next version of ComfyUI Portable will add the ability to include keyframes anywhere, which will make generation less of a slot-machine and more of an assembly tool.







































































