MiniMax-H3 Turbo LoRA β€” few-step audio-video generation

A LoRA for MiniMax-H3 that renders joint video + synchronized stereo audio in as few as 4 sampling steps instead of the usual ~20 β€” a ~5Γ— sampling speedup β€” and keeps getting better as you add steps.

Which checkpoint β€” v4 (step 600) or v1 (850)?

For most work, use minimax_h3_turbo_v4_step600_ema.safetensors. It's the strongest checkpoint we've released: much better static and small-motion shots, markedly better micro-detail (faces, fingers, fine texture), and the over-sharpening / plastic look of the earlier v1 (~850) line is fully resolved.

v4 introduced a static-frame enhancement β€” a big win for static and small-motion content. The one trade-off shows up only at 4 steps with large, fast motion, where v4 can produce motion-smear / trailing ghosting (we're actively fixing this). Two things address it:

  • Use 6–8 steps. This largely removes the smear and is where v4 looks its best. v4 also tolerates higher step counts better than v1, which tends to over-sharpen at high steps + strength 1.0.
  • For the specific case of 4 steps and heavy motion, the older v1 ~850 checkpoint can still be the friendlier pick.
Using 6–8 steps?        ── yes ──►  v4-600  (recommended)
   β”‚ no (4 steps)
   β–Ό
Heavy / fast motion?    ── no  ──►  v4-600  (recommended)
   β”‚ yes
   β–Ό
                                    v1-850  (friendlier at 4-step heavy motion)

Still a preview β€” training continues; the two areas still being improved are audio and behaviour under fast, intense motion.

Steps and strength β€” read this

  • 4 steps is the recommended minimum; 4–8 is the useful range. 6–8 steps look noticeably better than 4, so add steps if you can afford them. Past 8 steps it stops helping and can start to introduce over-sharp artifacts β€” there's no benefit to going higher, so stay in 4–8.
  • Keep strength at 1.0. It's tuned for 1.0 and holds up well across the 4–8 step range. Only reach for the strength dial if a specific clip misbehaves β€” then blurry ghosting / smear β†’ nudge up (~1.05–1.2), over-sharp grain β†’ nudge down (~0.8–0.95).
  • Keep the scheduler on simple.

Use it in ComfyUI (recommended)

Custom nodes: Larryvrh/ComfyUI-MiniMax-H3-Turbo β€” or search "MiniMax-H3 Turbo" in ComfyUI-Manager. (Keep the node updated; it evolves alongside these weights.)

  1. Install the nodes (Manager, or git clone into ComfyUI/custom_nodes) and put a .safetensors from this repo into ComfyUI/models/loras/. You also need the base MiniMax-H3 model, VAEs and text encoder β€” see the MiniMax-H3 tutorial.
  2. Start from the official MiniMax-H3 workflow (t2v or i2v) and make two changes:
    • insert MiniMax-H3 Turbo LoRA between the model loader and the sampler;
    • feed SamplerCustomAdvanced from MiniMax-H3 Turbo Sampler, and set the scheduler to simple at β‰₯ 4 steps.

Everything else stays as in the official graph, so both text-to-video and image-to-video work. A ready-made t2v workflow ships in the node repo (and here as minimax_h3_t2v_turbo.json) β€” drag it in.

  • Base model: any MiniMax-H3 base β€” full (bf16, int8_convrot) and the pruned/curve variants (pruned_int8, pruned_fp8). The node auto-detects a pruned base and re-injects the time-conditioning at run time, so one LoRA file covers every base.
  • low_vram switch: off applies the LoRA at run time (sharpest, recommended); on merges it into the weights for the lowest peak VRAM (a bit softer on quantized bases). Turn it on only if you run out of memory.
  • The custom sampler auto-adapts to your ComfyUI version: MiniMax-H3 runs video and audio on two different flow schedules; recent ComfyUI handles that natively (ModelSamplingAV) and older ComfyUI doesn't β€” the Turbo Sampler detects which and does the right thing either way, so nothing to change when you update ComfyUI.

Weights

All bf16, ~744 MB, applied as a plain low-rank update (W_eff = W + lora_B @ lora_A, alpha = rank, so no extra scaling). Prefer the EMA files; the non-EMA ones are for comparison.

file notes
minimax_h3_turbo_v4_step600_ema.safetensors recommended β€” current best. Strong static/small-motion, good micro-detail, no over-sharpening.
minimax_h3_turbo_v4_step600.safetensors v4-600 non-EMA (comparison).
minimax_h3_turbo_v4_step150_ema.safetensors earlier v4 checkpoint.
minimax_h3_turbo_4step_ema_ckpt850.safetensors v1 line (~850) β€” over-sharpened / plastic in general, but the friendlier pick for 4-step heavy motion (see above).
minimax_h3_turbo_4step_ema_ckpt500.safetensors older v1 (~500), softer.
minimax_h3_turbo_4step_ema.safetensors initial release (~200).

Naming: v4 is the current training recipe and stepN is the training step. Older files carry the previous 4step_ckptN naming, where 4step referred to the sampler-step count.

Standalone (no ComfyUI graph)

generate.py is a single self-contained file β€” it loads the base DiT + a LoRA, encodes the prompt, runs the few-step dual-schedule sampler, decodes and muxes an mp4. It still needs a ComfyUI checkout for the H3 model / VAE / text-encoder definitions:

git clone https://github.com/comfyanonymous/ComfyUI
cd ComfyUI && pip install -r requirements.txt && cd ..
pip install -r requirements.txt          # this repo: torch, safetensors, imageio-ffmpeg

# base weights from Comfy-Org/MiniMax-H3 into a models/ tree, then:
python generate.py \
  --comfyui ./ComfyUI \
  --base   models/diffusion_models/minimax_h3_fl2va_bf16.safetensors \
  --lora   minimax_h3_turbo_v4_step600_ema.safetensors \
  --te     models/text_encoders/qwen3vl_32b_minimax_h3_int8_convrot.safetensors \
  --video-vae models/vae/minimax_h3_video_vae_fp16.safetensors \
  --audio-vae models/vae/minimax_h3_audio_vae_fp32.safetensors \
  --prompt "A corgi in a chef hat flipping a pancake, sizzling sounds and a cheerful bark." \
  --width 1344 --height 768 --frames 124 --steps 6 --out corgi.mp4

Notes

  • Resolution / duration: width and height are multiples of 32 (short edge typically 768). Frame count is at 24 fps and snaps to the model's 17Β·k+5 grid (124 β‰ˆ 5 s). Validated range 124–362 frames (5–15 s).
  • VRAM: the base model is large (~33 B); an 80 GB GPU is comfortable at the largest resolutions. The ComfyUI node streams the base and adds the low_vram switch, so it runs on much smaller GPUs. In the standalone script, --offload-adaln trades ~13 GB of VRAM for CPU RAM.
  • Audio: 32 kHz stereo, aligned to the video; the two streams ride different flow schedules and are integrated each on its own clock. (Audio is one of the two areas still being improved β€” see the top.)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ 1 Ask for provider support

Model tree for larryvrh/MiniMax-H3-Turbo-Lora

Adapter
(7)
this model

Spaces using larryvrh/MiniMax-H3-Turbo-Lora 7