SageAttention is not new enough version — MiniMax H3
SageAttention is not new enough version or could not determine CUDA architecture, cannot apply MiniMax H3 memory efficient sage attention patch. The workflow can still run; this page covers H3 setup, models, and the fallback.
MiniMax H3 is an open-weights video model available through ComfyUI's built-in templates and nodes. For the official templates, you do not need a separate wrapper custom node.
If you searched sageattention is not new enough version or could not determine cuda architecture, cannot apply minimax h3 memory efficient sage attention patch, that is a performance warning, not a missing-module failure. The workflow can still run. For No module named 'sageattention', use the SageAttention missing-module page.
SageAttention is not new enough version
That log means the MiniMax H3 memory-efficient SageAttention patch did not apply. ComfyUI usually continues with standard attention. Update SageAttention only if you want the patch; switch the node to sdpa or comfy if you just want a clean queue. Setup, model folders, and the rest of H3 live on this page.
H3 is an omni-modal video model: feed it text, images, video, or audio, and it generates video with real stereo sound in the same pass (not audio bolted on afterward), up to 2K resolution and 15 seconds per clip. With official int8/FP8 quantizations and dynamic VRAM offloading, it runs locally on a GPU like an RTX 3060.
Fast answer
Update ComfyUI to a current release, download the workflow template from ComfyUI's template library (T2V, I2V, or R2V), follow the note inside the workflow to download the model files into the right folders, and queue. If you see SageAttention is not new enough version ... cannot apply minimax h3 memory efficient sage attention patch, that is a performance warning, not a hard failure — the workflow still runs.
What MiniMax H3 can do
| Mode | What you provide | What you get |
|---|---|---|
| Text-to-video (T2V) | A prompt only | Video generated from the prompt |
| Image-to-video (I2V) | One image + prompt | The image comes to life |
| First/last-frame | Start frame, end frame, or both | The model fills in the motion between them |
| Reference-to-video (R2V) | Reference images, video, and/or audio | A subject, motion, or voice carried through the clip |
| Native audio | Any of the above | Stereo audio generated with the video in the same pass |
H3 also understands multimodal context: you can describe the relationship between your input images, audio, and video in the prompt, and the model resolves the cross-modal work itself. Motion transfer is especially useful for iteration — a reference video can supply the camera move or performance while the subject and style come from elsewhere.
System requirements
The official ComfyUI blog reports the model works locally on a RTX 3060 when you use the smallest variants with dynamic VRAM offloading:
| Setup | VRAM | Notes |
|---|---|---|
| Full precision (BF16) | ~24 GB+ | Best quality; ~123.6 GB unquantized total before offloading |
| Int8 convrot (official) | 12-16 GB | Official int8 quantization, ~42.5 GB total for smallest variants |
| FP8 / pruned variants | 12-16 GB | Good quality with lower VRAM |
| NVFP4 / INT4 community quants | 8-12 GB | Community quants (rockerBOO, Abiray) for consumer GPUs |
Official int8 convrot quantization plus custom kernels reduce the total memory footprint by about 66% (from 123.6 GB full precision to 42.5 GB with the smallest variants), and dynamic VRAM offloading is what lets a 2K video model run on a 3060. Expect slower sampling with aggressive offloading.
Step 1: Update ComfyUI
MiniMax H3 support is built into current ComfyUI releases. Update before loading a template so the workflow's built-in nodes are available.
For a Git install:
cd ComfyUI
git pull
python -m pip install -r requirements.txtFor the Windows portable package, run the update from the portable root with the bundled Python. If you do not see the MiniMax H3 nodes after updating, restart ComfyUI completely and check that the frontend package updated too (see comfyui-frontend-package Not Installed or Version Outdated).
Step 2: Download the workflow template
The easiest path is the built-in template library:
- Open ComfyUI in the browser.
- Open the Workflow → Browse Templates (or template library) menu.
- Search for MiniMax H3.
- Load the T2V, I2V, or R2V template.
Each template contains a note that lists the exact model files to download and where to save them. The official weights live at Comfy-Org/MiniMax-H3 on Hugging Face.
Step 3: Download and place the models
The official repository contains these files. Download the ones your template asks for and place them in the matching folders.
Diffusion models — ComfyUI/models/diffusion_models/
| File | Notes |
|---|---|
minimax_h3_fl2va_bf16.safetensors | FL2VA (image-to-video) full precision |
minimax_h3_fl2va_int8_convrot.safetensors | FL2VA int8, good balance for 12-16 GB |
minimax_h3_fl2va_pruned_bf16.safetensors | FL2VA pruned BF16 |
minimax_h3_fl2va_pruned_fp8_scaled.safetensors | FL2VA pruned FP8 |
minimax_h3_fl2va_pruned_int8_convrot.safetensors | FL2VA pruned int8 |
minimax_h3_ref2va_bf16.safetensors | Ref2VA (reference-to-video) full precision |
minimax_h3_ref2va_int8_convrot.safetensors | Ref2VA int8 |
minimax_h3_ref2va_pruned_bf16.safetensors | Ref2VA pruned BF16 |
minimax_h3_ref2va_pruned_fp8_scaled.safetensors | Ref2VA pruned FP8 |
minimax_h3_ref2va_pruned_int8_convrot.safetensors | Ref2VA pruned int8 |
Text encoders — ComfyUI/models/text_encoders/
| File | Notes |
|---|---|
qwen3vl_32b_minimax_h3_bf16.safetensors | Full precision Qwen3-VL 32B encoder (large) |
qwen3vl_32b_minimax_h3_int8_convrot.safetensors | Int8 encoder, recommended for local setups |
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors | NVFP4-AWQ encoder, lowest VRAM |
VAE — ComfyUI/models/vae/
| File | Notes |
|---|---|
minimax_h3_audio_vae_fp32.safetensors | Audio VAE (stereo audio) |
minimax_h3_video_vae_fp16.safetensors | Video VAE |
LoRA (optional, community) — ComfyUI/models/loras/
The official T2V template can use Kijai's LightX2V turbo LoRAs from Kijai/MiniMax-H3_comfy on Hugging Face for faster sampling:
| File | Notes |
|---|---|
minimax_h3_fl2v_lightx2v_turbo_4step_v0.1_comfy.safetensors | 4-step turbo LoRA (v0.1) |
minimax_h3_fl2v_lightx2v_turbo_4step_v1.0_768p_resized_avg_rank_31_bf16.safetensors | 4-step 768p v1.0 |
minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_resized_avg_rank_24_bf16.safetensors | 8-step v1.0 |
minimax_h3_ref2v_lightx2v_turbo_4step_v0.1_resized_avg_rank_20_bf16.safetensors | Ref2V 4-step |
Folder mistakes to avoid
Do not put the text encoder in diffusion_models or the diffusion model in checkpoints. H3 files are named after the model family, not the loader. If a dropdown is empty after moving files, restart ComfyUI or refresh the model list, and confirm you edited the same ComfyUI install that is running. See Where to Put Safetensors in ComfyUI for the general folder map.
Community quantizations for lower VRAM
If you are on 8-12 GB VRAM, community quantized variants are widely used:
- rockerBOO/minimax-h3-nvfp4-convrot — NVFP4 and INT4/INT8 mixed variants, including
minimax_h3_fl2va_pruned_nvfp4.safetensorsandminimax_h3_ref2va_pruned_nvfp4.safetensors - Abiray/MiniMax-H3-nvfp4-INT4-INT8-Convrot — INT4/INT8/NVFP4 mixed quantizations plus quantized text encoders (
qwen3vl_32b_minimax_h3_int4_convrot.safetensors) and the VAE files
These are third-party conversions: verify the repo and checksums before use, and expect quality/behavior differences from the official weights.
Step 4: Run the workflow
- Load the template (T2V, I2V, or R2V) and confirm no nodes are red.
- Check the loader nodes point at the files you downloaded.
- For I2V, connect your starting image. For R2V, connect your reference image/video/audio inputs.
- Write a detailed prompt. H3 resolves multimodal references against the prompt, so describing how the inputs relate (for example, "use
<Picture 1>as the subject and<Audio 1>exactly as it is") gives better results. - Queue and wait. With native audio enabled, output is a video with stereo sound.
Common errors
| Error | What it means | Fix |
|---|---|---|
SageAttention is not new enough version or could not determine cuda architecture, cannot apply minimax h3 memory efficient sage attention patch | SageAttention is missing, too old, or CUDA arch was not detected; ComfyUI falls back to standard attention | Usually a warning — the workflow still runs. To silence it, update SageAttention or disable the patch and use sdpa/comfy attention. See No module named 'sageattention' |
No module named 'sageattention' | Optional acceleration backend missing | Not fatal; switch the node to sdpa, comfy, torch, or auto attention. See SageAttention missing |
No module named 'triton' | Triton layer missing or wrong for the installed Torch/CUDA | Fix Triton before SageAttention. See Triton missing or unavailable |
Out of memory / Not enough GPU memory | VRAM exhausted | Use the int8/pruned variants, lower resolution or frame count, close GPU-heavy apps, or enable --lowvram. See CUDA Out of Memory |
| Empty model dropdown | File in the wrong folder, or ComfyUI not restarted | Check the folder map above, restart, and confirm the running install path |
| Missing nodes in a downloaded workflow | Workflow built with newer ComfyUI or needs a custom node | Update ComfyUI; check Workflow Missing Nodes |
FAQ
Does MiniMax H3 run on an RTX 3060?
Yes. The official ComfyUI blog reports local 2K video generation on a 3060 using the smallest official variants with dynamic VRAM offloading. Expect slower sampling than on higher-VRAM cards.
Do I need a custom node pack for MiniMax H3 in ComfyUI?
No. ComfyUI 0.30.0+ has built-in MiniMax H3 nodes. Third-party packs exist (for example API-based wrappers or unified workflow packs) but the official templates use built-in nodes.
Where do MiniMax H3 models go in ComfyUI?
Diffusion models in models/diffusion_models/, text encoders in models/text_encoders/, VAE files in models/vae/, and LoRAs in models/loras/ — all under the active ComfyUI/models/ folder.
Is the SageAttention patch error fatal?
No. sageattention is not new enough version or could not determine cuda architecture, cannot apply minimax h3 memory efficient sage attention patch is a fallback warning: ComfyUI continues with standard attention. Update SageAttention or disable the patch if you want it gone.
Does H3 generate audio?
Yes — natively. H3 generates stereo audio in the same pass as the video; it is not a post-processing step. The audio VAE (minimax_h3_audio_vae_fp32.safetensors) must be placed in models/vae/ for audio workflows.
Related Guides
- Where to Put Safetensors in ComfyUI — folder map for model files
- SageAttention is not new enough version — LTX-2 — the same warning family on LTX-2
- No module named 'sageattention' in ComfyUI — SageAttention warnings and the H3 patch
- Triton missing or unavailable — Triton errors on Windows
- CUDA Out of Memory — VRAM fixes
- ComfyUI Failed to Get Custom Node List — if template downloads fail
- ComfyUI Multi GPU — running on GPU 1 or two instances
Source References
Start free with Wonderful Launcher if this affects your real ComfyUI environment. It keeps launcher-native repair, task logs, and runtime checks in one place; credits are only for image generation and metered tools.
Download Wonderful LauncherSee credit packagesDid this fix your issue?
Your answer helps prioritize verified ComfyUI repairs.