【ComfyUI】MiniMax H3の主要高速化ノードを組み合わせ別で比較してみた【動画生成AI】

ついにローカルにもMiniMax H3というOmni-Reference(オムニリファレンス)に対応した動画生成AIが登場しました。

しかし、高性能がゆえに生成時間も長くなってしまいます。
ですが、そんな生成時間を少しでも短くできる高速化ノードが存在します。

なので今回は、26/8/8時点での主要な高速化ノードを使い生成時間がどれだけ変わるのか。生成される動画の品質にはどれほど影響があるのかを、FL2VAとRef2VAの両方で試してみることにしました。

実行環境

今回の検証で使用した実行環境です。

OS : Windows 11
GPU : NVIDIA GeForce RTX 5090 32GB
CPU : Intel Core Ultra 7 265K
メモリ : 128GB
ComfyUIバージョン : 0.30.1
Pythonバージョン : 3.13.6
PyTorchバージョン : 2.9.1+cu130

使用したモデル

使用したモデルです。

FL2VA
diffusion_models
minimax_h3_fl2va_int8_convrot.safetensors

text_encoders
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors

vae
minimax_h3_video_vae_fp16.safetensors

vae
minimax_h3_audio_vae_fp32.safetensors

Ref2VA
diffusion_models
minimax_h3_ref2va_int8_convrot.safetensors

text_encoders
qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors

vae
minimax_h3_video_vae_fp16.safetensors

vae
minimax_h3_audio_vae_fp32.safetensors

使用する高速化ノード

使用する高速化ノードは以下の3つです。

これらはComfyUI標準のノードではなく、カスタムノードです。
それぞれ以下のカスタムノードで導入することが出来ます。

・MiniMax H3 Mem Eff Sage Attention Patch
ComfyUI-KJNodes

・Patch Sol-Attn
ComfyUI-SolAttn_triton

・Spectrum Apply MiniMax H3
ComfyUI-Spectrum-MiniMax-H3

生成条件

生成条件は以下の通りです。

FL2VAでは576×736 (10秒) 、Ref2VAでは864×480 (10秒)の動画をシード値固定で8つの条件で生成します。
条件は以下です。

Mem Eff SageSol-AttnSpectrum
なしなしなし
ありなしなし
なしありなし
なしなしあり
ありありなし
ありなしあり
なしありあり
ありありあり

パラメーターのある高速化ノードは、全てデフォルトで行います。

FL2VAの開始フレームで指定した画像とプロンプトは以下。

The scene shifts to the city streets.
The town's art style is anime-style.
A video of the character from <Picture 1> dancing in the city.
The only audio is the background music.

[Shot 1] The character from <Picture 1> is standing still with both arms held at their sides. The scene is filmed from behind, with the camera gradually moving closer.

[Shot 1] At 00:03.000, the camera cuts to a frontal view of the character from <Picture 1>.
The character in <Picture 1> performs an acrobatic dance.

[Shot 2] At 00:05.000, switch the camera to an upward angle.

[Shot 3] 00:06.500, Cut the camera to the side of the character in <Picture 1>.

[Shot 4] At 00:08.500, switch the camera to the side view.
The character in <Picture 1> forcefully extends their right arm upward, holds that pose, and finishes the dance.

Ref2VAで指定した画像とプロンプトは以下。

A fight scene.
The location is a multi-story parking garage.

The person in <Picture 1> and the person in <Picture 2> are facing each other, poised for combat.

[Shot 1] At 00:01.000, the two are engaged in a stylish hand-to-hand fight.

[Shot 2] At 00:05.000, the two unleash their finishing moves, triggering intense visual effects.

[Shot 3] At 00:08.000, the two shake hands, acknowledging each other's valiant effort.

スポンサーリンク

生成結果

それでは結果を書いていきます。

音声はミュートにしてあります。音声を聞きたい場合はミュートを解除してください。
その際は音量に注意してください。

FL2VAの結果

高速化ノード : なし

生成時間 : 約3分37秒

高速化ノード : Mem Eff Sage

生成時間 : 約2分50秒

高速化ノード : Sol-Attn

生成時間 : 約2分44秒

高速化ノード : Spectrum

生成時間 : 約2分19秒

高速化ノード : Mem Eff Sage + Sol-Attn

生成時間 : 約2分10秒

高速化ノード : Mem Eff Sage + Spectrum

生成時間 : 約1分30秒

高速化ノード : Sol-Attn + Spectrum

生成時間 : 約1分48秒

高速化ノード : Mem Eff Sage+ Sol-Attn + Spectrum

生成時間 : 約1分30秒

スポンサーリンク

Ref2VAの結果

高速化ノード : なし

生成時間 : 約3分44秒

高速化ノード : Mem Eff Sage

生成時間 : 約2分13秒

高速化ノード : Sol-Attn

生成時間 : 約2分41秒

高速化ノード : Spectrum

生成時間 : 約2分22秒

高速化ノード : Mem Eff Sage + Sol-Attn

生成時間 : 約2分15秒

高速化ノード : Mem Eff Sage + Spectrum

生成時間 : 約1分30秒

高速化ノード : Sol-Attn + Spectrum

生成時間 : 約1分49秒

高速化ノード : Mem Eff Sage+ Sol-Attn + Spectrum

生成時間 : 約1分32秒

最後に

生成速度が向上していますが、プロンプトの効きにも影響が出ていますね。

FL2VAでは途中のカットでキャラクターを側面から映すというプロンプトを入力しているのですが、Spectrum Apply MiniMax H3を適用した場合、側面から映すというプロンプトは反映されていません。

高速化ノードはそのあたりも吟味して使用する必要がありそうです。

それでは!

 

スポンサーリンク

コメント