Native MiniMax-H3 inference for Apple Silicon. The project is being built as a sequence of working vertical slices: deterministic host/model metadata first, then portable Metal block parity, prompt encoding, prompt-to-video/audio, and first/last-frame conditioning and then ordered references.
Prompt-to-video/audio, first/last-frame conditioning, and ordered Ref2VA image/video/audio references work end to end. The current work is incremental H3-specific Metal performance and memory optimization on M3 Max and M5 Max.
Tutorial
1. Build and inspect the model
The examples assume that the Hugging Face snapshot is in ./MiniMax-H3 and that FFmpeg and FFprobe are available on PATH .
make -j8 mkdir -p outputs ./h3 --info -d ./MiniMax-H3
--info checks the model layout and prints the selected Metal device without mapping all weights or generating media. Run ./h3 --help for the complete CLI reference.
Without -p , the same binary starts an Iris-style interactive session:
./h3 -d ./MiniMax-H3 --width 512 --height 512 --steps 6
Type a prompt to generate a numbered video. The session keeps the exact BF16 prompt conditioning, prepared DiT, and video decoder in memory, so repeating a prompt with another seed avoids loading and encoding them again. Useful commands are !status , !seed random , !seconds 2 , !show , !save output.mp4 , and !cache . Use !help for the full, short list.
... continue reading