FastVideo/Wan2.2-Syn-121x704x1280_32k
Viewer • Updated • 33.3k • 18.2k • 7
3-step text-to-video, INT8 pre-quantized for Apple Silicon.
The mid-tier FastMetal model — a DMD2-distilled Wan2.2 TI2V 5B with a quantization-aware-trained INT8 DiT. 720p-native, pre-quantized: no startup quantization.
| Path | Contents |
|---|---|
mlx_dit.safetensors / mlx_dit.json |
INT8 (affine, group-64) DiT |
text_encoder/, vae/, tokenizer/, scheduler/ |
everything needed to run standalone |
Requires macOS with Apple silicon (MPS) and Python 3.11+:
pip install torch transformers mlx safetensors av imageio imageio-ffmpeg
git clone https://github.com/FastVideo/FastVideo.git
cd FastVideo
python examples/inference/basic/mlx_wan22_generate.py \
--text-encoder-root ./FastMetal-5B-QAD \
--mlx-checkpoint ./FastMetal-5B-QAD \
--vae-root ./FastMetal-5B-QAD/vae \
--prompt "a river winding through a fantasy valley at golden hour" \
--fast
| Base model | Wan 2.2 TI2V 5B |
| Distillation | DMD2, 3 denoising steps |
| Quantization | affine INT8, group size 64, QAT-trained |
| Resolution | 704×1280 (720p), 121 frames |
| Flow shift | 5.0 |
| DiT weights | ~5 GB (INT8) |
DMD2 distillation of the Wan 2.2 TI2V 5B teacher onto an INT8 student on
NVIDIA GB200 clusters, with quantization-aware training (affine INT8,
group 64). Training corpus: FastVideo/Wan2.2-Syn-121x704x1280_32k.
| Model | Tier |
|---|---|
| FastMetal-1.3B-QAD | Entry — 16 GB+ class Macs |
| FastMetal-5B-QAD | Mid — 720p |
| FastMetal-14B-QAD | Quality — 24 GB+/ Ideally 36 Macs |
Base model
Wan-AI/Wan2.2-TI2V-5B-Diffusers