Use Text, Images, Video, and Audio Together
Build a richer creative brief with more than a prompt. Add supported visual and audio references so H3 can use subjects, motion, scenes, sound, and other context when creating your video.

Create AI videos from text, images, video, and audio references with MiniMax H3. Guide motion, scenes, and sound in natural language.
MiniMax H3 is a general-purpose multimodal generation model from MiniMax. It can understand context from text, images, video, and audio, then use those references to generate or edit audiovisual content.
H3 supports workflows such as text to video, image to video, reference-based generation, video editing, and V2V motion transfer. It can generate video with native stereo audio and supports output up to 15 seconds at up to 2K resolution.
For creators, marketers, e-commerce teams, and designers, this means you can describe a complete creative brief and use several types of reference media in the same workflow.
Build a richer creative brief with more than a prompt. Add supported visual and audio references so H3 can use subjects, motion, scenes, sound, and other context when creating your video.
Tell H3 what each input should do. Ask it to take camera movement from a reference video, a subject from an image, and sound direction from audio instead of treating every creative task as a separate workflow.
MiniMax H3 jointly generates audiovisual content and supports native stereo audio. This makes it useful for scenes that need voice, sound effects, music, or environmental sound as part of the generated result.
MiniMax states that H3 can generate videos up to 15 seconds at up to 2K resolution. Higher-resolution output can be useful for product visuals, branded scenes, interface concepts, dynamic posters, and other detail-sensitive content.
Use video as a motion reference for V2V motion transfer and other video-to-video workflows. Describe which movement, camera behavior, or action you want H3 to carry into your new scene.
Create short-form AI video for TikTok, Instagram Reels, YouTube Shorts, music visuals, personal projects, and storytelling concepts. Use text to video when you want to start from an idea. Add image, video, and audio references when you need more control over characters, motion, camera direction, or sound.
Build ad concepts, campaign hooks, product reveals, dynamic posters, social promos, and branded video drafts from a richer set of creative references. H3 can combine visual, motion, and audio context, making it useful when your campaign brief contains more than a single text prompt.
Turn product images and creative references into product videos, launch concepts, Shopify assets, promotional clips, and branded social content. Use image to video for simple motion, or add video and audio references when you want to guide camera behavior, pacing, sound, or overall presentation.
Explore product design presentations, UI/UX motion, game interfaces, concept videos, and visual communication for more complex workflows. H3's multimodal context lets teams describe relationships between reference assets in natural language before generating an audiovisual draft.
Use MiniMax H3 in your browser through the PWZ generator. No local GPU, model weights, or ComfyUI setup is required.
Select the MiniMax H3 generation workflow available on PWZ.
Describe the scene you want. Upload supported image, video, or audio references that should guide the subject, camera, movement, style, or sound.
Write the relationship directly in your prompt. For example, specify which image defines the subject, which video guides movement, and which audio provides the sound direction.
Choose the available output settings, then generate your video. Review the result and adjust your prompt, references, or settings when you want another version.
Review the required PWZ credit cost before generation so you can decide whether to continue.
These are MiniMax H3 model capabilities. Features available through PWZ depend on the API or inference workflow implemented on this page.
Start free and buy PWZ credit packs when you need more MiniMax H3 generations. Credits never expire and include watermark-free downloads and commercial usage rights.
600 credits · $0.062 per credit
Perfect for trying out PWZ — get started with a credit balance for your first generations.
2,000 credits · $0.052 per credit
For creators generating regularly — better per-credit value for daily PWZ workflows.
4,000 credits · $0.045 per credit
Best value for teams and content pipelines — more generations at the lowest per-credit cost.
MiniMax H3 is a general-purpose multimodal generation model from MiniMax. It understands text, images, video, and audio context and can generate audiovisual content with native stereo audio.
Yes. PWZ provides third-party browser-based access to supported MiniMax H3 video generation workflows without requiring you to deploy the model locally.
Yes. H3 supports text-to-video generation. You can describe the scene, motion, camera behavior, and other creative direction in your prompt.
Yes. H3 supports multimodal reference workflows using images, video, and audio. The exact number, size, and file formats accepted by PWZ depend on the implemented generation mode.
Yes. H3 supports native stereo audio as part of its audiovisual generation design rather than requiring every workflow to add audio only after video generation.
MiniMax states that H3 supports video up to 2K resolution and up to 15 seconds. Resolution and duration options available through PWZ may vary by generation mode.
H3 is designed around multimodal context and task generalization. It can use relationships between text, images, video, and audio instead of relying only on a single text prompt.
Bring your prompt, images, video, and audio references into one online workflow. Describe how they should work together, then generate your next AI video with MiniMax H3 on PWZ.
Try MiniMax H3 Free