MiniMax H3 prompts work best as timelines. Define the opening frame, visible action, camera path, sound, and ending state in that order. For image or reference workflows, describe what each input controls instead of repeating details the model can already see.
MiniMax H3 prompt guide: quick answer
- Prompt structure: opening frame, action, environmental response, camera, audio, and final frame.
- Text-to-video: define the whole scene because no source image establishes it.
- Image-to-video: preserve the fixed details and describe only what changes after frame one.
- Reference-to-video: assign one clear job to each image, video, or audio file.
- Audio: separate dialogue, physical sound, ambience, and background music.
- C Dance AI settings: 4 to 15 seconds, with 768P and 2K output.
The examples below are writing templates rather than controlled benchmark results. The community section identifies its test material separately.
What is MiniMax H3?
MiniMax H3, also known as Hailuo 3.0, is an AI video model that generates video and native stereo audio from text and visual references. This guide applies the official prompt structure to MiniMax H3 in C Dance AI.
| Property | Current article scope |
|---|---|
| Main workflows | Text-to-video, image-to-video, first/last-frame, and reference-to-video |
| Prompt length in C Dance AI | 1 to 7,000 characters |
| Duration in C Dance AI | Whole numbers from 4 to 15 seconds |
| Resolution in C Dance AI | 768P or 2K |
| Native audio | Dialogue, ambience, physical sound, and background music |
| Reference contract | Up to 9 images, 3 videos, 3 audio files, and 12 files total |
Which MiniMax H3 mode should you use?
| Mode | What the input already defines | What the prompt must define | Common mistake |
|---|---|---|---|
| Text-to-video | Nothing visual | Opening composition, subject, action, camera, sound, and ending | Starting with motion before establishing the scene |
| Image-to-video | Appearance and first-frame composition | Preserved details and movement after frame one | Rewriting the image instead of describing change |
| First/last-frame | Opening and ending states | The physical transition between both frames | Describing two still images with no connecting action |
| Reference-to-video | Assigned identity, motion, style, or timing cues | One role per asset and the intended transformation | Adding references that conflict with one another |
How should a MiniMax H3 prompt be structured?
A useful first draft answers six questions in order:
- What is in the opening frame?
- What does the subject do?
- How does the environment react?
- Where does the camera move?
- What can be heard?
- What does the final frame show?
For a single-shot product clip, that might become:
Live-action studio product video. A cold citrus drink can stands upright on dark stone. A narrow stream of water strikes the surface beside the can, and the splash curls around its base while condensation moves down the metal. The camera makes one slow clockwise arc, keeping the label facing the lens. Crisp rim light, restrained reflections. The shot ends with the water settled and the can centered.
The prompt describes cause and effect. Water hits the surface, the splash moves around the can, and the droplets settle. "Realistic liquid physics" is less useful because it names a desired quality without defining the event.
Open the C Dance AI workspace with MiniMax H3 selected when you are ready to test a prompt.
How does the prompt change by H3 mode?
H3 supports several workflows, and each one needs different information. Repeating the same prompt across every mode wastes the control supplied by the input assets.
Text-to-video: define the whole shot
Text-to-video has no opening image to establish the subject or composition. Describe the initial frame before the action begins.
Live-action, medium-wide shot inside a quiet bakery before sunrise. A baker opens the wooden shutters, places a warm loaf on the counter, then cuts one slice as steam rises. The camera slowly pushes toward the bread without changing height. Street ambience stays faint behind the scrape of the shutters and the crisp cut through the crust. Soft acoustic guitar plays under the scene. End on a close view of the sliced loaf.
Keep the opening concrete: framing, subject position, setting, and light. A prompt that begins with movement but never establishes the scene forces the model to decide too much at once.
Image-to-video: describe what changes after frame one
The uploaded image already supplies appearance and composition. Do not spend half the prompt repeating every visible detail. Identify what must stay fixed, then describe the next movement.
Use the uploaded image as the exact opening frame. Preserve the woman's face, blue cardigan, silver necklace, seat position, and the train carriage layout. She lifts her gaze from the folded letter toward the rain-covered window, then folds the paper along its existing crease. The camera moves slowly to the right while her reflection crosses the glass. Rain taps the window above the low rhythm of the train wheels. No background music.
If you provide both a first and last frame, describe the physical path between them. Avoid writing two static image descriptions. Explain how the pose, objects, camera, and light move from the opening state to the ending state.
Reference-to-video: give every asset one job
The H3 provider contract wired into C Dance AI supports up to nine images, three videos, and three audio files, with 12 files total. The public workspace currently exposes text-to-video and image-to-video; full-reference controls depend on the interface available when you generate. More references do not guarantee better continuity. Each file should resolve a specific uncertainty.
Use reference roles such as:
- Image 1 defines the person's face and hair.
- Image 2 defines the outfit and color palette.
- Video 1 supplies the body movement and camera rhythm.
- Audio 1 supplies the voice or timing reference.
Then write the transformation directly:
Use the person from Image 1 as the only performer. Preserve her facial proportions, short black hair, and red jacket. Follow the body movement and shot timing from Video 1, but place the action in the subway platform shown in Image 2. Keep the camera path and performer pose sequence from Video 1. Use Audio 1 for timing; do not add background music.
Conflicting references are a common source of drift. If two images show different outfits or face shapes, H3 must decide which details to keep. Remove the weaker reference before adding more prompt text.
What do community H3 tests show?
Official prompt rules explain the syntax. Community tests show which variables deserve an early check.
One public benchmark assembled 1,000 H3 generations across many subjects and styles. Its sample grid covers faces, typography, crowds, physical contact, animation, camera movement, and ordinary scenes. That range provides more evidence than a single attractive clip, while still requiring individual outputs to be inspected for failures.

Broad variety can hide individual failures. For a focused evaluation, keep the input, duration, mode, resolution, and one changed variable beside each attempt.
A separate community reference-to-video test combined character, environment, and object references in one moving scene. Two images define the performers, one defines the mountain setting, and one defines the bicycles. Those roles remove four visual ambiguities from a generic prompt for "two anime characters riding bicycles."
Reddit user irmemon225 created this reference-workflow sample. It documents one output from one reference set; it does not establish an average success rate.
A second community test used low resolution and a low step count to refine prompts before final rendering. The draft pass checks subject count, camera direction, reference roles, and shot order. Increase quality after those elements are stable.
How should camera motion be written?
Choose one dominant camera path for a short shot. H3 understands moves such as push in, pull out, pan, truck, tilt, pedestal, arc, tracking, static framing, shake, point of view, zoom, and roll.
Write the movement inside the scene:
The camera tracks beside the cyclist at street level while maintaining the same distance.
Avoid a stack of labels such as:
dynamic camera, drone shot, orbit, whip pan, zoom, rack focus
Those instructions compete for the same seconds. If the shot needs a reveal, state the starting view, the movement, and what the camera reveals:
A close shot begins on flour covering the baker's hands. The camera pulls back slowly, revealing the full worktable and six finished loaves while the baker continues kneading.
Use a cut only when the next shot contributes new information. A slight change in distance is usually better handled by camera motion.
How do you write a multi-shot H3 prompt?
A 6-second clip cannot carry the same story as a 15-second clip. Assign time to each beat before adding detail.
This 10-second product ad uses three shots:
[Shot 1] Live-action commercial, close view of a runner tying one shoe on a wet track before dawn. The camera stays low and static. Fabric pulls tight and the lace snaps softly.
[Shot 2] At 00:03.000, cut to a side tracking shot as the runner accelerates through the first bend. Each footfall throws a small spray from the track. Breathing and foot impacts rise above the distant city ambience.
[Shot 3] At 00:07.000, cut to the shoe landing beside the finish line, then push in slowly as the runner stops. A short bass pulse ends with the final footstep.
The cut times leave three seconds for the setup, four for the main action, and three for the finish. If dialogue is added, it must fit inside the same budget.
Do not add timestamps to every movement in a single continuous shot. Use them to mark distinct cuts or essential timed events.
How do you prompt dialogue and native audio?
H3 generates video and native stereo sound together. A prompt becomes easier to control when it distinguishes three audio layers:
- Dialogue: exact words spoken by an identified person.
- Soundscape: ambience, footsteps, impacts, fabric, rain, breathing, or machinery.
- Background music: music heard by the audience but not by people in the scene.
For two speakers, keep their identities stable:
[Shot 1] Live-action, medium shot in a small train compartment. The woman by the window, speaking softly (Speaker 1), says in English: "I get off at the next station." The conductor at the door, with a low voice (Speaker 2), replies in English: "We arrive in two minutes." Their mouths move only during their own lines.
Soundscape: Steady wheel noise, light rain against the glass, and one soft ticket punch.
Background music: Sparse piano notes at a slow tempo, fading before the conductor replies.
Keep spoken copy short. Read it aloud with a timer. If a sentence takes five seconds to say, it cannot fit naturally beside another line and several actions in a 6-second video.
For voiceover, say that it is off-screen and that the on-screen subject's lips remain closed. This reduces the chance that H3 assigns narration to the visible character.
Which MiniMax H3 prompt examples can you adapt?
These prompts cover different production problems. Replace the subject and brand details, but preserve the action logic.
A vertical creator demonstration
Vertical live-action phone video in a bright apartment. A woman faces the lens while holding a compact handheld vacuum in her right hand. She points to the nozzle with her left index finger, bends toward the sofa, and removes one line of crumbs in a single pass. The camera has light handheld movement but keeps her hands and the product in frame. Natural window light. The motor starts when the nozzle reaches the sofa, then stops at the end. No cuts and no background music.
A first-frame fashion animation
Use the uploaded portrait as the exact first frame. Preserve the model's face, braided hair, green coat, and the position of the street behind her. She turns her shoulders toward the road, takes two measured steps, and looks back over her left shoulder. The camera tracks backward at walking speed and keeps a full-body composition. Wind moves the coat hem in the same direction as her turn. End with her face visible in three-quarter profile.
A physical action scene
Live-action full-body shot in an empty warehouse. A stunt performer runs toward a low wooden barrier, plants both feet, clears it, lands with bent knees, rolls over the right shoulder, and rises into a run. Dust lifts on impact and settles behind him. A stabilized handheld camera follows from three meters back without overtaking him. Hard side light through tall windows. One continuous shot with footsteps, clothing movement, the landing impact, and no music.
A reference-led character replacement
Use Video 1 as the source for shot timing, camera angles, body movement, and background action. Replace only the performer in Video 1 with the person defined by Image 1. Preserve the face, hair, clothing, and body proportions from Image 1 throughout the clip. Keep every other person, object, camera move, and environmental sound from Video 1 unchanged. Do not copy the background from Image 1.
The final example needs Reference mode. A normal image-to-video prompt does not have a source video's timing and motion to preserve.
Why does MiniMax H3 ignore prompt instructions?
When an output fails, change the smallest instruction that explains the failure.
| Failure | Likely prompt problem | First revision |
|---|---|---|
| Random cuts | Several camera moves or scene changes compete | Request one continuous shot and one camera path |
| Wrong person speaks | Speakers are not identified consistently | Assign each speaker a stable label and shorten the exchange |
| Dialogue becomes rushed | The line is too long for the duration | Time the line aloud or increase the clip length |
| Product changes shape | Action, camera, and styling overload the shot | Lock the product orientation and remove secondary motion |
| Reference identity drifts | References conflict or have unclear roles | Keep the clearest identity image and state each remaining role |
| First and last frames do not connect | The physical transition is missing | Describe the intermediate pose, object, and camera changes |
| Sound feels unrelated | Audio is expressed only as a mood | Name the source, timing, and change in volume |
Do not rewrite the entire prompt after each failure. If the camera was wrong but the subject stayed consistent, keep the subject language and revise the camera sentence. This makes each retry informative.
Which duration and resolution should you choose?
The current C Dance AI integration accepts whole-number durations from 4 through 15 seconds. It offers 768P and 2K output, and the default is a 6-second 2K generation.
Current base credit calculation:
| Setting | Credits |
|---|---|
| 768P | 50 per output second |
| 2K | 80 per output second |
| Each reference image after the first five | 25 additional credits |
A 6-second text-to-video generation therefore starts at 300 credits in 768P or 480 credits in 2K. Under the reference contract, reference-video duration and extra reference images can increase the total. The workspace shows the calculation before submission.
Draft at the shortest duration that can hold the action and dialogue. Use 768P for early composition tests when final detail is not yet the question. Move to 2K after the prompt's timing and camera behavior are stable.
What should you check before generating?
Read the prompt once and confirm:
- The opening frame is clear.
- The clip has one main action.
- Cause and visible effect are both described.
- Camera instructions do not conflict.
- Cuts fit inside the selected duration.
- Each speaker and reference keeps one identity.
- Dialogue can be spoken in time.
- Ambience and music have separate jobs.
- The ending state is visible and achievable.
Then return to C Dance AI or generate with MiniMax H3 already selected. Save the prompt and settings with each attempt, record the failure in plain language, and adjust one variable before retrying.
MiniMax H3 prompt FAQ
What should a MiniMax H3 prompt include?
Start with the subject and visible action, then define the setting, camera path, shot order, dialogue, physical sound, and background music that the clip needs. Keep each instruction compatible with the selected duration.
How long can a MiniMax H3 video be in C Dance AI?
The current C Dance AI integration accepts whole-number durations from 4 through 15 seconds and offers 768P and 2K output.
Does MiniMax H3 generate audio?
Yes. MiniMax H3 can generate native stereo audio with the video. Prompts can direct dialogue, ambience, physical sound effects, and background music.
Can MiniMax H3 use image, video, and audio references?
The H3 provider contract supports image, video, and audio references. C Dance AI currently exposes text-to-video and image-to-video in the public workspace; full-reference controls depend on the interface available when you generate.
How should I prompt dialogue in MiniMax H3?
Identify each speaker consistently, write the exact spoken line, specify the language, and leave enough time for the line. Separate dialogue from ambience and background music so their roles remain clear.
Why does MiniMax H3 ignore parts of a prompt?
The prompt may contain too many actions, conflicting camera moves, unclear reference roles, or more dialogue than the clip can fit. Remove secondary instructions and test one change at a time.

