MiniMax H3 Video Generation API Integration Guide

This article introduces the integration and usage of the MiniMax H3 Video Generation API. This interface supports text-to-video generation, control of the first and last frames, and multi-modal reference video generation, using a unified V2 multi-modal content structure to create tasks.

Application Process

To use the MiniMax H3 Video Generation API, first go to the YuJun Console to obtain your API Token for future use.

If you are not logged in or registered, you will be automatically redirected to the login page to invite you to register and log in, and will return to the current page automatically after completion.

One API Token can call all services on the platform, no need to apply separately for each service. The first application will grant free credits for a trial experience; when credits are insufficient, you can recharge the general balance in the console.

📘 Complete Documentation: MiniMax H3 Video Generation API →

It is recommended to save the Token as an environment variable and not to write it into the source code or submit it to the version control system:

export ACEDATACLOUD_API_KEY="YOUR_API_KEY"

Interface Overview

  • Base URL: https://api.acedata.cloud
  • Endpoint: POST /minimax/videos
  • Authentication Method: Include authorization: Bearer {token} in the HTTP Header
  • Request Headers:
    • accept: application/json
    • content-type: application/json
  • Model: MiniMax-H3
  • Input Structure: Pass text, images, videos, and audio through the unified content
  • Output Mode: By default, wait synchronously for generation to complete and return the complete task; change to asynchronous by passing async: true or callback_url
  • Result Query: Asynchronous mode or historical tasks can query status and completed videos through the MiniMax H3 Task Query API
  • Asynchronous Callback: Optional, receive the final task result via callback_url

You do not need to pass action to select the generation mode; the interface will automatically determine the purpose based on the material types and role in content.

Suitable Scenarios

Scenario Input Combination Common Uses
Text-to-Video Text Advertising creativity, storyboard rehearsal, short videos, atmospheric shots
First Frame to Video Text + First Frame Image Make product images, posters, character photos, or illustrations come to life naturally
Last Frame / First and Last Frame Video Text + Last Frame, or Text + First Frame + Last Frame Control the opening and closing, transitions, growth changes, and before-and-after comparisons
Multi-modal Reference Video Text + Reference Image / Video / Audio Maintain consistency in characters and products, replicate actions, camera movements, audio tones, or editing rhythms

Calling Process

In the default synchronous mode, after calling /minimax/videos, keep the connection open; when generation is complete, directly obtain the video from the response's task.content.url.

To release the connection immediately, pass async: true or callback_url:

  1. Save the task_id and trace_id from the immediate response.
  2. If no callback is configured, call /minimax/tasks every 10 seconds to check.
  3. When task.status changes to succeeded, obtain the video from task.content.url.
  4. Stop polling when the status is failed or cancelled, and read task.error.

Top-Level Request Parameters

Parameter Type Required Default Value Description
model string Yes - Fixed as MiniMax-H3
content object[] Yes - Multi-modal content array, must contain one non-empty text item
resolution string Yes - 768P or 2K
duration integer Yes - Generation duration, an integer between 4-15 seconds
ratio string Conditionally Required adaptive adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
async boolean No false false returns complete results synchronously; true returns task identifier immediately, then results can be obtained through the task query interface
callback_url string No - Public callback URL to receive the final task result; automatically enables asynchronous mode when provided

The rules for ratio depend on the workflow:

  • Text-to-Video: Required and cannot be adaptive.
  • First Frame, Last Frame, or First and Last Frame Video: The aspect ratio is determined by the input image; it is recommended to omit or pass adaptive.
  • Multi-modal Reference Video: Can be omitted, default is adaptive; a fixed ratio can also be explicitly specified.

The interface does not accept legacy or compatible fields such as prompt, image_urls, audio_urls, messages, and first_frame_image. When receiving such parameter errors, please remove the old fields and migrate to content; for example, change "prompt": "A cat waving" to "content": [{"type": "text", "text": "A cat waving"}]. Do not send both new and old formats simultaneously.

Content Item Parameters

Each content item must have type, and the other fields are determined by the type:

type Data Field role Description
text text Not passed Each request must contain one non-empty text item, up to 7000 characters
image_url image_url.url first_frame First frame image; if there is only one image and role is omitted, it is also treated as the first frame
image_url image_url.url last_frame Last frame image; can be used alone or combined with first_frame to control the start and end
image_url image_url.url reference_image Reference subject, character, product, clothing, scene, or style
video_url video_url.url reference_video Reference action, camera movement, performance, or editing structure
audio_url audio_url.url reference_audio Reference audio tone, dialogue, music, or rhythm

Media addresses support three forms:

  • Publicly accessible HTTPS URL, recommended for large files.
  • mm_file://{file_id}, referencing files that have already been uploaded or have existing results.
  • Base64 data URI corresponding to the media type. Base64 will increase the size by about one-third, please ensure the entire request body does not exceed 64 MB.

Material Specifications and Quantity Limits

Material Format Single File Limit Size / Duration Quantity Limit
Image JPG, JPEG, PNG, WEBP, HEIC, HEIF Up to 30 MB Width and height both between 256-5760 px; aspect ratio 0.4-2.5 Up to 1 first frame, up to 1 last frame, up to 9 reference images
Video MP4, MOV; H.264/AVC or H.265/HEVC; audio track AAC or MP3 Up to 50 MB Each segment 2-15 seconds, total not exceeding 15 seconds; width and height both between 256-5760 px; aspect ratio 0.4-2.5; 23.976-60 fps Up to 3 segments of reference video
Audio WAV, MP3 Up to 15 MB Each segment 2-15 seconds, total not exceeding 15 seconds Up to 3 segments of reference audio

In multimodal reference scenarios, the total number of images, videos, and audios is limited to 12 files. First and last frame scenarios and reference material scenarios are mutually exclusive: once reference_image, reference_video, or reference_audio is used, first_frame or last_frame cannot be used, and vice versa.

Production-Level Capability Demonstration

The following are not concept images or placeholder materials, but real reference inputs and actual video outputs of the MiniMax H3 official production-level capability samples. The three sets of cases cover brand short films, live-action narratives, and fashion e-commerce, suitable for evaluating the model's most critical capabilities in commercial production.

Capability Key Observations
Consistency of Characters and Faces Stability of facial features, hairstyle, makeup, and character temperament after multiple camera switches
Facial Performance Eye contact, micro-expressions, emotional tension, and natural head movement in close-ups
Product Structure Retention Contours, materials, wearing relationships, and mirror reflections of products like glasses and handbags
Brand Visual Execution Consistency of scene atmosphere, film grain, colors, logo, and editing rhythm
Cinematic Narrative Ability to form complete segments through changes in shot distance, character staging, camera movement, rhythm, and sound

Here, "facial capability" refers to the consistency of character appearance, facial details, and performance control in video generation, not identity recognition, face comparison, or face-swapping interfaces.

High-End Brand Short Film: Unity of Characters, Products, and Brand Assets

Production Goal: 16:9 high-end fashion brand film. Establish a stark atmosphere with a desert road and vintage cars, maintaining the appearance of the female lead and the structure of the black handbag, and naturally incorporating the brand logo at the end. This case focuses on testing cross-shot character consistency, product retention, cinematic quality, and brand cohesion.

Atmosphere and Scene Reference Character Reference
Desert road and vintage car brand film atmosphere reference Brand film female lead reference
Handbag Product Reference Brand Logo Reference
Black handbag product reference Brand logo reference

Directly open or download the brand short film

The corresponding content organization is as follows:

{
  "model": "MiniMax-H3",
  "content": [
    {
      "type": "text",
      "text": "15 seconds, 16:9 high-end fashion brand film. A vintage car parked by a desert road, the female lead takes a black handbag from the trunk, briefly makes eye contact with the male lead, and leaves alone. Maintain consistency of characters, handbag, and brand visuals; stark and high-end, cinematic grain, crisp editing, and naturally present the brand logo at the end."
    },
    {
      "type": "image_url",
      "image_url": { "url": "https://cdn.acedata.cloud/uploads/6e65f865-f1c2-4f80-8b51-9a98d4d930b1" },
      "role": "reference_image"
    },
    {
      "type": "image_url",
      "image_url": { "url": "https://cdn.acedata.cloud/uploads/88d89cc3-e6cb-42b4-ab4c-1bbbf6c9f7c8" },
      "role": "reference_image"
    },
    {
      "type": "image_url",
      "image_url": { "url": "https://cdn.acedata.cloud/uploads/e91f7fff-f8e3-4da5-b882-87edbc3c9473" },
      "role": "reference_image"
    },
    {
      "type": "image_url",
      "image_url": { "url": "https://cdn.acedata.cloud/uploads/b68dac43-fb14-42b5-bf8b-fd4d65506520" },
      "role": "reference_image"
    }
  ],
  "resolution": "2K",
  "duration": 15,
  "ratio": "16:9"
}

Live-Action Vertical Short Drama: Face Consistency and Emotional Performance

Production Goal: 15 seconds, 9:16 dark romantic short drama trailer. Lock in character appearance through reference images of the male and female leads, and constrain the space with reference images of a castle; use medium close-ups and facial close-ups to express eye confrontations, fear, restraint, and a sense of danger. This case is suitable for observing the stability of human facial features, micro-expressions, gaze relationships, and continuous performances.

Male and Female Lead References Castle Scene References
Live-action short drama male and female lead references Dark castle scene references

Directly open or download the live-action short drama

The prompt should clearly define the character relationships, emotions, and shot types, rather than just describing "dialogue between a man and a woman":

15 seconds, 9:16 live-action dark romantic short drama trailer. The female lead accidentally enters a forbidden castle, awakening a sleeping vampire noble; he approaches dangerously and restrained, while she is fearful but does not yield. Maintain consistency in the facial features, hairstyles, and costumes of both characters, using medium close-ups and facial close-ups to express eye confrontations and emotional tension, with dark **###** lighting and a tight rhythm.

Fashion Glasses Advertisement: Maintain Facial Details and Product Structure

Production Goal: 9:16 high-end fashion glasses advertisement. Full-body images of characters are responsible for body shape and walking, facial reference images are responsible for facial features and makeup, and product images are responsible for surrounding curves, lens reflections, temples, and cat-eye outlines. This case tests facial close-ups, consistency among multiple people, wearing relationships, and product geometric structures.

Model and Styling References Facial Detail References Glasses Product References
Fashion advertisement model and styling references Model facial detail references Glasses product structure references

Directly open or download the fashion glasses advertisement

In product advertisements, the prompt should clearly separate the responsibilities of character references and product references: character materials constrain the face, makeup, body shape, and temperament; product materials constrain the outline, material, reflection, and wearing position. This is more stable than vaguely writing "generate a glasses advertisement."

Text-to-Video

When there is only one text item, it is a text-to-video. It is suitable for directly generating images from creative, script, or shot descriptions. The prompt can be organized in the order of "subject + action + scene + shot + lighting + sound."

curl -X POST 'https://api.acedata.cloud/minimax/videos' \
  -H "Authorization: Bearer $ACEDATACLOUD_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "MiniMax-H3",
    "content": [
      {
        "type": "text",
        "text": "15 seconds cinematic perfume advertisement: On the black rocks of the coast in the morning, a transparent perfume bottle is surrounded by mist and waves. A close-up showcases water droplets on the bottle and glass refraction, with the camera slowly pulling up from the product close-up to the vast sea; silver-blue tones, real natural light, high-end restraint, ending with a freeze on the product."
      }
    ],
    "resolution": "2K",
    "duration": 15,
    "ratio": "16:9"
  }'

Default synchronous response:

{
  "task": {
    "id": "f5977217-ed2c-40da-adbe-93d08235618f",
    "model": "MiniMax-H3",
    "status": "succeeded",
    "content": {
      "url": "https://cdn.acedata.cloud/minimax/f5977217-ed2c-40da-adbe-93d08235618f.mp4"
    },
    "resolution": "2K",
    "duration": 15,
    "usage": {
      "total_seconds": 15,
      "input_seconds": 0,
      "output_seconds": 15,
      "input_image_count": 0
    },
    "ratio": "16:9",
    "task_type": "generation",
    "modality": "video"
  }
}

If the request includes "async": true, the interface immediately returns:

{
  "task_id": "f5977217-ed2c-40da-adbe-93d08235618f",
  "trace_id": "trace_7f8c2b1a"
}

First Frame Image to Video

Mark the image as first_frame, and the model will start generating from that frame. It is suitable for making posters, product images, character design images, and photographic works come to life naturally.

{
  "model": "MiniMax-H3",
  "content": [
    {
      "type": "text",
      "text": "The character breathes naturally and looks out the window, the corners of the clothes are blown by the breeze, the camera slowly moves in."
    },
    {
      "type": "image_url",
      "image_url": {
        "url": "https://cdn.acedata.cloud/b1c82e4937.png"
      },
      "role": "first_frame"
    }
  ],
  "resolution": "2K",
  "duration": 5,
  "ratio": "adaptive"
}

Last Frame and First and Last Frame Video

Providing only last_frame allows the model to naturally generate to the specified image; at the same time, providing both first_frame and last_frame can clearly control the starting and ending points. Suitable for transitions, shape changes, growth processes, or product comparisons before and after.

{
  "model": "MiniMax-H3",
  "content": [
    {
      "type": "text",
      "text": "The girl naturally grows from childhood to youth, time flows smoothly, and the character remains in the center of the frame."
    },
    {
      "type": "image_url",
      "image_url": { "url": "YOUR_FIRST_FRAME_URL" },
      "role": "first_frame"
    },
    {
      "type": "image_url",
      "image_url": { "url": "YOUR_LAST_FRAME_URL" },
      "role": "last_frame"
    }
  ],
  "resolution": "2K",
  "duration": 5,
  "ratio": "adaptive"
}

The dimensions and aspect ratios of the first and last frames should be as consistent as possible, and the differences in subject position, composition, and lighting should not be too large, making it easier to achieve a natural transition.

Multi-modal Reference Live Video

Reference materials can be used in combination: reference images control the appearance of characters or products, reference videos control actions and camera movements, and reference audio controls dialogue tone, music, or editing rhythm. The prompt should clearly specify what each type of material is meant to control, avoiding the upload of materials without providing the relationship.

{
  "model": "MiniMax-H3",
  "content": [
    {
      "type": "text",
      "text": "Keep the facial features, hairstyle, and clothing of the reference character consistent, and complete the fashion short film according to the performance actions in the reference video; the camera rhythm follows the reference audio, with close-ups highlighting natural facial expressions."
    },
    {
      "type": "image_url",
      "image_url": { "url": "YOUR_CHARACTER_IMAGE_URL" },
      "role": "reference_image"
    },
    {
      "type": "video_url",
      "video_url": { "url": "YOUR_PERFORMANCE_VIDEO_URL" },
      "role": "reference_video"
    },
    {
      "type": "audio_url",
      "audio_url": { "url": "YOUR_AUDIO_URL" },
      "role": "reference_audio"
    }
  ],
  "resolution": "2K",
  "duration": 5,
  "ratio": "adaptive"
}

Callback Notification

Passing in callback_url will automatically enable asynchronous mode: the creation interface immediately returns task_id and trace_id, and after the task is completed, it POSTs the final result to that address, with a structure consistent with the task query response.

The final status in the callback is succeeded, failed, or cancelled. Even when using callbacks, it is recommended to save task_id for proactive queries or to compensate for missed notifications.

Common Errors

HTTP Status Code Meaning Handling Suggestions
400 Parameter error or invalid material combination Check required fields, role, material quantity, and format
401 Token missing or invalid Check Authorization: Bearer ...
402 Insufficient balance or quota Supplement general balance in the console
422 Content security check failed Adjust prompts or materials and resubmit
429 Too many requests Retry after exponential backoff; task polling is recommended to be about 10 seconds apart
500 Service temporarily unavailable Retain request information and retry later

In synchronous responses, task.status: succeeded indicates that the video has been generated; asynchronous confirmation only means the task has entered the queue. Only when the task is ultimately successful will it be billed; querying the task itself is free and will not incur repeated charges.