Your first H3 Max project should have a clear finish line: one short shot that you can download, review, and place in an edit. A product on a table with a small camera movement is a useful starting brief because mistakes are easy to identify.

This H3 Max tutorial takes you through text-to-video and image-to-video on fal. The settings were checked against official pages on September 4, 2026. The prompts are original examples to try, rather than results from a generation test.

AI-generated example reference image of a blue ceramic mug

AI-created reference-image example for the mug exercise. This is not a fal interface screenshot or a video frame generated by H3 Max.

1. Choose your input and open the right playground

Use H3 Max text to video when you need a new scene from a written brief. Use H3 Max image to video when you already have the photograph and composition you want to animate.

Check that the route says H3 Max. Sign in if prompted, then review the current price or allowance before submitting. Any fal charge is separate from your FreyaVideo balance. H3 Max is not currently integrated into FreyaVideo; the supported-model option appears at the end of this guide.

Choose one placement now: a horizontal website section or a vertical social post. This decision should shape your source image and framing before you generate.

2. Set up a small first attempt

In the text-to-video input form, replace the example prompt and use this starting setup:

Control First attempt What you are deciding
Duration 5 seconds One action to review
Resolution 768P A consistent review size
Prompt Expansion Mode balanced A quick initial prompt rewrite
Aspect Ratio 16:9 or 9:16 Horizontal or vertical placement

The text-to-video schema lists balanced rewriting at about one second and quality at up to roughly 30 seconds. That is preparation time, separate from rendering. Try quality later as a deliberate comparison; its name does not promise a better result for your particular shot.

The official model overview lists 5–15-second clips, native 480P or 768P, and native synchronized audio. Do not select settings from a standard H3 tutorial and assume they belong to H3 Max.

Leave the first attempt's seed unspecified. Save the prompt and settings so later revisions have a baseline.

3. Build a prompt around a shot you can judge

Write the subject first, followed by action, camera movement, lighting, and sound. For this product concept, paste:

A matte blue ceramic mug sits on an oak kitchen table beside a window. A thin stream of steam rises gently. The camera makes a slow, short push toward the mug, keeping its shape and handle visible. Soft morning light, quiet room ambience, no speech. One continuous shot.

AI-generated storyboard showing three progressively closer views of the same mug

Illustrated storyboard for a gradual push-in, not H3 Max output. The three views explain the intended framing progression; they are not frames extracted from a generated video.

Set three acceptance rules before submitting: the mug keeps one attached handle, the camera advances smoothly, and no person enters the scene. These are your review criteria, not a guarantee of what the model will produce.

Submit once and keep that request open while it runs. If it takes longer than expected, check its status before starting another paid attempt. A longer wait alone does not tell you whether the first request failed.

4. Use a photo when the starting appearance matters

For image-to-video, prepare a clear photograph with the entire subject visible. Crop it to your intended layout, remove distracting edges, and leave room for the planned movement. A tiny product surrounded by clutter makes detail checks difficult.

In the image-to-video playground, add the photo under Image URL. The image-to-video schema maps this to image_url: the starting image controls the output aspect ratio. Leave End Image URL empty on this first attempt.

Use this original motion prompt for a bottle photograph:

Keep the bottle upright in its current position. Make a very small camera push-in while a soft reflection moves across the glass. Preserve the bottle silhouette, cap, and existing label layout. Keep the background still. Quiet studio ambience, no dialogue, one continuous shot.

The image supplies the starting appearance; the text describes the movement you want. Avoid asking the camera to orbit behind a product when you only have a front view and need its unseen details to remain accurate.

5. Add an ending frame only for a planned transition

The optional end_image_url guides the last frame. Choose an opening and ending view of the same subject with a plausible connection: for example, a medium shot and a modestly closer composition.

Prepare both images at the same aspect ratio. Compare object position, background, and lighting before uploading. If all three change sharply, you are asking for a scene transformation as well as movement.

Add the second image under End Image URL, then describe how the shot should travel between the two views. Keep the prompt consistent with those images. A supplied close-up ending conflicts with a request to pull away into a wide shot.

Review the middle of the transition as carefully as its endpoints. Matching the destination is insufficient if the product visibly bends on the way there.

6. Check picture, sound, and the downloaded file

Play the complete result once without stopping. Then scrub the beginning, middle, and end to inspect the silhouette, label, handle, cap, and background. Check the composition at the size you intend to publish.

Listen separately for unwanted speech, abrupt noises, or distracting music. Add a specific sound instruction on your next attempt, such as “quiet room tone, no dialogue.” If you will use a different soundtrack, mute or replace the generated track in your editor.

When your request completes, use its download link. Open the saved file locally and play it from beginning to end; a working browser preview is only the first check. Confirm duration, framing, and audio. Save the prompt, source images, settings, and request identifier alongside it.

7. Make one targeted correction

Problem Next change
Product bends or drifts Reduce movement; use a clearer, less cluttered starting image
Camera travels too far Replace the camera sentence with a small push-in or locked view
First and last frames clash Choose a closer ending composition and keep lighting consistent
Label or logo changes Reduce motion; add an accurate graphic in editing if necessary
Unexpected sound Simplify the audio instruction, or replace the track in editing
Request still waiting Check the existing status before deciding whether to submit again
File fails to download Retry the completed result's download before paying for a new generation

Keep each previous attempt so you can compare the actual defect. Changing the photo, action, camera, and duration together makes the improvement harder to explain or repeat.

Continue with a useful next step

Read H3 Max vs Kling for a model comparison, or the fastest AI video generator guide for measuring total waiting time.

To animate a photo with a model currently offered in FreyaVideo, open image to video, choose an available model, and check its settings and credit estimate. H3 Max itself remains a fal workflow for now.

Frequently asked questions

Where can I use H3 Max?

Use fal's H3 Max text-to-video or image-to-video playground. H3 Max is not currently available in FreyaVideo's model selector.

What settings should I try first?

Start with a five-second clip, 768P resolution, and balanced prompt expansion. Keep one subject and one camera movement so you can identify what needs changing.

How do the first and last images work?

On the image-to-video route, image_url supplies the first frame and sets the output aspect ratio. The optional end_image_url guides the final frame. These controls do not guarantee that every product detail stays exact.

Does H3 Max generate native 1080p or 2K video?

The H3 Max endpoints checked for this guide offer native 480P and 768P, with durations from 5 to 15 seconds. Standard MiniMax H3 is a different model with different settings.