H3 Max /

H3 Max Reference to Video

Combine authorized image, video, or audio references when identity, movement, camera rhythm, voice, or atmosphere needs stronger guidance.

Give every reference one job

Assign a clear purpose to each asset: identity from an image, movement from a clip, voice from audio, or atmosphere from another source. Redundant and conflicting references make the requested result harder to interpret.

Connect references in the prompt

Do not assume the model knows why an asset was attached. State which subject, action, camera quality, voice, or sound should come from each reference and identify anything that must not transfer.

Understand reference token credits

Reference-to-video includes 4,096 reference tokens. Each additional started block of 1,000 tokens costs 8 credits. The interface estimates this separately from the duration and resolution charge before submission.

Use only material you can authorize

Confirm that you have permission to use every uploaded face, voice, recording, logo, and creative asset. Avoid impersonation, deceptive editing, and references whose ownership or consent is unclear.

Questions

Frequently asked questions

What references can H3 Max use?

The fal reference-to-video route can accept image, video, and audio references, subject to its current input requirements.

How many references should I attach?

Use the smallest set that clearly defines the result. More references are not automatically better and can add both ambiguity and token cost.

Are reference credits included in the video estimate?

The interface shows the base generation estimate and accounts for additional reference-token blocks when the included allowance is exceeded.