How to prompt Wan 3.0—name the camera in every shot, or it cuts for you

Wan 3.0 reads an additive formula—entity, scene, motion, then the look, then the sound—and it renders audio in the same pass. Leave the camera unnamed and it will cut inside a clip you wanted as one take. Five seconds runs from 400 credits at 480p.

Published:

Two widescreen film frames float above a glowing amber timeline in a dark studio: the left one whole and steady on a small brass tripod head, the right one split into three misaligned slivers with nothing holding it.

Wan 3.0 is Alibaba's joint audio-video model, and on LUVI it is the largest model family in the catalog. It runs as text-to-video, image-to-video and reference-to-video, from 2 to 30 seconds, at 480p, 720p or 1080p. A five-second clip came to 400 credits at 480p, 800 at 720p and 1,600 at 1080p when we quoted it on September 22, 2026; the higher-fidelity Wan 3.0 Prime models quoted 2,520 for five seconds at 1080p. What follows is how the model reads a prompt—starting with the failure people write to us about most.

Why did my clip cut in the middle?

Wan cuts because nothing in your prompt told it how to hold the frame. Alibaba's guide asks you to name the camera treatment of every shot, and LUVI's own prompting guide for the family puts the consequence plainly: left unnamed, Wan edits on your behalf and cuts inside a single clip. The fix is one clause per shot—"fixed camera", "slow push-in", "tracking shot"—and one main camera action, not two.

That matters because Wan has a sanctioned way to give you several shots, and it is not silence. Alibaba's "what not to prompt" list names "rapid scene changes within a single clip" and explains why: "Cuts happen between clips, not within them. One clip = one continuous shot." If you want shots, you number them.

How do I write more than one shot?

You write the multi-shot formula, and Wan 3.0 honors the timestamps in it. Alibaba gives the shape as "Overall description + Shot number + Timestamp + Shot content", with the example form Shot 1 [0–3 s] xxxx, Shot 2 [3–6 s] xxx. The overall description carries the story and the subject's look once, so you are not re-describing the person in every shot; each shot then states its own framing, its own camera move and its own transition.

Overall: a quiet winter morning in a mountain village, warm and unhurried, one woman throughout. Shot 1 [0–4 s] A woman in a thick grey wool coat unlatches a wooden gate with one gloved hand, her breath clouding; medium shot, soft overcast light from the left, fixed camera. Shot 2 [4–8 s] Hard cut transition. The same woman walks down a snow-dusted lane toward the camera; wide shot, slow push-in, low winter sun behind her. Sound effects: boots compressing fresh snow, the iron latch clicking open, distant crows calling. No dialogue. No background music.

Four to six seconds per shot is the working range. Timestamps that reach past the clip length are wasted, so keep the last one inside the duration you set.

What does the rest of the prompt look like?

It is an additive formula, read in order, not a pile of keywords. Alibaba publishes it as "Prompt = Entity (description) + Scene (description) + Motion (description) + Aesthetic control + Stylization", and the guide is blunt about what happens when you stop after the first three: "Most weak prompts cover entity and motion but skip aesthetic control and stylization — producing a static camera in an undefined void."

  • Entity. Who or what, resolved to appearance: "a young woman with wavy auburn hair in a vintage floral dress".
  • Scene. The environment, foreground and background.
  • Motion. What moves, how far, how fast: "swaying violently", "moving slowly".
  • Aesthetic control. Light source, lighting, shot size, camera angle, lens, camera movement. This is the slot the weak prompts skip.
  • Stylization. One style word—cyberpunk, line-art illustration, claymation, felt, pixel—not a stack of them.

How do I write the sound?

You describe it as content, because Wan 3.0 generates audio in the same pass and the audio parameter is on by default. Turning it off does not make the clip cheaper: the model's own contract says the price is the same either way. Alibaba splits sound into three formulas—voice as "Character's lines + Emotion + Tone + Speed + Timbre + Accent", a sound effect as "Source material + Action + Ambient sound", and background music as score plus style. Spoken lines go inside double quotes, and so does any text that should appear on screen.

There is no negative prompt on the Wan video models, so exclusions are written as sentences at the tail: "No dialogue." "No background music." "No on-screen text, no subtitles." If you are unsure which models in the catalog have the field at all, the help center keeps a list of where a negative prompt exists.

How do I point at an image or a video I uploaded?

By type and number. Alibaba's rule is literal: "For images, write 'Image n'. For videos, write 'Video n'." Images and videos are counted separately, in upload order, and the identifier is used as a noun in the sentence—"The cat in Image 1 plays in the room in Image 2." In English the guide asks for a capital letter and a space between the word and the number.

LUVI's Workspace inserts its own handle, @Image 1, when you attach a reference. Keep it, and write the Wan form beside it the first time it appears. A reference video lends appearance, motion and voice timbre at once, so say which of the three you actually want from it.

What belongs in the parameters instead of the prompt?

Four things, and writing them into the text wastes words: duration (2 to 30 seconds, 5 by default), resolution (480p, 720p, 1080p), ratio (adaptive by default, or 16:9, 4:3, 1:1, 3:4, 9:16) and audio. Never type "16:9", "1080p" or "10 seconds" into the prompt itself.

Two of them move the price. Length is billed per second, and resolution multiplies that per-second rate: our three quotes for the same five seconds stepped 400, 800 and 1,600 credits as the resolution went up. The help center explains how resolution and duration change the price across the catalog.

How we checked

  • Models: Wan 3.0 Text-to-Video, Image-to-Video and Reference-to-Video, plus Wan 3.0 Prime Text-to-Video.
  • Settings: five-second clips at 480p, 720p and 1080p, 16:9.
  • Date and what was read: September 22, 2026—the models' parameter contracts and live credit quotes in LUVI, LUVI's own prompting guide for the Wan family, and Alibaba Cloud's text-to-video prompt guide (last updated September 4, 2026).
  • Results: 400, 800 and 1,600 credits for five seconds at 480p, 720p and 1080p; 2,520 credits for Wan 3.0 Prime at 1080p. No outputs were generated for this post, so it makes no claim about output quality.
  • Drawbacks: the four-to-six-second-per-shot range and the behavior of unnamed camera moves come from LUVI's prompting guide rather than a published benchmark.

In LUVI

Pick a Wan model from the Wan family page, write the formula in order, and check it before you spend anything: the feather button beside the prompt box runs LUVI's prompting guide for whichever model you selected and rewrites the prompt in that model's dialect. Attach references the way the reference media guide describes, then read the credit estimate above the Generate button—it is the ceiling, not a guess.

If you would rather drive it from a chat window, LUVI's connector exposes the same models and the same prompting guides to Claude and ChatGPT; the MCP documentation covers the setup. New accounts start with 1,000 credits, so you can create a LUVI account and run the example above without buying anything first.

Sources

More from Guides