Hailuo takes its camera direction as a bracketed command from a closed list of fifteen, placed inline where the move happens—[Push in], [Tracking shot], [Static shot]. It is also picture only, so every word you spend on sound is wasted. Six seconds start at 380 credits.
Wan 3.0 reads an additive formula—entity, scene, motion, then the look, then the sound—and it renders audio in the same pass. Leave the camera unnamed and it will cut inside a clip you wanted as one take. Five seconds runs from 400 credits at 480p.
On Hunyuan3D's image-to-3D modes the text field is not read at all—the model's own contract says it accepts no prompt. Every bit of control you have is in how you prepare the input image. Meshes start at 600 credits on Rapid and 1,000 on Pro.
Seedream 5.0 Pro has no aspect ratio parameter. The shape lives in size, written as WIDTH*HEIGHT with an asterisk, and a ratio typed into the caption is ignored. Here are the thirteen legal sizes, the thinking switch, and the caption it rewards. From 144 credits per image.
A multi-model AI platform runs video, image, audio and 3D models from many labs in one workspace, on one balance. Choose one by the exact model versions it carries, whether it prices each run before it starts, and how it helps you prompt each model.
Nano Banana Pro re-renders the entire image on every edit, so an instruction without a preservation clause changes things you never asked about. Say what changes, say what stays, and point at references as Image 1. Edits start at 175 credits at 1k.
HappyHorse generates picture and sound together, but it does not infer sound from what it sees: a prompt with no audio line returns a near-silent video. It is also one of the few video models that honors timecoded shots, which is the opposite of the Seedance 2.5 rule.
FLUX 3 Video renders picture and sound in one pass, so the audio is written into the prompt rather than switched on. Quote every line and give it a visible speaker, or it can appear as text burned into the frame. On LUVI, clips start at 1,870 credits for five seconds at 720p.
Kling 3 reads a shot as subject, movement and scene first, then camera, light and atmosphere: 60 to 100 words, one camera move, one action. The whole prompt is hard-capped at 2,500 characters and an overrun fails the submission before anything renders.
Seedance 2.5 reads integer-second timestamps, typed audio channels and one camera move per shot. Most bad results come from prompting it like Seedance 2.0. Here is the structure that works, taken from the guide we run inside LUVI.
Write Veo 3.1 prompts camera first: cinematography, subject, action, context, style. Put dialogue, effects and ambience into the shot, and switch on Audio Generation, which is off by default on LUVI. Veo 3.1 Lite has no audio. Clips start at 800 credits for eight seconds at 720p.