Which LUVI video models generate sound—and the one line each of them needs

Most video models on LUVI now render sound in the same pass as the picture, but they disagree about almost everything else: whether the switch is on, whether it costs more, and how you write the sound. Here is the audio switch and the audio syntax for each family, checked against their schemas.

Published:

A row of toggle switches on a black console: three glow amber with a different shape of sound rising above each, one sits dark with nothing above it, and the last position is a blank plate with no switch at all.

Most of the video families in the LUVI catalog now generate their own sound—dialogue, effects, ambience and music, rendered in the same pass as the picture. What they do not share is how you ask. Some ship with the audio switch on, one ships with it off and charges more when you turn it on, and several have no switch at all, which means the prompt is the only thing deciding whether you hear anything. Google DeepMind's own Veo guide puts the general rule plainly: "Explicitly define the sounds you want to hear, to match the audio to your visuals." This post is the per-family version of that sentence.

Which models have no audio switch at all?

Five families expose no audio parameter on LUVI, so what you hear is decided entirely by the prompt: MiniMax H3, Grok Imagine Video, Gemini Omni 1.1 Flash, HappyHorse and Veo 3.1 Lite. Four of them always generate sound. Veo 3.1 Lite is the exception in the other direction—it has no audio parameter to switch on.

"No switch" does not mean "same behavior", and the difference matters when you forget to write an audio line:

  • HappyHorse returns a near-silent clip. It does not infer sound from what it sees.
  • MiniMax H3 invents sound. An omitted audio field is filled in, never left silent.
  • Grok Imagine Video gives you silence or random audio, with no way to ask for either.

Three families, three different punishments for the same omission.

Is the audio on by default, and does it cost extra?

Six families ship with the switch on, and on the ones we measured it changes nothing about the price. Seedance 2.5, Wan 3.0, FLUX 3 Video, Kling, PixVerse and Vidu Q3 all default to audio on. Two of them say so in the parameter itself: Seedance 2.5's generate_audio is documented as "Does not affect price", and Wan 3.0's audio as "Same price either way". We checked both claims against live estimates—a five-second Seedance 2.5 clip at 720p estimates at 3,030 credits with the audio on and 3,030 credits with it off, and a five-second FLUX 3 Video clip at 720p estimates at 1,870 credits either way.

Veo is the exception on both counts. Its generate_audio parameter defaults to false and its own description says it "Increases the per-second cost." An eight-second Veo 3.1 Fast clip at 720p estimates at 1,280 credits with the audio off and 1,536 credits with it on (checked September 22, 2026). So the model that arrives "already scored" is the one family that ships silent and bills you for turning the sound on.

There is one trap inside a family that otherwise defaults to on. On Kling's Omni O3 rows, sound defaults to true for text-to-video and image-to-video but to false for reference-to-video, which also carries a separate keep_original_sound for the audio riding along in the reference. If a Kling reference shot came back mute, that is why.

And one family is picture only: Hailuo renders no dialogue, no sound design and no music, so every word you spend on audio in a Hailuo prompt is wasted.

How does each family want the sound written?

This is where the dialects diverge hardest. All of these come from the prompting guides LUVI runs for each family.

  • Veo 3.1. Audio goes into the shot, labeled: the spoken line in quotes attributed to a speaker, SFX: thunder cracks in the distance, Ambient noise: the quiet hum of a starship bridge, music named in prose. Silence is directable as content—"no music, only the wind".
  • Seedance 2.5. Typed channels, with the character widths exactly as shown: music in fullwidth (…), sound effects in ASCII <…>, dialogue in ASCII {…}, on-screen subtitles in lenticular 【…】. Ambient beds ride in plain prose.
  • MiniMax H3. A fielded document. Always emit both overall_soundscape: and non_diegetic_music:—an omitted field is filled with invented sound. Dialogue is a tagged block: <d>[English] I get off at the next station.</d>.
  • Gemini Omni 1.1 Flash. One sentence behind the label Sound design: …, covering effects, ambience and dialogue. Music has to be asked for explicitly, and by its source and processing character rather than its genre—"a low tinny radio broadcast in the background, playing a song".
  • FLUX 3 Video. Four layers—speech, ambience, effects tied to a visible action, music with its place in the mix. One or two is usually right; a busy mix is a documented failure.
  • HappyHorse. A named audio line in three tiers: foreground (dialogue, hero effect), mid (Foley on something visible), background (ambience, room tone, music).
  • Wan 3.0. The sound clause comes last in the additive formula, split into voice, sound effect and background music, each described rather than listed.
  • Vidu Q3 and PixVerse. Sound written as content next to the picture that motivates it: audio includes soft room tone and faint spoon clink.

Why do quotation marks mean opposite things?

Because on two of these families, quotation marks are an instruction to put words on the screen rather than in a mouth—and both of them are good enough at typography for that to be a real risk.

On Gemini Omni 1.1 Flash, dialogue uses a colon and no quotation marks: In a crisp, thoughtful tone, Clara says: It has to be here. Quotation marks are reserved for text that is meant to be visible, as in a street sign that reads something. Carrying the Veo habit across is how a line of dialogue becomes a caption.

On FLUX 3 Video, quotes are correct, but only with a visible speaker. A quoted line with nobody on screen to say it, and no "voiceover" or "narration" cue, can render as text burned into the frame—which follows from Black Forest Labs advertising that FLUX 3 renders typography as a natural part of the scene.

Everywhere else—Veo, Kling, Wan, Vidu Q3, HappyHorse—a quoted line is speech. Kling wants the physical beat first so it knows whose mouth to move, and Wan wants the voice described around the line. Two Google models, two opposite conventions, is the single most expensive thing to get wrong in this list.

How do you ask for silence?

Every family answers this differently, and "write the word silence" is wrong on all of them.

  • A parameter. Veo, Seedance 2.5, Wan 3.0, FLUX 3, Kling, PixVerse and Vidu Q3 all have a switch. That is the reliable way.
  • An omitted channel. On Seedance 2.5 you leave out the () music channel rather than asking for no music.
  • A literal value. On MiniMax H3 silence is written as N/A in the music field. Leaving the field out produces invented sound instead.
  • A sentence. Wan suppresses by sentence at the tail—"No dialogue." "No background music."—and Gemini Omni places bare fragments last: "No dialogue." "No music, just realistic real world sound."
  • Not at all. Grok Imagine Video has no silence flag. If you need a mute clip from it, strip the audio afterwards.

One rule crosses all of them: asking for silence in prose on a model that has a switch tends to produce static or dead air, because you are still asking an audio model to generate something. Partial silence is different, and it is fine everywhere as content: "for the final two seconds, only rain against the window."

How we checked

  • Models: the live text-to-video, image-to-video and reference-to-video rows for Veo 3.1 and 3.1 Fast and Lite, Seedance 2.5, Wan 3.0, FLUX 3 Video, Kling v3.0 and Omni O3, PixVerse v6 and c1, Vidu Q3, MiniMax H3, Gemini Omni 1.1 Flash, HappyHorse 1.1 and Hailuo.
  • What was read: every one of those rows' parameter schemas and default values, the prompting guides LUVI runs for each family, and the Google pages linked below.
  • Date and runs: September 22, 2026. Credit figures come from live estimates on LUVI at the settings named in the text.
  • Results: Seedance 2.5 estimated 3,030 credits for five seconds at 720p with audio on and off; FLUX 3 Video 1,870 credits either way; Veo 3.1 Fast 1,280 credits with audio off and 1,536 with it on, for eight seconds at 720p.
  • Drawbacks: No outputs were generated for this post, so it makes no claim about output quality. The switch defaults and prices above are a snapshot—read the estimate in the Workspace before you run, and check the parameter panel, since a provider can change a default.

Which one should I pick?

Start from what the sound has to do, not from the picture.

  • A character saying specific words goes to a family whose dialogue convention you are willing to follow exactly: Veo 3.1 for quoted lines with named effects and ambience, HappyHorse if the line is in Mandarin, Cantonese, Japanese, Korean, German or French and needs phoneme-level lip-sync, MiniMax H3 if you want speakers, soundscape and score in separate fields.
  • A mood bed under a wordless shot wants a music cue and a named place in the mix. On FLUX 3 Video that is not optional: a mix built only from ambience and effects drifts toward inaudible.
  • Effects that land on a visible action work everywhere, and they work best when you attach them to the thing the camera can see rather than naming the sound alone.
  • A clip you will score yourself should use the switch, not a prompt sentence—and should probably avoid Grok Imagine Video, which has no switch.

In the Workspace, the feather button applies the right dialect for whichever model is selected, which is the fastest way to stop carrying one family's audio syntax into another: what the feather button does and what a prompting guide is. You can also ask for the manual from Claude or ChatGPT through the LUVI connector.

New here? Create a LUVI account and test your audio line on the cheapest short clip in the family you picked.

Sources

More from Comparisons