Slatesslates

docs

Prompting Guide

Every model in Slates reads a prompt differently. This page is what each one pays attention to, what it quietly ignores, and the mistakes that cost you a generation. It is generated from the same source the Slates CLI, the MCP server and Studio Agent are built on, so it cannot drift from the app.

For everything else about Slates, including credit costs, shortcuts and troubleshooting, see the complete reference. You can paste it into any LLM.

Video

Seedance 2.0

Seedance 2.0 is a multimodal director: it reads your text, images, video and audio at once and splits them into a "spatial layer" (what is in frame) and a "temporal layer" (how it changes). So a good prompt is an engineering-style instruction, not a piece of copywriting. Audio is always generated alongside video at no extra cost.

ByteDance's official advanced formula has 8 slots: precise subject + action details + scene/environment + lighting & color tone + camera movement + visual style + image quality + constraints. Sweet spot 60-150 words for a single shot, longer for multi-shot.

Pin the subject in the first sentence

A matte black earbud case sits on a polished obsidian surface...

The first 20-30 words are the identity anchor. If the subject isn't locked in immediately, Seedance will hallucinate new subjects mid-generation.

Shot 1 / Shot 2 / Shot 3 — never time stamps

Shot 1: Side shot of the alley; the man slowly starts running.
Shot 2: He knocks over a fruit stand; the camera shakes and cuts to his face.
Shot 3: He climbs a low wall; the camera pulls back onto the empty street.

Critical. ByteDance: write a "Shot 1 / Shot 2 / Shot 3" storyboard in the order events occur, then merge it into one prompt. Do NOT write "At 4 seconds" or "0:00–0:03" and do not set per-shot durations — Seedance 2.0 does not respond to timestamps at all, and forcing them "may lead to abnormal generation results." Let the plot set the pacing. (Seedance 2.5 is the exception: it does read integer-second timestamps.)

Order inside each shot

camera move → action + expression → position change → audio

ByteDance's recommended per-shot order. Lead with the camera ("slowly push in from a wide shot", "fixed camera position", "cut to..."), then what the subject does, then where they end up, then the sound.

Lighting is a top quality lever

A cool-white diagonal beam from upper left, dust particles drifting through...

Lighting & color tone has its own slot in the official formula. Describe it before or alongside the subject.

Standard camera terms — including shot size

medium shot · close-up · wide shot · slow push-in · smooth lateral tracking · fixed shot

ByteDance: the model has a strong understanding of camera terminology, so use it directly — this is an open vocabulary, not a fixed list, and shot size counts as camera direction. Only ONE camera movement per shot: asking for push, pull, pan and move at once increases image instability.

Slow, gentle, continuous movement

slowly raise a hand · quickly turn the head · walk slowly · sit down naturally with the motion

Official rule: name the body part and quantify range, speed and force — and prefer small continuous movement over sprints, big jumps and violent rolls. Slow-motion is supported in natural language; "fast" is a known quality-degrading word.

Externalize emotion

❌ she looks very sad
✅ head lowering, shoulders trembling slightly, eyes reddening, fingers clutching the corner of her clothing

Replace abstract emotion words with the physical detail that shows them. This is the single highest-leverage habit in ByteDance's guide — the model renders bodies, not adjectives.

Separate camera from subject motion

The earbud rises smoothly. The camera tracks upward.

Two different sentences. Mixing them ("the camera speed ramps as the earbud rises") is a common cause of shaky, glitchy output.

Images, clips and audio in ONE generation

Marcus (image 1) performs the motion from video 1 and uses the voice timbre from audio 1.

Critical. Attaching a clip does NOT mean "edit this clip". A video or audio attachment is a REFERENCE, numbered in the rail exactly like an image, and it sits alongside your images in the same generation — the composer cites them as "image N", "video N", "audio N", in rail order, and shows you the exact sentence before you press Generate. Reorder the tiles to change what those numbers mean. To actually rewrite a clip, use Edit with AI instead — that is a different, deliberate choice. EACH MODALITY IS NUMBERED SEPARATELY, from 1 — two images and one clip are "image 1", "image 2" and "video 1", never a single running count, so an audio reference is "audio 1" no matter how many images sit in front of it.

Give a character a voice

Sarah (image 1) uses the voice timbre from audio 1. She says, "We open in ten minutes."

Critical. An audio reference can mean five different things to Seedance — music, dialogue, voice, tone or timbre — so SAY WHICH. Name it as the voice timbre and the clip supplies the voice while your prompt supplies the words; leave it unroled and the model falls back to dialogue, re-transcribes the clip and speaks ITS words instead (a real take came back as "a map called Slates" for "an app called Slates"). Bind each speaker in a sentence rather than by attachment order — position carries nothing: "Images 1-2 are Character 1 and correspond to Audio 1; Images 3-4 are Character 2 and correspond to Audio 2." Verbatim from ByteDance. Audio-alone works on 2.5; on 2.0 pair it with at least one image or video.

Multi-character shots — forbid twins

Throughout the video, characters with completely identical appearance, clothing, and accessories are prohibited. Do not generate duplicate avatars or a twin effect.

With several characters in frame, Seedance can render the same person twice. ByteDance's fix: bind each character to its image ("Marcus (image 1)"), append that constraint verbatim at the end, and prefer single-person reference photos. Past 4 reference people, stability drops — compose a group still first.

💡 Quality and constraint slots have their own official vocabulary: ask for "HD, rich details, cinematic texture, natural colors, soft lighting" — not "8K / masterpiece / trending on artstation." Seedance has no negative-prompt field, so constraints go inline: "keep it subtitle-free", "do not generate a logo", "do not generate a watermark".

💡 Style block at the end: one primary anchor plus 2-3 supporting details. End with "Single continuous take" if you want one shot with no cuts. Never write "no cut" or "seamless transition" — those aren't in the training vocabulary.

💡 Multi-modal: up to 9 images, 3 videos and 3 audio references — 12 files in total, with the reference video capped at 15 seconds combined and the audio at 15. An audio reference on 2.0 needs at least one image or video alongside it (2.5 accepts audio on its own). Cite them by type and index — "Zhang San@Image 1", or the "Marcus (image 1)" form Slates composes from your @mentions. Never cite an asset ID instead of the image number; the model can't associate the two. Max length: 4,000 characters.

💡 A reference VIDEO changes the price: it bills input seconds PLUS output seconds, summed across every clip attached. Two 5-second references on an 8-second generation bills 18 seconds, not 8. The Generate button and the duration menu both show that total before you commit. Over the cap is refused rather than trimmed, precisely so you are never charged for a clip the model never saw.

💡 Don't cross-pollinate image-model syntax: named lenses, apertures and film stocks ("85mm f/1.4", "Kodak Portra 400") are a Nano Banana lever and a Seedance anti-pattern. Translate them into shot size, depth of field and colour tone instead.

Seedance 2.5

Seedance 2.5 is the default video model. Against 2.0 it buys one 30-second take instead of 15, up to 30 image references (plus 10 video and 10 audio), audio-only references and integer-second timestamps — and it gives up 4K. It runs at 480p, 720p or 1080p, and it costs more than 2.0 at every tier they share, so pick 2.0 when you want the same resolution cheaper or you want 4K.

ByteDance's official advanced formula has 8 slots: precise subject + action details + scene/environment + lighting & color tone + camera movement + visual style + image quality + constraints. Sweet spot 60-150 words for a single shot, longer for multi-shot.

Do not write edit instructions here

❌ a wide shot of the workshop, remove the tripod
✅ the workshop bench, clear and uncluttered

Critical. With references attached, "add", "insert", "remove", "delete", "modify", "replace", "change", "extend" and "continue" make Seedance 2.5 treat the request as a video EDIT, and it then fails on constraints it never set — after the job has queued. Describe the finished frame instead. To actually edit a clip, attach it and pick Seedance 2.5 Edit.

Timestamps work here — they do not on 2.0

0-3 seconds: he steps off the curb, rain starting.
3-7 seconds: headlights sweep across him; he turns.
7-15 seconds: he runs; the camera falls behind.

Critical. Seedance 2.0 ignores timing and answers only to "Shot 1 / Shot 2"; 2.5 acts on whole-second timestamps, and that is what makes a 30-second take controllable rather than just long. Three forms work: intervals ("0-3 seconds…3-7 seconds"), a point ("at the 5-second mark"), or relative ("after 3 seconds"). Whole seconds only, no gaps between intervals, and never to choreograph fast repeated motion. They work on edits too, where a range scopes the change: "…from 4-6 seconds…".

Pin the subject in the first sentence

A matte black earbud case sits on a polished obsidian surface...

The first 20-30 words are the identity anchor. If the subject isn't locked in immediately, Seedance will hallucinate new subjects mid-generation.

Order inside each shot

camera move → action + expression → position change → audio

ByteDance's recommended per-shot order. Lead with the camera ("slowly push in from a wide shot", "fixed camera position", "cut to..."), then what the subject does, then where they end up, then the sound.

Lighting is a top quality lever

A cool-white diagonal beam from upper left, dust particles drifting through...

Lighting & color tone has its own slot in the official formula. Describe it before or alongside the subject.

720p is not the cheap one here

30s · 720p · Face route = 484 credits
15s · 1080p · Seedance 2.0 Face = 411 credits

Critical. Length is what moves the price, and 2.5 doubles the length ceiling — so a 30-second 720p clip can cost more than a 15-second 1080p one, against a 1,000-credit starting balance. Explore at short LENGTH rather than low resolution: cut the seconds to 4-8 while you are finding the shot, and stay at the resolution you actually want. A 480p pass does not de-risk a 720p render — generation is stochastic, so the 720p run is a different take, not the same shot rendered better. The Generate button always shows the exact number first.

Audio-only references

Reference the timbre in audio 1 to generate...

2.5 accepts an audio reference on its own — a voice line, a music bed, a room tone — with no image or video alongside it. 2.0 could not. Audio references never cost extra on any Seedance route.

Standard camera terms — including shot size

medium shot · close-up · wide shot · slow push-in · smooth lateral tracking · fixed shot

ByteDance: the model has a strong understanding of camera terminology, so use it directly — this is an open vocabulary, not a fixed list, and shot size counts as camera direction. Only ONE camera movement per shot: asking for push, pull, pan and move at once increases image instability.

Slow, gentle, continuous movement

slowly raise a hand · quickly turn the head · walk slowly · sit down naturally with the motion

Official rule: name the body part and quantify range, speed and force — and prefer small continuous movement over sprints, big jumps and violent rolls. Slow-motion is supported in natural language; "fast" is a known quality-degrading word.

Externalize emotion

❌ she looks very sad
✅ head lowering, shoulders trembling slightly, eyes reddening, fingers clutching the corner of her clothing

Replace abstract emotion words with the physical detail that shows them. This is the single highest-leverage habit in ByteDance's guide — the model renders bodies, not adjectives.

Separate camera from subject motion

The earbud rises smoothly. The camera tracks upward.

Two different sentences. Mixing them ("the camera speed ramps as the earbud rises") is a common cause of shaky, glitchy output.

Images, clips and audio in ONE generation

Marcus (image 1) performs the motion from video 1 and uses the voice timbre from audio 1.

Critical. Attaching a clip does NOT mean "edit this clip". A video or audio attachment is a REFERENCE, numbered in the rail exactly like an image, and it sits alongside your images in the same generation — the composer cites them as "image N", "video N", "audio N", in rail order, and shows you the exact sentence before you press Generate. Reorder the tiles to change what those numbers mean. To actually rewrite a clip, use Edit with AI instead — that is a different, deliberate choice. EACH MODALITY IS NUMBERED SEPARATELY, from 1 — two images and one clip are "image 1", "image 2" and "video 1", never a single running count, so an audio reference is "audio 1" no matter how many images sit in front of it.

Give a character a voice

Sarah (image 1) uses the voice timbre from audio 1. She says, "We open in ten minutes."

Critical. An audio reference can mean five different things to Seedance — music, dialogue, voice, tone or timbre — so SAY WHICH. Name it as the voice timbre and the clip supplies the voice while your prompt supplies the words; leave it unroled and the model falls back to dialogue, re-transcribes the clip and speaks ITS words instead (a real take came back as "a map called Slates" for "an app called Slates"). Bind each speaker in a sentence rather than by attachment order — position carries nothing: "Images 1-2 are Character 1 and correspond to Audio 1; Images 3-4 are Character 2 and correspond to Audio 2." Verbatim from ByteDance. Audio-alone works on 2.5; on 2.0 pair it with at least one image or video.

Multi-character shots — forbid twins

Throughout the video, characters with completely identical appearance, clothing, and accessories are prohibited. Do not generate duplicate avatars or a twin effect.

With several characters in frame, Seedance can render the same person twice. ByteDance's fix: bind each character to its image ("Marcus (image 1)"), append that constraint verbatim at the end, and prefer single-person reference photos. Past 4 reference people, stability drops — compose a group still first.

💡 30 image references is a budget, not a target — 2-4 strong references still beat both extremes, one per role. ByteDance's own ceilings for 2.5: 1-8 subjects bound by image reference stay stable (9-12 works but needs re-rolls), 1-5 subjects bound by video or audio reference, and 5-10 seconds is the sweet spot for a reference clip. Unlike 2.0, a multi-view turnaround sheet can be a single subject reference here — past 5 subjects, go back to one view per image. The larger budget is for long multi-shot takes and for video plus audio references alongside images.

💡 A reference VIDEO bills input seconds PLUS output seconds, and 2.5 accepts references up to 30s combined — so a 20-second reference driving a 20-second output bills 40 seconds. The Generate button shows the total.

💡 Quality and constraint slots have their own official vocabulary: ask for "HD, rich details, cinematic texture, natural colors, soft lighting" — not "8K / masterpiece / trending on artstation." Seedance has no negative-prompt field, so constraints go inline: "keep it subtitle-free", "do not generate a logo", "do not generate a watermark".

💡 Style block at the end: one primary anchor plus 2-3 supporting details. End with "Single continuous take" if you want one shot with no cuts. Never write "no cut" or "seamless transition" — those aren't in the training vocabulary.

💡 Frames and reference images stay mutually exclusive, and on a first/last-frame generation Seedance 2.5 chooses the aspect ratio itself — the ratio control shows "Adaptive" because the start frame decides the shape.

Seedance 2.5 Edit

Seedance 2.5 Edit changes an existing clip: attach the clip, describe only what should be different, and the original motion, framing and timing are kept. It is the only editor in Slates that takes a clip longer than 15 seconds — 4 to 30s, against Kling O3 Edit's 3-15s and Omni Flash Edit's 3-10s.

Output length and aspect ratio follow the SOURCE clip, so there is no duration or ratio control — the clip you attach is the quote. Output is 480p, 720p or 1080p with native audio.

Name the change, keep the rest

Strictly edit the clip, and change the blue jacket from navy to red.

Critical. The clip already carries its composition, motion, timing and performance — re-describing them fights the model. Say the change as "from A to B" rather than as an outcome: naming what it currently is tells the model what to overwrite. One change per pass; chain passes for compound edits. Never write "reference the video" in an edit: that phrasing gets the request re-read as a fresh generation inspired by your clip instead of an edit of it.

Scope the edit in time

Change the man's action from drinking coffee to mopping the floor from 4-6 seconds, and leave the rest of the clip unchanged.

Critical. Seedance 2.0 ignores timing and answers only to "Shot 1 / Shot 2"; 2.5 acts on whole-second timestamps, and that is what makes a 30-second take controllable rather than just long. Three forms work: intervals ("0-3 seconds…3-7 seconds"), a point ("at the 5-second mark"), or relative ("after 3 seconds"). Whole seconds only, no gaps between intervals, and never to choreograph fast repeated motion. They work on edits too, where a range scopes the change: "…from 4-6 seconds…".

Turn Face on when a face is visible

Critical. The default provider blocks character faces outright — this is not a price optimisation, it is whether the job runs at all. There is no consented-real-face route for editing; real-person footage the Face route rejects has to go to Kling O3 Edit.

An edit costs about double a generation

Every provider bills an edit on the input clip AND the output, so a 20-second edit is priced like 40 seconds of generation. Read the number on the Generate button rather than reasoning from the generation rate.

When to use it instead of the others

Length is the reason: it is the only engine that accepts a clip over 15 seconds. Inside the others' range, choose on fidelity — Omni Flash Edit is the prompt-only fidelity winner and the cheapest seat, and Kling O3 Edit is the one that takes subject and style reference images.

It edits the audio too

Only edit the man's dialogue: change it to "Don't come over here," in an American accent. Keep everything else unchanged.

The same engine rewrites what is heard while the picture stays put — change a line, change an accent, translate the dialogue and re-fit the lip movement, or strip and replace music and sound effects. Priced like any other edit, on the source clip's length.

Prompt and clip only

No character or style reference images on this engine in Slates today. If the edit needs a reference image to lock an identity, that is Kling O3 Edit's job.

💡 Trim before you edit, not after: the bill is the source clip's length rounded up, so a 30-second clip you only needed 8 seconds of costs nearly four times what it had to.

💡 Clips under 20 seconds edit most reliably. Up to 30 is accepted, and the returned clip can land within about a third of a second of the source length — only transition frames are compressed, nothing is cut.

Kling 3.0

Kling 3.0 features native audio-visual co-generation with dialogue, sound effects, and music (Omni tier). Define your core subjects clearly at the beginning of the prompt and keep descriptions consistent across shots.

Dialogue

Character says, "exact words here"

Use quotation marks for precise speech. Languages (Omni only): English, Chinese, Japanese, Korean, Spanish.

Voice Quality

with a trembling voice, "I'm scared"

Describe emotional tone, pitch, or speaking style before the dialogue. No pronouns or synonyms after a character's first introduction — they cause voice drift.

Sound Effects

SFX: heavy boots on wet pavement, distant siren wailing

Use the "SFX:" prefix, with physical-cause specificity — "SFX: footsteps" is too vague.

Multi-Character Dialogue (Omni)

Alice says in English, "Hello!" Immediately, Bob replies in Spanish, "¡Hola!"

The "Immediately" keyword makes lines back-to-back; without it Kling adds a natural conversational beat.

Ambient Noise & Music

Ambient noise: city traffic, birds chirping
Background music: tense orchestral strings

Set the background soundscape and request specific music styles or moods.

Multi-shot

Shot 1: ... Shot 2: ...

Max 6 cuts, 15s total. One primary action and ONE camera move per shot; describe the subject identically in every shot block.

💡 Keep dialogue concise (under 10 seconds per line). Use the Language and Accent settings in Audio Controls to control speech characteristics.

Kling O3 Edit

Kling 3.0 features native audio-visual co-generation with dialogue, sound effects, and music (Omni tier). Define your core subjects clearly at the beginning of the prompt and keep descriptions consistent across shots.

Dialogue

Character says, "exact words here"

Use quotation marks for precise speech. Languages (Omni only): English, Chinese, Japanese, Korean, Spanish.

Voice Quality

with a trembling voice, "I'm scared"

Describe emotional tone, pitch, or speaking style before the dialogue. No pronouns or synonyms after a character's first introduction — they cause voice drift.

Sound Effects

SFX: heavy boots on wet pavement, distant siren wailing

Use the "SFX:" prefix, with physical-cause specificity — "SFX: footsteps" is too vague.

Multi-Character Dialogue (Omni)

Alice says in English, "Hello!" Immediately, Bob replies in Spanish, "¡Hola!"

The "Immediately" keyword makes lines back-to-back; without it Kling adds a natural conversational beat.

Ambient Noise & Music

Ambient noise: city traffic, birds chirping
Background music: tense orchestral strings

Set the background soundscape and request specific music styles or moods.

Multi-shot

Shot 1: ... Shot 2: ...

Max 6 cuts, 15s total. One primary action and ONE camera move per shot; describe the subject identically in every shot block.

Video edit — name the change, keep the rest

Replace the man in @Video1 with @Element1, keeping his walk cycle, the camera move, and the rain unchanged.

Critical. @Video1 is your clip; attached subject refs compile to @Element1..; style refs to @Image1.. (max 4 combined). One edit intent per pass — chain passes for compound changes. Original audio is preserved verbatim.

💡 Keep dialogue concise (under 10 seconds per line). Use the Language and Accent settings in Audio Controls to control speech characteristics.

Gemini Omni Flash

Gemini Omni Flash is the cheap 720p tier with native synced audio included — dialogue, SFX, and ambient generate WITH the video at no extra cost. 3-10s, 16:9 or 9:16. Text-to-video, one start frame, or up to 7 reference images. No last frame, no video/audio references.

Structure like a shot brief

subject + action + setting + camera + lighting + tone

Descriptive prompts are fine for generation (the short-prompt rule is edit-only).

Dialogue

The barista says, "Your usual?"

Audio is prompt-driven — there are no audio parameters. Dialogue in quotes.

Sound in plain language

rain patters on the tin roof · distant traffic hum

Describe sounds directly in the prose — no SFX: prefix needed.

Name references inline

Marcus (image 1) walks into the cafe...

Up to 7 reference images merge into one list — refer to them by number in the prompt.

Negatives as plain instructions

Do not show text.

No negative-prompt field — write what to avoid as a direct instruction.

Know its seat

Cheap drafts, iteration volume, and audio-in-one-gen at low cost. For hero shots, Seedance 2.5 (the default) or Seedance 2.0 (4K, cheaper) still win.

Omni Flash Edit

Omni Flash Edit changes what the prompt names in an existing 3-10s clip, footage-synced — prop, effect, environment, and lighting swaps. Prompt + source clip only: no reference images (identity swaps that need refs → Kling O3 Edit). 720p output; voice editing unsupported.

One short change — MANDATORY

Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.

Critical. Google's own doc: simple prompts work best; overly descriptive prompts cause unintended changes. Long "keep every frame identical" preambles make drift WORSE. One change, then the magic phrase.

Always end with the preservation phrase

...Keep everything else the same.

The one documented preservation lever. Every edit prompt ends with it.

Never name objects as metaphors

❌ a candle-like flame  →  ✅ small magical flames on his fingertips

"Candle-like" renders a literal candle in his hand. Describe the effect itself.

No conditional timing cues

❌ ...appears WHEN he calls it, perches AS he walks

Beat-by-beat stage directions cued to moments in the footage hard-fail the request. Collapse to one continuous action; the model syncs it to the footage's own motion.

Frame effects as harmless VFX

❌ his fingertips catch fire  →  ✅ magical flames appear on his fingertips

Google's safety filter is strict about harm-to-person phrasing. Magical/harmless framing passes.

Expect a possible tail artifact

Occasional jitter or a doubled final speech beat in the last ~0.5s. Trim the tail on the timeline — don't burn a re-roll on it.

💡 Ship via segment-splice: edit only the seconds where the change happens (Trim / Split first), then splice back over the original on the timeline with the original audio underneath. Chain edits one change at a time — each edit saves as a new clip linked to its parent.

MiniMax H3

MiniMax H3 generates picture and sound in one pass — 24fps, 32kHz stereo, 5-15 seconds, 11 stably-supported languages. It is the only video model in Slates where audio is AUTHORED rather than switched on: synchronised dialogue and action sounds go in the body of the prompt, ambience goes in a soundscape section, and audience-only music goes in a score section. Put a sound in the wrong section and it is dropped, doubled, or attributed to the wrong source.

Two seats that differ in LADDER and PRICE, not in what they accept. Base H3 runs 480p / 768p / 2K / 4K; H3 Max is fal's faster post-train and runs 480p / 768p / 1080p, dearer than base H3 at the tier they share - a deliberate speed pick, never the cheap one. BOTH read up to 9 reference images plus 3 video and 3 audio clips (12 files total, and audio never travels alone), and both animate a start frame and an end frame. 768p is the default on both because it is the tier the model natively generates; base H3's 2K and 4K are upscales of a 768p base. Reference images past the free allowance are billed and the allowances DIFFER: 5 free on base H3, 4 on Max.

Three audio layers, three places

body: "First batch of the morning."
Soundscape: shutters scrape, trays clink
Score: solo piano, slow, no swell

Critical. Dialogue, singing and diegetic music (a radio in the scene) go in the BODY on the beat they land. Ambience goes in the soundscape. The score is audience-only — name instruments and tempo, not moods.

Shots and cut times

[Shot 2] At 00:03.500, the camera cuts to...

The first shot carries no timestamp; later shots open with the bracket and a rising cut time inside the clip length. Transition verbs: cuts to / transitions to / changes to / switches to.

Write the camera into the sentence

The camera pushes in with small amplitude at slow speed toward the letter in her hands.

Named moves (push in, pull out, arc, tracking, POV, roll) with amplitude and speed modifiers. Never stack them as labels.

Voiceover needs both halves

says in an off-screen voiceover: "..." — his lips remain completely closed.

The off-screen phrase alone still animates a mouth. State the closed lips explicitly.

Cite references by number

Marcus (image 1) walks into the workshop (image 2)...

H3 takes references as typed slots and expects plain numbered prose — image 1, video 1, audio 1, each modality counted separately from 1. fal's own prompt-field description: "Refer to reference assets by their modality and order in the reference lists: Image 1, Image 2, Video 1, Audio 1, and so on." Do not hand-write angle-bracket tags — MiniMax's guides use <Subject N> / <Audio N> as DOCUMENTATION notation and typing them puts literal brackets in the prompt. An audio reference binds as a voice timbre to a named speaker, the same primitive Seedance uses. Name each reference inline; never write role essays. Slates does this for you: @mention a subject or environment and it composes Marcus (image 1) in the cafe (image 2), citing them in the exact order it sends them. One canonical identity image avoids competing facial renderings; a "Reference Image Instructions" block drags reference lighting into your scene. Start with 2-3 focused refs.

Say how much of a reference survives

Give the man in image 3 the weathered leather texture of the jacket in image 4.

H3 is the only seat that understands transferring a characteristic onto a DIFFERENT subject. State each reference's job and how much of it should carry through — kept whole, kept in part, transferred, or a loose echo.

Reference images past the fifth cost extra

Critical. The first 5 are free; each one after that adds 4 credits at every resolution and length, and the model takes 9. Four extra images on a 10s 768p clip add 16 credits to a 30-credit generation. Attach what the shot needs, not the ceiling.

2K and 4K are upscales, not bigger renders

Critical. Only 480p and 768p are generated natively; 2K and 4K enlarge a finished 768p take. In our own testing the 2K pass showed MORE artifacting than the 768p original while costing 33 credits for a 5-second take against 15. Generate and judge at 768p; step up only when a delivery spec demands the pixels.

Frames or references, never both

A start and/or end frame runs on a different endpoint from references — the reference endpoint has no frame slots. Slates refuses the combination rather than dropping one side. An audio reference also cannot travel alone: pair it with an image or video.

💡 Slates disables the provider's prompt expander, so what you write is what the model reads — nothing will pad a thin prompt. Aim for a 350-500 word body on a reference-carrying shot, and let dialogue-heavy scenes run longer if that is what fits the spoken timeline.

LTX-2.5

LTX-2.5 scores the picture on the same pass that draws it, so SOUND IS THE FIRST THING YOU WRITE, not the last. Lightricks ranks the six parts of a prompt in this order: sound, camera, character detail, shot type and scene, then scene dressing — and scene dressing is the first thing to cut when a prompt sprawls. Everything goes in ONE flowing paragraph, not a list of labelled sections.

Two seats. Base LTX-2.5 is the distilled build: 720p / 1080p / 1440p / 4K and clips from 6 to 20 seconds, and it is the cheapest native 1080p second in Slates. LTX-2.5 Pro is the full diffusion build ("Diffusion Fidelity Rendering" spends extra compute on busy frames) but reaches a SHORTER ladder — 1080p and 10 seconds maximum — while costing about a third more. Pro is for a dense final render; base is for iteration, long takes and 4K.

Anchor every sound to something in frame

the rope creaks against the cleat, gulls somewhere off the port bow

Critical. Write the audio line last, then check each cue has a visible or at least locatable source. Anything unanchored gets invented for you. "Not visible but locatable" passes — a whistle is fine if you name the marshal post it comes from.

Never write mood words for sound

Bad: "tense atmosphere, sense of dread"
Good: "a loose shutter knocks twice against the frame"

Atmosphere adjectives produce nothing. If a scene feels thin, add one more MOVING OBJECT with a sound attached to it rather than another adjective.

Dialogue takes quotes, language and accent

"We should not have come back," in English with a slight German accent.

Give the character a beat of stillness before they speak so the lip sync has something to lock against. Describe the beat: looks, waits, speaks, looks away.

Emotion is physical, not abstract

Bad: "she looks anxious"
Good: "her jaw sets, she turns the ring on her finger twice"

The model renders actions, not adjectives. Tension in the jaw, weight shifts, fidgeting hands — these read; "anxious" does not.

Multishot: two to four shots, and re-establish at every cut

wide establishing shot — hard cut — macro close-up — match cut — medium shot

Critical. One generation can carry several connected shots holding character, light and voice across the cuts. Name the edit ("hard cut", "dissolve") in the prose, then RESET scale, angle, lens and light. Two to four is the working range.

Re-identify characters at every cut

Good: "the woman in the bronze gown"
Bad: "she"

Pronouns lose the character across a cut. Repeat the original descriptor every time. Also state what the SOUND does at the cut — silence is not assumed.

Write the camera into the sentence

a slow push-in settles as she reaches the door, then holds

Slates does not expose the camera_motion enum, and prose is the better tool anyway: a written move can be tied to a specific moment, an enum value cannot. Name lens, framing and the moment the move resolves.

Durations are even numbers only, from six

Critical. 6, 8, 10, 12, 14, 16, 18 or 20 seconds — there is no 5s or 7s LTX clip. And the long end is 1080p-and-below only: at 1440p and 4K the ceiling drops to 10s. Slates always sends an explicit length rather than letting the model pick one, so what you choose is what you are billed for.

Do not ask for text on screen

Neither the spelling nor its stability frame to frame can be relied on. Signage, labels and captions belong in post.

💡 Frames, not references. LTX takes a start frame and an optional end frame (which generates a transition between the two) — it has no reference endpoint at all, so identity, style and environment reference images are not available on this model. For character consistency across separate shots, use MiniMax H3 or Kling.

💡 Aspect ratios are 16:9 and 9:16 only, and native audio is included free at every resolution — there is no sound surcharge on either seat.

💡 In image-to-video, do not cut away from the opening frame too early: you have paid for that frame, so let it play before the first move.

Veo 3.1

Veo 3.1 generates synchronized audio directly with video. Aspect ratio: 16:9 only (for 9:16 vertical, use Kling or Seedance). Native single-clip duration: 4, 6, or 8 seconds — longer durations require chaining clips via last-frame reuse.

Official Cloud formula: [Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]. Sweet spot ~50-150 words.

Dialogue

Character says, "exact words"

Use quotation marks for exact speech. Keep voice direction terse: "says in a weary voice", "whispers", "shouts". 2-3 speakers max — sync degrades past that.

Sound Effects — with cause

SFX: thunder cracks in the distance

Always specify direction or distance — "SFX: thunder" alone is too vague.

Ambient is mandatory

Soft office ambience. · Wind on the open ridge.

Include an ambience line in every scene — without it the audio mix feels dead.

No subtitles — MANDATORY

The founder says, "..." (no subtitles). Soft office ambience.

Critical. Without (no subtitles) after every dialogue line, Veo bakes subtitle text into the video. This is genuinely critical and underspecified in most guides.

Cinematography vocabulary

85mm · shallow depth of field · Rembrandt lighting · dolly in · whip pan

Veo responds to real lens, lighting, and camera-move terms — lead the prompt with them.

💡 First-frame + last-frame is Veo's strongest workflow. Generate a start frame, generate an end frame, then animate with both as anchors. Motion-Lock hack: keep ~60% of the same background pixels between start and end to prevent latent drift.

💡 Keep dialogue under one natural breath — lines fit the 8s clip ceiling. Texture-realism phrases: fine skin pores, visible fabric weave, subtle contrast, no gloss or sharpening.

Image

Nano Banana 2

Nano Banana 2 is a language model that outputs pixels. Brief it like a creative director, not like a Stable-Diffusion tag tool. The biggest realism lever: specificity that mimics how real photographers describe their work.

Google's 4 official rules: Be specific. Use positive framing (describe what you want, not what you don't). Control the camera with cinematic terms. Iterate conversationally.

Cinematic prompt formula

Film still from [Director] [genre]. Shot on [camera] with [lens]. [Subject + action]. [3-5 details]. [Lighting]. [Color palette]. [Film stock].

Specific gear beats generic descriptors. "ARRI Alexa 65 with Panavision anamorphic" outperforms "cinematic camera."

Named lenses + apertures

85mm f/1.4 · 135mm f/2.8 · 50mm f/1.2 · 35mm f/2 · Panavision anamorphic · 400mm telephoto

135mm f/2.8 is the cheat code for skin texture and intimate compression. Anamorphic for cinematic width + horizontal flares.

Named film stocks (one per prompt)

Kodak Portra 400 · Fuji Velvia 50 · Ilford HP5 Plus · CineStill 800T

Portra = natural skin warmth. Velvia = saturated landscape. HP5 = gritty B&W grain. CineStill 800T = tungsten night with halation. Never mix stocks.

Don't carry lens + stock into a video prompt

85mm f/1.4, Portra 400
→ close-up, shallow depth of field, warm natural colors, cinematic texture

Lenses, apertures, film stocks and camera bodies are an image-model lever and a video-model anti-pattern — ByteDance's Seedance guide never mentions f-stops, lens millimetres, fps or shutter angle. When you animate a frame you made here, translate the look into shot size, depth of field and colour tone instead of pasting the gear list across.

Physics-based lighting

Single key light at 45 degrees from upper left. Color temperature 4500K. Crisp catchlights in the eyes.

Direction + Kelvin temp + named source. "Single key light at 10 o'clock" beats "soft lighting" every time.

Imperfection vocabulary

visible pores · peach fuzz · ISO noise · sweat beading · slight hyperpigmentation · unretouched raw photography

Forces the model away from AI-clean skin. The default is too smooth — you have to ask for the imperfections that real photos have.

❌ The anti-list — avoid these

8k · masterpiece · hyperrealistic · ultra-detailed · trending on ArtStation · perfect skin · flawless · airbrushed · cinematic (alone)

Critical. Tag-soup phrases from the Stable-Diffusion era. Measured success ~60-70% with these vs ~95%+ with positive description. Always specify which cinema — director, lens, era, stock.

No negative-prompt field

✅ "empty street" not "no cars"
✅ "without people, vehicles, or signage"
❌ "not anime, not cartoon, not 3D"

Reframe positively first. Use inline "without" / "free of" only when positive framing can't suppress the unwanted element.

Reference images — name them, never label roles

Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3).

Up to 14 refs (10 object + 4 character — caps don't trade). Name each reference inline; never write role essays. Slates does this for you: @mention a subject or environment and it composes Marcus (image 1) in the cafe (image 2), citing them in the exact order it sends them. One canonical identity image avoids competing facial renderings; a "Reference Image Instructions" block drags reference lighting into your scene. Start with 2-3 focused refs.

Common fixes

Hands → "five fingers, natural proportions"
Text → quote-wrap "HEADLINE" + specify font
Left/right → "from the character's perspective"

Default left/right is the viewer's perspective. Surreal prompts trip uncanny valley — the model drags toward realism. For surrealism, lean hard into "painted" / "illustrated".

Resolution tactics

1k = drafts · 2k = hero · 4k = print/final

Pick by need. 2K+ allocates more tokens to surface detail, so texture vocab (pores, fabric weave, grain) compounds at higher resolution.

💡 Boring vs Cinema. Boring: "Wide shot of man on dock looking at forest." Cinema: "Direct overhead drone shot on weathered dock. Single figure climbing up frame bottom. Boot prints leading toward shore. Pale winter light. Anamorphic flare. Desaturated blue/slate palette. Kodak Portra 400 grain. Map of threat."

💡 3-strike rule. If three iterations on the same prompt haven't landed, stop. The slot machine doesn't converge — the prompt structure is wrong, not the seed.

Nano Banana 2 Lite

Nano Banana 2 is a language model that outputs pixels. Brief it like a creative director, not like a Stable-Diffusion tag tool. The biggest realism lever: specificity that mimics how real photographers describe their work.

Google's 4 official rules: Be specific. Use positive framing (describe what you want, not what you don't). Control the camera with cinematic terms. Iterate conversationally.

Cinematic prompt formula

Film still from [Director] [genre]. Shot on [camera] with [lens]. [Subject + action]. [3-5 details]. [Lighting]. [Color palette]. [Film stock].

Specific gear beats generic descriptors. "ARRI Alexa 65 with Panavision anamorphic" outperforms "cinematic camera."

Named lenses + apertures

85mm f/1.4 · 135mm f/2.8 · 50mm f/1.2 · 35mm f/2 · Panavision anamorphic · 400mm telephoto

135mm f/2.8 is the cheat code for skin texture and intimate compression. Anamorphic for cinematic width + horizontal flares.

Named film stocks (one per prompt)

Kodak Portra 400 · Fuji Velvia 50 · Ilford HP5 Plus · CineStill 800T

Portra = natural skin warmth. Velvia = saturated landscape. HP5 = gritty B&W grain. CineStill 800T = tungsten night with halation. Never mix stocks.

Don't carry lens + stock into a video prompt

85mm f/1.4, Portra 400
→ close-up, shallow depth of field, warm natural colors, cinematic texture

Lenses, apertures, film stocks and camera bodies are an image-model lever and a video-model anti-pattern — ByteDance's Seedance guide never mentions f-stops, lens millimetres, fps or shutter angle. When you animate a frame you made here, translate the look into shot size, depth of field and colour tone instead of pasting the gear list across.

Physics-based lighting

Single key light at 45 degrees from upper left. Color temperature 4500K. Crisp catchlights in the eyes.

Direction + Kelvin temp + named source. "Single key light at 10 o'clock" beats "soft lighting" every time.

Imperfection vocabulary

visible pores · peach fuzz · ISO noise · sweat beading · slight hyperpigmentation · unretouched raw photography

Forces the model away from AI-clean skin. The default is too smooth — you have to ask for the imperfections that real photos have.

❌ The anti-list — avoid these

8k · masterpiece · hyperrealistic · ultra-detailed · trending on ArtStation · perfect skin · flawless · airbrushed · cinematic (alone)

Critical. Tag-soup phrases from the Stable-Diffusion era. Measured success ~60-70% with these vs ~95%+ with positive description. Always specify which cinema — director, lens, era, stock.

No negative-prompt field

✅ "empty street" not "no cars"
✅ "without people, vehicles, or signage"
❌ "not anime, not cartoon, not 3D"

Reframe positively first. Use inline "without" / "free of" only when positive framing can't suppress the unwanted element.

Reference images — name them, never label roles

Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3).

Up to 14 refs (10 object + 4 character — caps don't trade). Name each reference inline; never write role essays. Slates does this for you: @mention a subject or environment and it composes Marcus (image 1) in the cafe (image 2), citing them in the exact order it sends them. One canonical identity image avoids competing facial renderings; a "Reference Image Instructions" block drags reference lighting into your scene. Start with 2-3 focused refs.

Common fixes

Hands → "five fingers, natural proportions"
Text → quote-wrap "HEADLINE" + specify font
Left/right → "from the character's perspective"

Default left/right is the viewer's perspective. Surreal prompts trip uncanny valley — the model drags toward realism. For surrealism, lean hard into "painted" / "illustrated".

Resolution tactics

1k only on Lite

Lite outputs 1K only — use it for iteration volume and drafts, then switch to Nano Banana 2 for 2K/4K finals.

💡 Boring vs Cinema. Boring: "Wide shot of man on dock looking at forest." Cinema: "Direct overhead drone shot on weathered dock. Single figure climbing up frame bottom. Boot prints leading toward shore. Pale winter light. Anamorphic flare. Desaturated blue/slate palette. Kodak Portra 400 grain. Map of threat."

💡 3-strike rule. If three iterations on the same prompt haven't landed, stop. The slot machine doesn't converge — the prompt structure is wrong, not the seed.

Audio

Seed Audio 1.0

Seed Audio builds a whole audio scene — dialogue, effects and ambience together — from one plain sentence. Write it the way you would describe the moment to a person standing next to you, not the way you would write a video prompt.

It has no duration setting. Length comes from the words, so Slates appends your chosen duration to the prompt ("… 15 seconds") and bills exactly that. Set the duration control to what you actually want and let the sentence stay clean.

One plain sentence

nature soundscape, wide open field cicadas and birds and a loon.

No shot language, no production jargon, no formatting. Plain description outperforms anything that reads like a spec sheet.

Duration lives in the prompt

tiny applause of 2 or 3 people at an open mic. 15 seconds

Critical. The duration control writes this for you. Do not also type a different length into your sentence — the two will fight and you pay for the one you selected.

Say the crowd size

tiny applause of 2 or 3 people  ·  a packed arena roaring

"Applause" alone returns a full room. Scale words are the single highest-leverage edit on any crowd, traffic or nature bed.

Cut beds longer than the shot

Ask for a few seconds more than the clip needs so the edit has handles to fade in and out of. Beds that end exactly on the cut always sound clipped.

No Kling syntax here

✗ SFX: heavy boots
✓ heavy boots on wet pavement, a siren far off

Critical. The "SFX:" and "Ambient noise:" prefixes belong to Kling video prompts. Seed Audio treats them as words in the scene and the result gets worse.

Dialogue in quotes

a tired bartender says, "we closed twenty minutes ago", glasses clinking behind him

Speech goes in quotes inside the same sentence as the room. Pick a preset voice for a specific speaker, or leave it unset and let the scene cast itself.

References

match the room tone of @Audio1

Up to 3 audio clips (max 30s each), referenced as @Audio1–@Audio3 — OR one image to score what is in frame. Never both in the same generation.

Know its seat

Scenes, beds, room tone and dialogue in one pass. For a single effect that has to land on a specific frame, use Sound Effects.

ElevenLabs Sound Effects

Sound Effects makes one short sound with an exact length — the lane for a hit that has to land on a specific frame, or a seamless loop you can lay under a whole scene.

Duration is always sent explicitly (0.5–22s). Billing is per second, so the length you pick is the price you pay.

Describe the cause, not the label

✗ door sound
✓ heavy oak door slams shut in a stone hallway

Critical. Material, weight and room are what separate a usable effect from a stock-library shrug. Name all three.

One sound per generation

This surface makes a single event. A door, then footsteps, then a siren is three generations layered on the timeline — or one Seed Audio scene.

Duration is the edit

0.8s for an impact · 4s for a whoosh · 22s for a bed

Ask for roughly the length you need. A 4-second request for a door slam pads the tail with room tone you then have to trim.

Loops

steady rain on a canvas tent  (loop on, 12s)

Turn loop on for anything continuous — rain, engine hum, crowd murmur — and it will tile without a seam.

Prompt influence

0.3 default · 0.7 literal

Higher hugs your wording with less variation between takes; lower explores. Raise it when a re-roll keeps wandering off the brief.

Know its seat

One precise effect on a known frame. Full rooms and layered scenes are cheaper and better in one Seed Audio pass.

Inworld TTS-2

The voice seat: one named voice saying one line. Unlike every other surface in Slates, the prompt is not a description of what you want — it IS the words that get spoken, verbatim, and its length is what you are billed for.

A voice belongs to a character, the same way a face does. Build it once from a clip or a description, then send it lines.

Direction goes in SQUARE BRACKETS

✗ (quietly) I hope nobody notices
✓ [whispering] I hope nobody notices

Critical. Brackets are read as direction and never spoken. PARENTHESES ARE SPOKEN ALOUD — a parenthetical stage direction comes back with the narrator saying the word "quietly". This is the single easiest way to ruin a take.

Plain English inside the brackets

[very quiet] · [very slow] · [say excitedly] · [whisper in a hushed style]

It is natural-language steering across emotion, volume, pitch, speed, articulation and vocal style — not a fixed vocabulary. Write the direction the way you would say it to an actor.

Non-verbals are their own tags

[laugh] · [sigh] · [breathe] · [cough] · [yawn] · [clear throat]

They land inline, where they occur in the line.

A mistyped tag fails silently

Critical. An instruction it does not recognise is still swallowed and still changes the delivery — it is never spoken and never errors. So a typo produces a strange read with no warning. If a take sounds off, suspect the tag before the voice.

Tags persist until changed

[very slow] First line. Second line is still slow. [reset] Third is normal.

A direction governs everything after it, across sentences. Use [reset] to return to a normal read rather than assuming the next sentence starts clean.

Punctuation is the timing

Wait. Stop.  ≠  Wait, stop.

Full stops buy a beat; commas do not. Write the line the way it is said.

One line, one take

Split a paragraph into separate generations so a bad clause costs one re-roll instead of the whole speech. Max 2,000 characters per take.

Spell out anything ambiguous

twenty twenty-six · Doctor Reyes

Numbers, dates and abbreviations are read literally. Write them as they should sound.

Know its seat

One voice, cleanly. Dialogue mixed with effects and room tone in one pass is Seed Audio; a single non-speech sound is Sound Effects.

Going deeper

The tips above are the short version. The full prompting craft ships as the Slates skills library with the community membership: per-model long-form guides, character identity, style prompting, cost discipline, the vision feedback loop and campaign blueprints. See pricing.

One-time purchase · 30-day money-back guarantee