Slatesslates

docs

Prompting Guide

Every model in Slates reads a prompt differently. This page is what each one pays attention to, what it quietly ignores, and the mistakes that cost you a generation. It is generated from the same source the Slates CLI, the MCP server and Studio Agent are built on, so it cannot drift from the app.

For everything else about Slates, including credit costs, shortcuts and troubleshooting, see the complete reference. You can paste it into any LLM.

Video

Seedance 2.0

Seedance 2.0 is a multimodal director: it reads your text, images, video and audio at once and splits them into a "spatial layer" (what is in frame) and a "temporal layer" (how it changes). So a good prompt is an engineering-style instruction, not a piece of copywriting. Audio is always generated alongside video at no extra cost.

ByteDance's official advanced formula has 8 slots: precise subject + action details + scene/environment + lighting & color tone + camera movement + visual style + image quality + constraints. Sweet spot 60-150 words for a single shot, longer for multi-shot.

Pin the subject in the first sentence

A matte black earbud case sits on a polished obsidian surface...

The first 20-30 words are the identity anchor. If the subject isn't locked in immediately, Seedance will hallucinate new subjects mid-generation.

Shot 1 / Shot 2 / Shot 3 — never time stamps

Shot 1: Side shot of the alley; the man slowly starts running.
Shot 2: He knocks over a fruit stand; the camera shakes and cuts to his face.
Shot 3: He climbs a low wall; the camera pulls back onto the empty street.

Critical. ByteDance: write a "Shot 1 / Shot 2 / Shot 3" storyboard in the order events occur, then merge it into one prompt. Do NOT write "At 4 seconds" or "0:00–0:03" and do not set per-shot durations — official docs say precise timing is unstable and forcing it "may lead to abnormal generation results." Let the plot set the pacing.

Order inside each shot

camera move → action + expression → position change → audio

ByteDance's recommended per-shot order. Lead with the camera ("slowly push in from a wide shot", "fixed camera position", "cut to..."), then what the subject does, then where they end up, then the sound.

Lighting is a top quality lever

A cool-white diagonal beam from upper left, dust particles drifting through...

Lighting & color tone has its own slot in the official formula. Describe it before or alongside the subject.

Standard camera terms — including shot size

medium shot · close-up · wide shot · slow push-in · smooth lateral tracking · fixed shot

ByteDance: the model has a strong understanding of camera terminology, so use it directly — this is an open vocabulary, not a fixed list, and shot size counts as camera direction. Only ONE camera movement per shot: asking for push, pull, pan and move at once increases image instability.

Slow, gentle, continuous movement

slowly raise a hand · quickly turn the head · walk slowly · sit down naturally with the motion

Official rule: name the body part and quantify range, speed and force — and prefer small continuous movement over sprints, big jumps and violent rolls. Slow-motion is supported in natural language; "fast" is a known quality-degrading word.

Externalize emotion

❌ she looks very sad
✅ head lowering, shoulders trembling slightly, eyes reddening, fingers clutching the corner of her clothing

Replace abstract emotion words with the physical detail that shows them. This is the single highest-leverage habit in ByteDance's guide — the model renders bodies, not adjectives.

Separate camera from subject motion

The earbud rises smoothly. The camera tracks upward.

Two different sentences. Mixing them ("the camera speed ramps as the earbud rises") is a common cause of shaky, glitchy output.

Images, clips and audio in ONE generation

Marcus (image 1) performs the motion from video 1, speaking the line in audio 1.

Critical. Attaching a clip does NOT mean "edit this clip". A video or audio attachment is a REFERENCE, numbered in the rail exactly like an image, and it sits alongside your images in the same generation — the composer cites them as "image N", "video N", "audio N", in rail order, and shows you the exact sentence before you press Generate. Reorder the tiles to change what those numbers mean. To actually rewrite a clip, use Edit with AI instead — that is a different, deliberate choice.

Multi-character shots — forbid twins

Throughout the video, characters with completely identical appearance, clothing, and accessories are prohibited. Do not generate duplicate avatars or a twin effect.

With several characters in frame, Seedance can render the same person twice. ByteDance's fix: bind each character to its image ("Marcus (image 1)"), append that constraint verbatim at the end, and prefer single-person reference photos. Past 4 reference people, stability drops — compose a group still first.

💡 Quality and constraint slots have their own official vocabulary: ask for "HD, rich details, cinematic texture, natural colors, soft lighting" — not "8K / masterpiece / trending on artstation." Seedance has no negative-prompt field, so constraints go inline: "keep it subtitle-free", "do not generate a logo", "do not generate a watermark".

💡 Style block at the end: one primary anchor plus 2-3 supporting details. End with "Single continuous take" if you want one shot with no cuts. Never write "no cut" or "seamless transition" — those aren't in the training vocabulary.

💡 Multi-modal: up to 9 images, 3 videos and 3 audio references — 12 files in total, with the reference video capped at 15 seconds combined and the audio at 15. An audio reference on 2.0 needs at least one image or video alongside it (2.5 accepts audio on its own). Cite them by type and index — "Zhang San@Image 1", or the "Marcus (image 1)" form Slates composes from your @mentions. Never cite an asset ID instead of the image number; the model can't associate the two. Max length: 4,000 characters.

💡 A reference VIDEO changes the price: it bills input seconds PLUS output seconds, summed across every clip attached. Two 5-second references on an 8-second generation bills 18 seconds, not 8. The Generate button and the duration menu both show that total before you commit. Over the cap is refused rather than trimmed, precisely so you are never charged for a clip the model never saw.

💡 Don't cross-pollinate image-model syntax: named lenses, apertures and film stocks ("85mm f/1.4", "Kodak Portra 400") are a Nano Banana lever and a Seedance anti-pattern. Translate them into shot size, depth of field and colour tone instead.

Seedance 2.5

Seedance 2.5 is a SECOND SEAT next to 2.0, not an upgrade of it. It buys one 30-second take instead of 15, up to 30 image references (plus 10 video and 10 audio), and audio-only references — and it gives up 1080p and 4K entirely. It is 480p or 720p, on every route. Everything below about writing the prompt is the same as 2.0.

ByteDance's official advanced formula has 8 slots: precise subject + action details + scene/environment + lighting & color tone + camera movement + visual style + image quality + constraints. Sweet spot 60-150 words for a single shot, longer for multi-shot.

Do not write edit instructions here

❌ a wide shot of the workshop, remove the tripod
✅ the workshop bench, clear and uncluttered

Critical. With references attached, "add", "remove", "replace", "change", "extend" and "continue" make Seedance 2.5 treat the request as a video EDIT, and it then fails on constraints it never set — after the job has queued. Describe the finished frame instead. To actually edit a clip, attach it and pick Seedance 2.5 Edit.

Pin the subject in the first sentence

A matte black earbud case sits on a polished obsidian surface...

The first 20-30 words are the identity anchor. If the subject isn't locked in immediately, Seedance will hallucinate new subjects mid-generation.

Shot 1 / Shot 2 / Shot 3 — never time stamps

Shot 1: Side shot of the alley; the man slowly starts running.
Shot 2: He knocks over a fruit stand; the camera shakes and cuts to his face.
Shot 3: He climbs a low wall; the camera pulls back onto the empty street.

Critical. ByteDance: write a "Shot 1 / Shot 2 / Shot 3" storyboard in the order events occur, then merge it into one prompt. Do NOT write "At 4 seconds" or "0:00–0:03" and do not set per-shot durations — official docs say precise timing is unstable and forcing it "may lead to abnormal generation results." Let the plot set the pacing.

Order inside each shot

camera move → action + expression → position change → audio

ByteDance's recommended per-shot order. Lead with the camera ("slowly push in from a wide shot", "fixed camera position", "cut to..."), then what the subject does, then where they end up, then the sound.

Lighting is a top quality lever

A cool-white diagonal beam from upper left, dust particles drifting through...

Lighting & color tone has its own slot in the official formula. Describe it before or alongside the subject.

720p is not the cheap one here

30s · 720p · Face route = 484 credits
15s · 1080p · Seedance 2.0 Face = 411 credits

Critical. Length is what moves the price, and 2.5 doubles the length ceiling — so a 30-second 720p clip can cost more than a 15-second 1080p one, against a 1,000-credit starting balance. Draft at 480p and 4-8 seconds; spend the length only on a take you already know works. The Generate button always shows the exact number first.

Audio-only references

Reference the timbre in audio 1 to generate...

2.5 accepts an audio reference on its own — a voice line, a music bed, a room tone — with no image or video alongside it. 2.0 could not. Audio references never cost extra on any Seedance route.

Standard camera terms — including shot size

medium shot · close-up · wide shot · slow push-in · smooth lateral tracking · fixed shot

ByteDance: the model has a strong understanding of camera terminology, so use it directly — this is an open vocabulary, not a fixed list, and shot size counts as camera direction. Only ONE camera movement per shot: asking for push, pull, pan and move at once increases image instability.

Slow, gentle, continuous movement

slowly raise a hand · quickly turn the head · walk slowly · sit down naturally with the motion

Official rule: name the body part and quantify range, speed and force — and prefer small continuous movement over sprints, big jumps and violent rolls. Slow-motion is supported in natural language; "fast" is a known quality-degrading word.

Externalize emotion

❌ she looks very sad
✅ head lowering, shoulders trembling slightly, eyes reddening, fingers clutching the corner of her clothing

Replace abstract emotion words with the physical detail that shows them. This is the single highest-leverage habit in ByteDance's guide — the model renders bodies, not adjectives.

Separate camera from subject motion

The earbud rises smoothly. The camera tracks upward.

Two different sentences. Mixing them ("the camera speed ramps as the earbud rises") is a common cause of shaky, glitchy output.

Images, clips and audio in ONE generation

Marcus (image 1) performs the motion from video 1, speaking the line in audio 1.

Critical. Attaching a clip does NOT mean "edit this clip". A video or audio attachment is a REFERENCE, numbered in the rail exactly like an image, and it sits alongside your images in the same generation — the composer cites them as "image N", "video N", "audio N", in rail order, and shows you the exact sentence before you press Generate. Reorder the tiles to change what those numbers mean. To actually rewrite a clip, use Edit with AI instead — that is a different, deliberate choice.

Multi-character shots — forbid twins

Throughout the video, characters with completely identical appearance, clothing, and accessories are prohibited. Do not generate duplicate avatars or a twin effect.

With several characters in frame, Seedance can render the same person twice. ByteDance's fix: bind each character to its image ("Marcus (image 1)"), append that constraint verbatim at the end, and prefer single-person reference photos. Past 4 reference people, stability drops — compose a group still first.

💡 30 image references is a budget, not a target. Every reference rule still holds: 2-4 strong references beat both extremes, one reference per role, one authoritative rendering per subject — and past 4 reference PEOPLE, output stability drops regardless of the cap. The larger budget is for long multi-shot takes and for video plus audio references alongside images.

💡 A reference VIDEO bills input seconds PLUS output seconds, and 2.5 accepts references up to 30s combined — so a 20-second reference driving a 20-second output bills 40 seconds. The Generate button shows the total.

💡 Quality and constraint slots have their own official vocabulary: ask for "HD, rich details, cinematic texture, natural colors, soft lighting" — not "8K / masterpiece / trending on artstation." Seedance has no negative-prompt field, so constraints go inline: "keep it subtitle-free", "do not generate a logo", "do not generate a watermark".

💡 Style block at the end: one primary anchor plus 2-3 supporting details. End with "Single continuous take" if you want one shot with no cuts. Never write "no cut" or "seamless transition" — those aren't in the training vocabulary.

💡 Frames and reference images stay mutually exclusive, and on a first/last-frame generation Seedance 2.5 chooses the aspect ratio itself — the ratio control shows "Adaptive" because the start frame decides the shape.

Seedance 2.5 Edit

Seedance 2.5 Edit changes an existing clip: attach the clip, describe only what should be different, and the original motion, framing and timing are kept. It is the only editor in Slates that takes a clip longer than 15 seconds — 4 to 30s, against Kling O3 Edit's 3-15s and Omni Flash Edit's 3-10s.

Output length and aspect ratio follow the SOURCE clip, so there is no duration or ratio control — the clip you attach is the quote. Output is 480p or 720p with native audio.

Name the change, keep the rest

Strictly edit the clip, and change the blue jacket to a red one.

Critical. The clip already carries its composition, motion, timing and performance — re-describing them fights the model. One change per pass; chain passes for compound edits. Never write "reference the video" in an edit: that phrasing gets the request re-read as a fresh generation inspired by your clip instead of an edit of it.

Turn Face on when a face is visible

Critical. The default provider blocks character faces outright — this is not a price optimisation, it is whether the job runs at all. There is no consented-real-face route for editing; real-person footage the Face route rejects has to go to Kling O3 Edit.

An edit costs about double a generation

Every provider bills an edit on the input clip AND the output, so a 20-second edit is priced like 40 seconds of generation. Read the number on the Generate button rather than reasoning from the generation rate.

When to use it instead of the others

Length is the reason: it is the only engine that accepts a clip over 15 seconds. Inside the others' range, choose on fidelity — Omni Flash Edit is the prompt-only fidelity winner and the cheapest seat, and Kling O3 Edit is the one that takes subject and style reference images.

Prompt and clip only

No character or style reference images on this engine. If the edit needs a reference image to lock an identity, that is Kling O3 Edit's job.

💡 Trim before you edit, not after: the bill is the source clip's length rounded up, so a 30-second clip you only needed 8 seconds of costs nearly four times what it had to.

Kling 3.0

Kling 3.0 features native audio-visual co-generation with dialogue, sound effects, and music (Omni tier). Define your core subjects clearly at the beginning of the prompt and keep descriptions consistent across shots.

Dialogue

Character says, "exact words here"

Use quotation marks for precise speech. Languages (Omni only): English, Chinese, Japanese, Korean, Spanish.

Voice Quality

with a trembling voice, "I'm scared"

Describe emotional tone, pitch, or speaking style before the dialogue. No pronouns or synonyms after a character's first introduction — they cause voice drift.

Sound Effects

SFX: heavy boots on wet pavement, distant siren wailing

Use the "SFX:" prefix, with physical-cause specificity — "SFX: footsteps" is too vague.

Multi-Character Dialogue (Omni)

Alice says in English, "Hello!" Immediately, Bob replies in Spanish, "¡Hola!"

The "Immediately" keyword makes lines back-to-back; without it Kling adds a natural conversational beat.

Ambient Noise & Music

Ambient noise: city traffic, birds chirping
Background music: tense orchestral strings

Set the background soundscape and request specific music styles or moods.

Multi-shot

Shot 1: ... Shot 2: ...

Max 6 cuts, 15s total. One primary action and ONE camera move per shot; describe the subject identically in every shot block.

💡 Keep dialogue concise (under 10 seconds per line). Use the Language and Accent settings in Audio Controls to control speech characteristics.

Kling O3 Edit

Kling 3.0 features native audio-visual co-generation with dialogue, sound effects, and music (Omni tier). Define your core subjects clearly at the beginning of the prompt and keep descriptions consistent across shots.

Dialogue

Character says, "exact words here"

Use quotation marks for precise speech. Languages (Omni only): English, Chinese, Japanese, Korean, Spanish.

Voice Quality

with a trembling voice, "I'm scared"

Describe emotional tone, pitch, or speaking style before the dialogue. No pronouns or synonyms after a character's first introduction — they cause voice drift.

Sound Effects

SFX: heavy boots on wet pavement, distant siren wailing

Use the "SFX:" prefix, with physical-cause specificity — "SFX: footsteps" is too vague.

Multi-Character Dialogue (Omni)

Alice says in English, "Hello!" Immediately, Bob replies in Spanish, "¡Hola!"

The "Immediately" keyword makes lines back-to-back; without it Kling adds a natural conversational beat.

Ambient Noise & Music

Ambient noise: city traffic, birds chirping
Background music: tense orchestral strings

Set the background soundscape and request specific music styles or moods.

Multi-shot

Shot 1: ... Shot 2: ...

Max 6 cuts, 15s total. One primary action and ONE camera move per shot; describe the subject identically in every shot block.

Video edit — name the change, keep the rest

Replace the man in @Video1 with @Element1, keeping his walk cycle, the camera move, and the rain unchanged.

Critical. @Video1 is your clip; attached subject refs compile to @Element1..; style refs to @Image1.. (max 4 combined). One edit intent per pass — chain passes for compound changes. Original audio is preserved verbatim.

💡 Keep dialogue concise (under 10 seconds per line). Use the Language and Accent settings in Audio Controls to control speech characteristics.

Gemini Omni Flash

Gemini Omni Flash is the cheap 720p tier with native synced audio included — dialogue, SFX, and ambient generate WITH the video at no extra cost. 3-10s, 16:9 or 9:16. Text-to-video, one start frame, or up to 7 reference images. No last frame, no video/audio references.

Structure like a shot brief

subject + action + setting + camera + lighting + tone

Descriptive prompts are fine for generation (the short-prompt rule is edit-only).

Dialogue

The barista says, "Your usual?"

Audio is prompt-driven — there are no audio parameters. Dialogue in quotes.

Sound in plain language

rain patters on the tin roof · distant traffic hum

Describe sounds directly in the prose — no SFX: prefix needed.

Name references inline

Marcus (image 1) walks into the cafe...

Up to 7 reference images merge into one list — refer to them by number in the prompt.

Negatives as plain instructions

Do not show text.

No negative-prompt field — write what to avoid as a direct instruction.

Know its seat

Cheap drafts, iteration volume, and audio-in-one-gen at low cost. For hero shots, Kling 3.0 (general default) or Seedance 2.0 (premium/physics) still win.

Omni Flash Edit

Omni Flash Edit changes what the prompt names in an existing 3-10s clip, footage-synced — prop, effect, environment, and lighting swaps. Prompt + source clip only: no reference images (identity swaps that need refs → Kling O3 Edit). 720p output; voice editing unsupported.

One short change — MANDATORY

Small magical flames appear on his fingertips when he snaps his fingers, and vanish when he blows on them. Keep everything else the same.

Critical. Google's own doc: simple prompts work best; overly descriptive prompts cause unintended changes. Long "keep every frame identical" preambles make drift WORSE. One change, then the magic phrase.

Always end with the preservation phrase

...Keep everything else the same.

The one documented preservation lever. Every edit prompt ends with it.

Never name objects as metaphors

❌ a candle-like flame  →  ✅ small magical flames on his fingertips

"Candle-like" renders a literal candle in his hand. Describe the effect itself.

No conditional timing cues

❌ ...appears WHEN he calls it, perches AS he walks

Beat-by-beat stage directions cued to moments in the footage hard-fail the request. Collapse to one continuous action; the model syncs it to the footage's own motion.

Frame effects as harmless VFX

❌ his fingertips catch fire  →  ✅ magical flames appear on his fingertips

Google's safety filter is strict about harm-to-person phrasing. Magical/harmless framing passes.

Expect a possible tail artifact

Occasional jitter or a doubled final speech beat in the last ~0.5s. Trim the tail on the timeline — don't burn a re-roll on it.

💡 Ship via segment-splice: edit only the seconds where the change happens (Trim / Split first), then splice back over the original on the timeline with the original audio underneath. Chain edits one change at a time — each edit saves as a new clip linked to its parent.

Veo 3.1

Veo 3.1 generates synchronized audio directly with video. Aspect ratio: 16:9 only (for 9:16 vertical, use Kling or Seedance). Native single-clip duration: 4, 6, or 8 seconds — longer durations require chaining clips via last-frame reuse.

Official Cloud formula: [Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]. Sweet spot ~50-150 words.

Dialogue

Character says, "exact words"

Use quotation marks for exact speech. Keep voice direction terse: "says in a weary voice", "whispers", "shouts". 2-3 speakers max — sync degrades past that.

Sound Effects — with cause

SFX: thunder cracks in the distance

Always specify direction or distance — "SFX: thunder" alone is too vague.

Ambient is mandatory

Soft office ambience. · Wind on the open ridge.

Include an ambience line in every scene — without it the audio mix feels dead.

No subtitles — MANDATORY

The founder says, "..." (no subtitles). Soft office ambience.

Critical. Without (no subtitles) after every dialogue line, Veo bakes subtitle text into the video. This is genuinely critical and underspecified in most guides.

Cinematography vocabulary

85mm · shallow depth of field · Rembrandt lighting · dolly in · whip pan

Veo responds to real lens, lighting, and camera-move terms — lead the prompt with them.

💡 First-frame + last-frame is Veo's strongest workflow. Generate a start frame, generate an end frame, then animate with both as anchors. Motion-Lock hack: keep ~60% of the same background pixels between start and end to prevent latent drift.

💡 Keep dialogue under one natural breath — lines fit the 8s clip ceiling. Texture-realism phrases: fine skin pores, visible fabric weave, subtle contrast, no gloss or sharpening.

Image

Nano Banana 2

Nano Banana 2 is a language model that outputs pixels. Brief it like a creative director, not like a Stable-Diffusion tag tool. The biggest realism lever: specificity that mimics how real photographers describe their work.

Google's 4 official rules: Be specific. Use positive framing (describe what you want, not what you don't). Control the camera with cinematic terms. Iterate conversationally.

Cinematic prompt formula

Film still from [Director] [genre]. Shot on [camera] with [lens]. [Subject + action]. [3-5 details]. [Lighting]. [Color palette]. [Film stock].

Specific gear beats generic descriptors. "ARRI Alexa 65 with Panavision anamorphic" outperforms "cinematic camera."

Named lenses + apertures

85mm f/1.4 · 135mm f/2.8 · 50mm f/1.2 · 35mm f/2 · Panavision anamorphic · 400mm telephoto

135mm f/2.8 is the cheat code for skin texture and intimate compression. Anamorphic for cinematic width + horizontal flares.

Named film stocks (one per prompt)

Kodak Portra 400 · Fuji Velvia 50 · Ilford HP5 Plus · CineStill 800T

Portra = natural skin warmth. Velvia = saturated landscape. HP5 = gritty B&W grain. CineStill 800T = tungsten night with halation. Never mix stocks.

Don't carry lens + stock into a video prompt

85mm f/1.4, Portra 400
→ close-up, shallow depth of field, warm natural colors, cinematic texture

Lenses, apertures, film stocks and camera bodies are an image-model lever and a video-model anti-pattern — ByteDance's Seedance guide never mentions f-stops, lens millimetres, fps or shutter angle. When you animate a frame you made here, translate the look into shot size, depth of field and colour tone instead of pasting the gear list across.

Physics-based lighting

Single key light at 45 degrees from upper left. Color temperature 4500K. Crisp catchlights in the eyes.

Direction + Kelvin temp + named source. "Single key light at 10 o'clock" beats "soft lighting" every time.

Imperfection vocabulary

visible pores · peach fuzz · ISO noise · sweat beading · slight hyperpigmentation · unretouched raw photography

Forces the model away from AI-clean skin. The default is too smooth — you have to ask for the imperfections that real photos have.

❌ The anti-list — avoid these

8k · masterpiece · hyperrealistic · ultra-detailed · trending on ArtStation · perfect skin · flawless · airbrushed · cinematic (alone)

Critical. Tag-soup phrases from the Stable-Diffusion era. Measured success ~60-70% with these vs ~95%+ with positive description. Always specify which cinema — director, lens, era, stock.

No negative-prompt field

✅ "empty street" not "no cars"
✅ "without people, vehicles, or signage"
❌ "not anime, not cartoon, not 3D"

Reframe positively first. Use inline "without" / "free of" only when positive framing can't suppress the unwanted element.

Reference images — name them, never label roles

Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3).

Up to 14 refs (10 object + 4 character — caps don't trade). Name each reference inline; never write role essays. Slates does this for you: @mention a subject or environment and it composes Marcus (image 1) in the cafe (image 2), citing them in the exact order it sends them. One canonical identity image avoids competing facial renderings; a "Reference Image Instructions" block drags reference lighting into your scene. Start with 2-3 focused refs.

Common fixes

Hands → "five fingers, natural proportions"
Text → quote-wrap "HEADLINE" + specify font
Left/right → "from the character's perspective"

Default left/right is the viewer's perspective. Surreal prompts trip uncanny valley — the model drags toward realism. For surrealism, lean hard into "painted" / "illustrated".

Resolution tactics

1k = drafts · 2k = hero · 4k = print/final

Pick by need. 2K+ allocates more tokens to surface detail, so texture vocab (pores, fabric weave, grain) compounds at higher resolution.

💡 Boring vs Cinema. Boring: "Wide shot of man on dock looking at forest." Cinema: "Direct overhead drone shot on weathered dock. Single figure climbing up frame bottom. Boot prints leading toward shore. Pale winter light. Anamorphic flare. Desaturated blue/slate palette. Kodak Portra 400 grain. Map of threat."

💡 3-strike rule. If three iterations on the same prompt haven't landed, stop. The slot machine doesn't converge — the prompt structure is wrong, not the seed.

Nano Banana 2 Lite

Nano Banana 2 is a language model that outputs pixels. Brief it like a creative director, not like a Stable-Diffusion tag tool. The biggest realism lever: specificity that mimics how real photographers describe their work.

Google's 4 official rules: Be specific. Use positive framing (describe what you want, not what you don't). Control the camera with cinematic terms. Iterate conversationally.

Cinematic prompt formula

Film still from [Director] [genre]. Shot on [camera] with [lens]. [Subject + action]. [3-5 details]. [Lighting]. [Color palette]. [Film stock].

Specific gear beats generic descriptors. "ARRI Alexa 65 with Panavision anamorphic" outperforms "cinematic camera."

Named lenses + apertures

85mm f/1.4 · 135mm f/2.8 · 50mm f/1.2 · 35mm f/2 · Panavision anamorphic · 400mm telephoto

135mm f/2.8 is the cheat code for skin texture and intimate compression. Anamorphic for cinematic width + horizontal flares.

Named film stocks (one per prompt)

Kodak Portra 400 · Fuji Velvia 50 · Ilford HP5 Plus · CineStill 800T

Portra = natural skin warmth. Velvia = saturated landscape. HP5 = gritty B&W grain. CineStill 800T = tungsten night with halation. Never mix stocks.

Don't carry lens + stock into a video prompt

85mm f/1.4, Portra 400
→ close-up, shallow depth of field, warm natural colors, cinematic texture

Lenses, apertures, film stocks and camera bodies are an image-model lever and a video-model anti-pattern — ByteDance's Seedance guide never mentions f-stops, lens millimetres, fps or shutter angle. When you animate a frame you made here, translate the look into shot size, depth of field and colour tone instead of pasting the gear list across.

Physics-based lighting

Single key light at 45 degrees from upper left. Color temperature 4500K. Crisp catchlights in the eyes.

Direction + Kelvin temp + named source. "Single key light at 10 o'clock" beats "soft lighting" every time.

Imperfection vocabulary

visible pores · peach fuzz · ISO noise · sweat beading · slight hyperpigmentation · unretouched raw photography

Forces the model away from AI-clean skin. The default is too smooth — you have to ask for the imperfections that real photos have.

❌ The anti-list — avoid these

8k · masterpiece · hyperrealistic · ultra-detailed · trending on ArtStation · perfect skin · flawless · airbrushed · cinematic (alone)

Critical. Tag-soup phrases from the Stable-Diffusion era. Measured success ~60-70% with these vs ~95%+ with positive description. Always specify which cinema — director, lens, era, stock.

No negative-prompt field

✅ "empty street" not "no cars"
✅ "without people, vehicles, or signage"
❌ "not anime, not cartoon, not 3D"

Reframe positively first. Use inline "without" / "free of" only when positive framing can't suppress the unwanted element.

Reference images — name them, never label roles

Marcus (image 1) sits across from the woman (image 2) in the cafe (image 3).

Up to 14 refs (10 object + 4 character — caps don't trade). Name each reference inline; never write role essays. Slates does this for you: @mention a subject or environment and it composes Marcus (image 1) in the cafe (image 2), citing them in the exact order it sends them. One canonical identity image avoids competing facial renderings; a "Reference Image Instructions" block drags reference lighting into your scene. Start with 2-3 focused refs.

Common fixes

Hands → "five fingers, natural proportions"
Text → quote-wrap "HEADLINE" + specify font
Left/right → "from the character's perspective"

Default left/right is the viewer's perspective. Surreal prompts trip uncanny valley — the model drags toward realism. For surrealism, lean hard into "painted" / "illustrated".

Resolution tactics

1k only on Lite

Lite outputs 1K only — use it for iteration volume and drafts, then switch to Nano Banana 2 for 2K/4K finals.

💡 Boring vs Cinema. Boring: "Wide shot of man on dock looking at forest." Cinema: "Direct overhead drone shot on weathered dock. Single figure climbing up frame bottom. Boot prints leading toward shore. Pale winter light. Anamorphic flare. Desaturated blue/slate palette. Kodak Portra 400 grain. Map of threat."

💡 3-strike rule. If three iterations on the same prompt haven't landed, stop. The slot machine doesn't converge — the prompt structure is wrong, not the seed.

Audio

Seed Audio 1.0

Seed Audio builds a whole audio scene — dialogue, effects and ambience together — from one plain sentence. Write it the way you would describe the moment to a person standing next to you, not the way you would write a video prompt.

It has no duration setting. Length comes from the words, so Slates appends your chosen duration to the prompt ("… 15 seconds") and bills exactly that. Set the duration control to what you actually want and let the sentence stay clean.

One plain sentence

nature soundscape, wide open field cicadas and birds and a loon.

No shot language, no production jargon, no formatting. Plain description outperforms anything that reads like a spec sheet.

Duration lives in the prompt

tiny applause of 2 or 3 people at an open mic. 15 seconds

Critical. The duration control writes this for you. Do not also type a different length into your sentence — the two will fight and you pay for the one you selected.

Say the crowd size

tiny applause of 2 or 3 people  ·  a packed arena roaring

"Applause" alone returns a full room. Scale words are the single highest-leverage edit on any crowd, traffic or nature bed.

Cut beds longer than the shot

Ask for a few seconds more than the clip needs so the edit has handles to fade in and out of. Beds that end exactly on the cut always sound clipped.

No Kling syntax here

✗ SFX: heavy boots
✓ heavy boots on wet pavement, a siren far off

Critical. The "SFX:" and "Ambient noise:" prefixes belong to Kling video prompts. Seed Audio treats them as words in the scene and the result gets worse.

Dialogue in quotes

a tired bartender says, "we closed twenty minutes ago", glasses clinking behind him

Speech goes in quotes inside the same sentence as the room. Pick a preset voice for a specific speaker, or leave it unset and let the scene cast itself.

References

match the room tone of @Audio1

Up to 3 audio clips (max 30s each), referenced as @Audio1–@Audio3 — OR one image to score what is in frame. Never both in the same generation.

Know its seat

Scenes, beds, room tone and dialogue in one pass. For a single effect that has to land on a specific frame, use Sound Effects.

ElevenLabs Sound Effects

Sound Effects makes one short sound with an exact length — the lane for a hit that has to land on a specific frame, or a seamless loop you can lay under a whole scene.

Duration is always sent explicitly (0.5–22s). Billing is per second, so the length you pick is the price you pay.

Describe the cause, not the label

✗ door sound
✓ heavy oak door slams shut in a stone hallway

Critical. Material, weight and room are what separate a usable effect from a stock-library shrug. Name all three.

One sound per generation

This surface makes a single event. A door, then footsteps, then a siren is three generations layered on the timeline — or one Seed Audio scene.

Duration is the edit

0.8s for an impact · 4s for a whoosh · 22s for a bed

Ask for roughly the length you need. A 4-second request for a door slam pads the tail with room tone you then have to trim.

Loops

steady rain on a canvas tent  (loop on, 12s)

Turn loop on for anything continuous — rain, engine hum, crowd murmur — and it will tile without a seam.

Prompt influence

0.3 default · 0.7 literal

Higher hugs your wording with less variation between takes; lower explores. Raise it when a re-roll keeps wandering off the brief.

Know its seat

One precise effect on a known frame. Full rooms and layered scenes are cheaper and better in one Seed Audio pass.

Going deeper

The tips above are the short version. The full prompting craft ships as the Slates skills library with the community membership: per-model long-form guides, character identity, style prompting, cost discipline, the vision feedback loop and campaign blueprints. See pricing.

One-time purchase · 30-day money-back guarantee