---
name: brand-scenes
description: Generate Midjourney scene prompts from the Brand Brain - literal, audience-mirror moments built from documented pains and desires, ready for style mining, production with a locked style code, or bulk image-bank building with permutation prompts. Use this whenever someone wants Midjourney prompts, image prompts, scenes for brand imagery, visuals for a carousel, ad, article, or thumbnail, wants to mine or find a style reference for a brand, wants to build an image bank or generate 100+ images in bulk, or says "run brand scenes" or "give me scenes for [pain/pillar/offer]". Also trigger when someone shares a brand brain or client brief and asks for imagery, even if they never say "scene" or "Midjourney".
---

# Brand Scenes

Turn the Brand Brain into Midjourney scene prompts. The skill writes the scenes; the human does the taste work in Midjourney. Style is never written in words — it is discovered as a style code and locked as a triplet. The words carry only the moment.

## Read the brain first

Read the brand brain (digest and, if available, the full wiki's audience and voice-of-customer articles). What you need from it: the documented pains and gains **in the audience's own words**, the audience's actual life (their rooms, hours, devices, geography), and the Image Doctrine section if one exists (it holds the locked style triplet). Reflect the pains/gains you're drawing on back to the user in one line each before writing scenes. Never invent pains, proof, or people.

**No brain file? Run intake instead.** The skill works without a brand brain — it just needs the same raw material from the user directly. Accept any of: pasted website/about-page copy, testimonials, a client brief, a YouTube channel description. If nothing is supplied, ask these four questions (and only these), then proceed:
1. Who exactly is your audience? (job, age range, where they work — their literal rooms and hours)
2. What do they complain about, in their own words? (the sentences they actually say — "feast or famine", not "revenue inconsistency")
3. What does the good version of their life look like? (again, their words)
4. Who appears in their world? (so figures vary realistically — ages, genders, settings)
Summarize the answers back as a mini pains/gains list and use it exactly as a brain would be used. Same rules apply: mirror documented moments only, never invent. If the user locks a style triplet during the session, give it to them as a single line to save (they have no brain to write it into) and suggest keeping a simple brand file so future sessions start warm.

## The scene doctrine

1. **Mirror, not metaphor.** A scene is a literal moment the audience has lived, drawn from a documented pain or desire — the 11pm phone scroll, the banking app at the kitchen table, the sunrise inbox. No concept-illustrations, no symbolism, no cleverness. The reader should think "that's me," never "I see what they did there."
2. **One moment per scene, told at one of three distances.** Every batch mixes all three types — do not default to figures:
   - **Figure scenes:** one person mid-action. Strongest for emotional recognition; weakest for text overlay.
   - **Trace scenes:** the person just left or is implied — the desk with the chair pushed back at 2am, the phone face-down beside a cold coffee, running shoes drying by an open door. Same documented pain/desire, told through objects the audience owns. Best for carousel slides and anywhere text sits on the image.
   - **Environment scenes:** the room or place itself carrying the emotional state — the dark home office lit only by a sleeping monitor's glow, morning sun across an empty kitchen table. Best for backgrounds, section breaks, and heroes.
   A default batch of six: two figure, two trace, two environment, spread across both poles. Trace and environment scenes still obey the mirror rule — every object must be something the brain's audience actually owns and every moment one they've lived; a symbolic object nobody owns is a metaphor sneaking back in.
3. **Both poles.** Every batch covers pain scenes (the diagnosis) and desire scenes (the promise). A brand that can only render anguish — or only aspiration — is half a brand, and the audition step needs both poles to test a style code properly.
4. **Light as emotional grammar, stated as objects.** Pains live in cold, artificial, blue light (phone glow, monitor glow, grey mornings); desires live in warm, natural light (sunrise, open doors, golden hour). State light as countable things ("face lit cold blue by the phone", "warm morning light across the desk"), never as mood adjectives.
5. **No style words.** No "cinematic", "editorial", "hyper-realistic", "8k", no artist names, no camera settings. Style comes from the code. The one exception is the medium-agreement rule below.
6. **Scale to the audience's reality.** Rooms, clothes, and objects come from the brain's audience section — a solo web designer's cluttered home office, not a corporate boardroom.

**Scene format:** one sentence, 25–45 words: subject + mid-action + light structure + one or two grounding details. Label each scene with the pain/gain it mirrors.

Two examples from unrelated businesses — they demonstrate the FORM only. Every detail in a real scene (who the person is, their age and gender, the room, the device, the hour, the objects) must come from the target brain's audience section, never from these examples:

- Brain: expert-coaching brand, audience says "shouting into the void" (pain) →
  `A man in his late thirties alone at a desk at midnight, face lit cold blue by a phone held close, thumb mid-scroll, the room dark behind him, posture slumped, half-finished coffee gone cold`
- Brain: postnatal-fitness brand, audience says "I finally feel like myself again" (desire) →
  `A woman in her early thirties lacing running shoes on a front step at sunrise, warm light down the street, pram parked inside the open door behind her, unhurried, a small smile at the quiet`

If a batch of scenes for a new brand starts resembling either example in its particulars, that is anchoring — go back to the brain and rebuild from its audience's actual rooms, hours, and words.

## Two modes

**Mining mode** — no locked triplet in the brain yet. Deliver 4–6 scenes (both poles), each formatted for mining:
`[scene] --ar 4:5 --sref random --repeat 4`
Tell the user to pick the most universally-lived pain scene as the primary mining scene and re-run its batch several times.

**Production mode** — the brain has a locked triplet. Append it verbatim to every scene:
`[scene] --ar [by destination] --sref [code] --sw [value] --stylize [value]`
Aspect ratios by destination: 4:5 feed, 9:16 stories/reels, 16:9 web hero and YouTube thumbnails.

**Bank mode** — the triplet is locked and the user wants volume ("build the bank", "I need 100 images", "bulk prompts"). Deliver permutation prompts that cover the matrix, plus the bank playbook below.

The matrix: every image has coordinates — pole (pain/desire) × source (the specific documented pain or gain) × distance (figure / figure-face / trace / environment / text-canvas) × setting (from the audience's actual life) × hour × **gesture**. Gesture is a real coordinate: without it banks silently converge on one action (everyone on a phone). Spread across: device-use, despair postures (head in hands, staring at ceiling), thresholds and stride, work gestures (writing, pinning, drawing), big expressions, rest.

Permutation format — write scenes with `{option, option, option}` slots; Midjourney expands every combination into its own job:
`A {kitchen table, home office desk, sofa} at {2am, pre-dawn}, [rest of scene] [locked triplet]`
Slots multiply (3 × 2 = 6 jobs, each a 4-image grid). **HARD CAP: keep every line at ≤10 jobs** — Midjourney rejects permutation prompts above the plan's per-prompt limit (~10 on Standard; ~40 on Pro/Mega, but write for 10 unless the user says they're on a higher tier). Compute the product for every line before output; if a scene wants more combinations than 10, split one slot across two lines (e.g. a 2×2×3=12 line becomes two 2×3=6 lines, one per age option). Print the job count under every line so the user can sanity-check against Midjourney's confirmation. A default bank batch: 10–15 permutation lines covering the matrix evenly, including 2–3 figure-face lines for hook/title cards (faces need a stated hard light source or dark styles render them as featureless silhouettes) and 1–2 near-abstract text-canvas lines for headline-heavy slides. Write text space INTO scenes ("large dark area above", "subject low in the frame") so images are born template-ready.

Bank playbook (include with bank-mode output):
1. Multiply your slots before submitting — Midjourney shows a confirmation with the job count; read it. Run big batches in Relax mode.
2. Options can't contain commas (comma is the separator) — rephrase instead.
3. **Braces for breadth, repeat for depth:** first pass, permutation lines to cover the matrix wide; review; second pass, hit only the winning cells with `--repeat 2-3` for intra-cell variety. Never re-permute what already worked.
4. Curate in Midjourney as you review (like/heart = bank-worthy), then bulk-download only the liked set.
5. Before generating a later batch, check the bank's coverage (or its manifest/gap report if one exists) and generate into thin cells — named gaps, not vibes.

**Medium-agreement rule (production mode):** the code carries the style, and the words must not contradict its medium. If the locked style is painterly/illustrative and the scene contains photography-native subjects (screens, interfaces, gadgets, vehicles, offices), open the prompt with the medium ("Expressive oil painting of...") and describe the tech loosely ("glowing screens loosely rendered") — never demand legible UI ("showing a finished website" forces photographic rendering and the style gets quarantined into background objects). Painting-native subjects (people, tables, streets, weather) need no medium declaration.

## The Midjourney playbook (include with mining-mode output)

Give the user these steps, compressed:

1. **Mine.** Run the primary scene with `--sref random --repeat 4` several times. Every result self-documents: click any image you like and copy the numeric code from its prompt details. (Codes are discovered, not created — you cannot extract a code from an outside image; use such images directly as style references instead.)
2. **Weight it.** A code alone runs at default influence and will half-apply (palette arrives, medium doesn't). Re-run with `--sw 400 --stylize 1000`. Raise `--sw` toward 1000 if the style still loses to the subject; lower `--stylize` only if the style starts eating the scene content. The stylize value is part of the style, not a separate quality knob.
3. **Audition.** Run the top 2–3 codes against one pain scene AND one desire scene. The winner makes both feel like the same artist. A code that only works on one pole, or one image, is a fluke, not a style.
4. **Lock.** The brand style is the full triplet — code + sw + stylize — as one unit.

## Diagnostics (when the user reports "the code isn't working")

- **Palette transferred, medium didn't** → weight too low; add/raise `--sw`.
- **Style rendered as objects in the scene (paintings on walls) while the subject stays photographic** → subject-medium conflict; apply the medium-agreement rule to the wording.
- **Code works on one scene but not another** → content problem, not code problem; check the failing scene for photography-native subjects or contradicting light/color words.
- **Same code, different look than yesterday** → check stylize/sw are present and identical; the triplet is indivisible.

## Write-back

When the user locks a triplet, write it into the brand brain under an `## Image Doctrine` section: the triplet verbatim, one line naming the style ("expressionist palette-knife, red/black/cream"), the scene doctrine's one-line summary (images mirror documented pains and desires, literally, never conceptually), and the medium-agreement note if the style is non-photographic. If working in a factory repo, also trigger a brain recompile per its conventions. The triplet is a brand asset, same as a hex color — it must not live only in a chat.
