The prompt doesn't do shit.The workflow does.
Everyone asked for the prompt behind my STORM parody. Here it is, but it only works because of what sits around it: the assets, the direction and the way I load it all into the model. This is the full build, in three steps.
Assets
A video model has no memory. If you describe your hero loosely, the next shot gives you a different face and a different jacket. Every asset is a pair: an image the model anchors to, and the same words describing it in every prompt.
Character sheets
- Three views on one sheet: a large portrait in three-quarter view, a full body from the front and a full body from the back.
- On the front full body, crop the head out. On wide shots the model otherwise grabs the tiny blurry face from the small figure instead of the portrait.
- Keep the sheet boring on purpose: flat neutral grey background, flat light, real skin with pores, no grain, no cinematic grade. The cinema look lives in the video prompt, not in the sheet.
- Check the eyes. Without a small catch-light in the pupil the face looks dead, and a dead face can't act.
The lead, built from my own photos
The one who talks back
The one who sits and rocks
Use the attached photos as the exact identity of the woman: same face, same features, same skin, same proportions. Create a clean character sheet on a flat neutral grey studio background with flat even light, no film grain, no cinematic color grade. On the right: a large portrait of her head and shoulders in a three-quarter view, face turned slightly away from the camera, calm neutral expression, a small catch-light in both eyes, real skin with visible pores and no retouching. On the left: two full-body figures standing neutral with arms relaxed, one seen from the front with the image cropped at the collar so no head is visible, one seen from the back. The outfit is identical in all three views: light blue striped oversized shirt, black knit V-neck vest, black tie, grey plaid pleated mini skirt, burgundy lace-up knee-high boots. Wet-look blonde bob. No text, no labels, no props.
A clean character sheet of a young man in a worn school uniform on a flat neutral grey studio background with flat even light, no film grain, no cinematic color grade. He has a very short dark buzz cut, a narrow face with sharp cheekbones and a slightly tired, unimpressed look. Outfit: faded navy blazer with a small blue crest patch, greyish untucked white shirt, loosened striped navy tie, dark trousers, scuffed black shoes. Layout: a large portrait of his head and shoulders in a three-quarter view with a small catch-light in both eyes and real skin with visible pores; a full body from the front with the image cropped at the collar so no head is visible; a full body from the back; a face profile. The outfit is identical in every view. No text, no labels.
For the second guy I only swapped the hair: a long, messy light-brown fringe falling into his eyes. Same uniform, same layout, so they read as one school.
Locations
- Give the model the place from several angles: the front, the reverse view and one or two views from above. Past the edges of a single frontal picture it invents a new world every time.
- Prefer three-quarter angles over flat frontal ones. They give the model depth to place people correctly.
- Keep one light logic across all angles: one sun, one direction of shadows.
- Leave an anchor you can stage against. Here it's the bleachers: "she stands on the lowest bench" works, "she stands in the yard" is a lottery.
Front: the facade and the bleachers
Reverse: what the crowd sees
From above, three-quarters
From above, the other side
Use Image 1 as the exact location: the same red brick school facade, the same black wooden bleachers, the same banners, stone paving and lamp posts. Show the same place from a new angle: standing on the top bench of the bleachers and looking back across the stone plaza toward the lawn and the trees, the edge of the top bench in the foreground. Keep the same time of day, the same soft late-afternoon light and the same direction of shadows as Image 1. Empty, no people. Photorealistic, natural lens, no film grain, no text.
Props
Every object that matters goes in as its own image, clean on white. If a prop only exists as a word in the prompt, the model redesigns it in every shot, and in a single take it can even change shape halfway through.
Use Image 1 as the exact object. Show this black pistol alone on a seamless pure white background, side view, soft even studio light, sharp focus over the whole object, a soft contact shadow under it. No hands, no text, no logos added, nothing else in the frame.
Motion reference
The dance is the one thing you can't describe in words. I tried breaking the choreography down second by second, and the model averaged it into mush every time. What worked was the full original clip as a video reference, plus one line in the prompt saying it's used only for the crowd's movements and not for its camera, cuts or pacing.
A tip that cost me a lot of takes: don't crop or trim the reference to make it "cleaner". The full clip with the full frontal composition gave the best result.
The motion reference: a frame from the original clip. And yes, I judge you for ripping other people's shots for AI slop.
The prompt
Don't write. Direct. Every second, every camera move, every emotion, exactly how you see it in your head. The prompt below runs about 1,400 words for 30 seconds, and that length is fine. What breaks a generation is an overloaded beat, not a long prompt.
- Write timed beats, each continuing from exactly where the camera and every person were at the end of the previous one.
- Describe the camera as a physical thing moving through space: running, backing away, circling. Shot names like "Medium shot" or "Wide shot" at the start of a beat read as an edit list and invite cuts.
- The next person must already be visible before the camera moves to them. "Meanwhile, off screen" is an invitation to cut.
- Write handheld language, not motion-control language. Words like "precise arcs" and "soft settle" give you a sterile, cartoon-smooth camera.
- Never write a word you don't want in the frame, even as a ban. "No drone visible" puts a drone in the shot.
- Keep acting small and human. "His eyes cross on the barrel" gets you a grimace. A glance, a swallow, a shift of weight gets you a person.
- Name the role of every reference in the prompt, and lock each character with both the slot and a visible detail: "@Image3, the boy with the buzz cut", every single time.
One continuous 30-second single take in real time. Grounded cinematic realism, 35mm film texture, low late-afternoon autumn sun, muted palette of navy, red brick and grey stone. CAMERA BEHAVIOR for the entire take: raw unstabilized handheld, as if an operator is running with the camera: constant micro-shake, footstep bounce, slight horizon tilt, imperfect reframing, brief focus hunting. Every big move (the 180-degree arcs, the FPV flight, the handheld pull-back) keeps this shake on top, with sharp starts, overshoot and rough recovery. Real-time 1x speed from the first to the last frame: no time-lapse, no speed ramp, no slow motion, no frame skipping, no cuts. No rig, drone or crew visible. REFERENCES: @Image1 is LEAD, the main character: her face, wet-look blonde bob, light-blue striped oversized shirt, black knit vest, black tie, grey plaid mini skirt, burgundy lace-up knee boots. @Image2 is SITTER, the boy with the long messy light-brown fringe. @Image3 is TALKER, the boy with the very short dark buzz cut. @Image4 is the school facade with wide black wooden bleacher steps. @Image5, @Image6 and @Image7 show the layout of the same campus plaza, lawn and lamp posts, for spatial continuity only. @Image8 is LEAD's pistol. @Video1 is used only for the crowd's dance movements and for the position of the young man in the white shirt standing in the middle of the crowd; do not take its camera, cuts or pacing. That white-shirt young man does not exist in this video: LEAD takes his exact place. @Audio1 is LEAD's voice and contains all her lines in order; use it for her timbre, delivery and exact words. CROWD: about forty young men, a compact tight cluster shoulder to shoulder on the central part of the bleachers. Different faces, builds and hair, none resembling @Image2 or @Image3, all in the same worn uniform as @Image2 and @Image3: faded navy blazers with a blue crest patch, greyish untucked white shirts, loosened striped navy ties, scuffed black shoes. Every boy wears his blazer. Rumpled, exhausted, worn down by life. 0-2s: Medium shot, LEAD faces camera, cigarette already lit between her lips. She turns her back to camera toward the bleacher steps of @Image4, where the cluster stood in formation; the cluster is already bursting apart in panic, boys scattering a short distance left and right, stumbling down the steps and colliding. The camera whips a fast, rough 180-degree arc around her, overshoots and jerks back into a close-up of her face with the chaos streaming past behind her. 2-4s: LEAD stares at the scattering boys and takes one slow, deep drag, the ember glowing brighter, cheeks slightly hollowing, eyes darting across the runners in rapid blink bursts. Then she pulls the cigarette from her lips between her right index and middle fingers and holds it beside her face, smoke spilling out of her mouth with the words. LEAD, @Audio1: "Wait... what the fuck?" As the camera launches away, she flicks the cigarette aside with a snap of her fingers, out of frame. 4-7s: The camera launches off her shoulder into a rough, jittery low FPV flight through the fleeing boys, weaving between running bodies on the steps, and catches SITTER crawling on all fours between their legs to the far edge of the lowest bench by the metal railing. He drops onto his butt with his back against the railing, arms locked around his knees, rocking, fingers digging into his shins. The camera brakes into a shaky handheld close-up of his face: a thousand-yard stare with sudden blink bursts, lips moving before any sound comes, as if he has stood in this shot a thousand times. 7-11s: SITTER says: "I can't stand there anymore... I just can't." Voice: "An 18-year-old English boarding-school boy with a soft London accent. Thin frayed tenor cracking into a hoarse whisper; slow broken delivery with shaky breaths between words; hollow and unhinged, jumping higher and breaking on the last word." While he speaks, the handheld camera keeps walking backward away from him, shaking, widening from his close-up into a wide shot of the bleachers, SITTER small at the railing in the foreground. In the middle ground of that wide shot, TALKER runs down the steps trying to escape; LEAD steps into his path and simply grabs the back of his collar with her right fist mid-stride, a plain physical grab: his feet skid on the wooden bench, his momentum yanks him half around, he nearly falls. No impact effect, no sudden camera push, just ordinary movement at real speed. 11-19s: The handheld camera hurries forward into a medium three-quarter profile, bouncing as it climbs the steps alongside LEAD while she drags TALKER back up the bleachers by the collar with her right fist, the @Image8 pistol in her left hand. She sweeps the pistol in a slow arc across the boys; they freeze in a staggered wave, hands half raised. Behind and around them, terrified boys crawl back up the bleachers on hands and knees, clambering over the wooden benches, pulling each other up by the sleeves, scrambling toward their old spots and glancing back at the gun, some already standing stiffly in rows. She jabs the pistol toward TALKER, breathless from climbing, furious and dead serious, as if her whole career depends on this trend. LEAD, @Audio1: "You are all motherfuckers supposed to stand still. It's a trend!" TALKER stumbles, clutching her wrist; his eyes flick to the gun once, then lock straight onto LEAD's eyes and stay there for his whole line, steady and cheeky, a smug half-smile creeping back despite the fear. TALKER, looking her in the eye, says: "Don't you think ripping off someone else's shot for an AI trend is kinda messed up?" Voice: "An 18-year-old posh English public-school boy. Light nasal tenor, breathless from being dragged; clipped quick delivery, scared but cheeky, words tumbling faster as fear rises, lifting up at the end of the question." LEAD meets his stare for one cold beat without slowing down. 19-21s: Still unbroken, the camera swings a rough handheld 180-degree arc around them into a frontal view of the @Image4 steps. LEAD shoves TALKER down into the row directly in front of her center spot, slightly to her right; he stumbles into place there, straightens his blazer and stays there, clearly visible in the front of the cluster. The last boys squeeze back in around him, SITTER among them. LEAD climbs into the middle of the cluster and stands in the exact spot of the white-shirt young man from @Video1: center of the crowd, a few benches up, boys packed around her on every side and in the rows below her. 21-24s: LEAD stands there inside the crowd, upright and still, chin up. She raises the pistol straight up in her left hand. LEAD, @Audio1: "I don't give a fuck." She fires once into the sky, a sharp crack echoing off the brick, the boys around her flinching in a wave, TALKER flinching hardest. LEAD, screaming, @Audio1: "I want trend!" 24-30s: As she lowers the smoking pistol to her hip, the boys around her perform the dance movements of @Video1 in silence, mouths closed: they fold forward into a deep bow, torsos parallel to the ground, heads and arms hanging loose; snap upright with heads thrown back to the sky, chests open, hands clawing at their lapels and ties; fold down again and whip back up in a rolling storm wave across the rows, heads rolling, blazers flapping. Tired and rumpled, but at full force. TALKER dances with everyone in his spot just below and to the right of LEAD, fully in sync, but steals a quick sideways glance up at her between moves. SITTER dances with the same empty stare, a fraction late. LEAD stays completely still and upright in the middle of the crowd, exactly as the white-shirt young man does in @Video1, the dancers folding and rising all around her, pistol hanging at her hip. The handheld camera drifts slowly backward, shaking, until the whole compact crowd fills the frame frontally below the facade, LEAD at its center, TALKER right below her. SOUND: no music, no chanting; the only voices are the spoken lines above. Boots and shoes on wooden steps, panicked shouts during the scatter, LEAD's sharp inhale on the drag and smoky exhale on her line, heavy breathing, FPV wind whoosh, the operator's footsteps backing away, TALKER's shoes skidding on wood and fabric tearing taut as she grabs his collar, one gunshot with brick echo, then synchronized stomps on wood and fabric snapping with every dance accent.
This version was written for my slot order. If your files sit in different slots, change the numbers in the prompt, not the files.
Generation
Open Higgsfield, go to Video, pick Seedance 2.5 and drop in all your references: images, the video and the audio. Paste the prompt and generate.
| Slot | File | Role |
|---|---|---|
@Image1 | Lead character sheet | Face, hair and outfit of the lead |
@Image2 | Sitter sheet | The guy with the long fringe |
@Image3 | Talker sheet | The guy with the buzz cut |
@Image4 | Front of the location | Facade and bleachers |
@Image5–7 | Other angles | Layout of the place only |
@Image8 | Pistol | The prop |
@Video1 | Original clip | Crowd movement only |
@Audio1 | Lead's lines | Voice, timing, exact words |
- Generate at 480p while you iterate. It's faster and cheaper.
- When a take works, upscale that take. Re-generating at a higher resolution doesn't enlarge your good version, it makes a brand new one.
- One continuous take of up to 30 seconds. Longer stories get split into takes that start where the last one ended.