The prompt is already written for you
Our generator does not ask for a prompt, but the scene still has to be described to the model - once, by us, and kept stable for everyone. Here is what is in it.
15s
15-second widescreen clip
16:9 / 720P
Ready for short-form posts
50
Technical failures are refunded
What the built-in prompt covers
It locks the scene first: a dim barn at night filled with loose hay, one lantern casting warm saturated light and deep shadows, a 16mm fairy-tale look with visible grain, soft focus and slight gate weave, a locked-off camera with a slight handheld sway - and the instruction to keep all of it unchanged. The warmth clause matters more than it sounds: without it, models tend to drift toward washed-out, flat lighting.
Then it assigns the cast: the first reference image is the little man in the dark velvet tuxedo who tiptoes forward, stops, raises a finger to his lips and dances with both arms out; the second is the maiden who watches from the hay and reacts. It closes with the wording that keeps identity stable - preserve both faces, hair and outfits from the reference images. One thing it never says: never ask for a face swap or a face replacement. That phrasing gets rejected by content-safety filters on the big video models, and the render simply fails.
The scene line
A 1970s European fairy-tale film still, shot on 16mm: a dim barn at night, loose hay, one lantern, deep shadows, warm saturated colour, visible grain, soft focus, light halation on the lamp - no modern objects and no text.
The motion line
He tiptoes four steps across the hay toward the camera, stops, looks into the lens, and slowly raises one finger to his lips. Camera locked off with a slight handheld sway, with the instruction to keep face, height, hair and suit identical to the reference sheet, and no extra fingers or morphing feet.
The finishing pass
Grade back toward the reference: a little grain, a light vignette, slight jitter, and audio laid back on top. That is the step that makes an AI clip read as a damaged 16mm print rather than a clean render - and it is why the beats matter more than the pixels.
You do not have to write any of this yourself - the generator uses it already.
Generate with the built-in Rumpelstiltskin AI prompt →Create your Rumpelstiltskin video
Photo 1 becomes the little man who tiptoes, shushes and dances. Photo 2 becomes the maiden who waits in the hay and reacts.
For the best result
- One person (or pet) per photo, facing the camera
- Face clear and well lit - no sunglasses, hats or hands over it
- Head to the waist or knees, so the outfit shows
- No mirror selfies - the phone can end up in the video
Choose quality
Every beat of the barn clip - the tiptoe, the shush, the arms-out dance - in one 15-second 16:9 video.
Output
Made from these 2 photos


The clip is silent, so you can lay the trending sound over it in TikTok or Reels. Finished videos land in My videos.
0/2 photos added
Video model: Seedance 2.0. AI-generated.
Rumpelstiltskin AI questions
Ready for your Rumpelstiltskin video?
Two photos in, one 15-second barn video out. Packs start at $6.99, credits never expire, and you only spend them when you press generate.
Make my video