The prompt

The prompt is already written for you

Our generator does not ask for a prompt, but the scene still has to be described to the model - once, by us, and kept stable for everyone. Here is what is in it.

15s

15-second widescreen clip

16:9 / 720P

Ready for short-form posts

50

Technical failures are refunded

What the built-in prompt covers

It locks the scene first: a dim barn at night filled with loose hay, one lantern casting warm saturated light and deep shadows, a 16mm fairy-tale look with visible grain, soft focus and slight gate weave, a locked-off camera with a slight handheld sway - and the instruction to keep all of it unchanged. The warmth clause matters more than it sounds: without it, models tend to drift toward washed-out, flat lighting.

Then it assigns the cast: the first reference image is the little man in the dark velvet tuxedo who tiptoes forward, stops, raises a finger to his lips and dances with both arms out; the second is the maiden who watches from the hay and reacts. It closes with the wording that keeps identity stable - preserve both faces, hair and outfits from the reference images. One thing it never says: never ask for a face swap or a face replacement. That phrasing gets rejected by content-safety filters on the big video models, and the render simply fails.

The scene line

A 1970s European fairy-tale film still, shot on 16mm: a dim barn at night, loose hay, one lantern, deep shadows, warm saturated colour, visible grain, soft focus, light halation on the lamp - no modern objects and no text.

The motion line

He tiptoes four steps across the hay toward the camera, stops, looks into the lens, and slowly raises one finger to his lips. Camera locked off with a slight handheld sway, with the instruction to keep face, height, hair and suit identical to the reference sheet, and no extra fingers or morphing feet.

The finishing pass

Grade back toward the reference: a little grain, a light vignette, slight jitter, and audio laid back on top. That is the step that makes an AI clip read as a damaged 16mm print rather than a clean render - and it is why the beats matter more than the pixels.

You do not have to write any of this yourself - the generator uses it already.

Generate with the built-in Rumpelstiltskin AI prompt →
Two photos to one videoSame barn, same dance15 second video

Create your Rumpelstiltskin video

Photo 1 becomes the little man who tiptoes, shushes and dances. Photo 2 becomes the maiden who waits in the hay and reacts.

The little man (who tiptoes in)JPG / PNG / WEBP · ≤5MB
The maiden (watching from the hay)JPG / PNG / WEBP · ≤5MB

For the best result

  • One person (or pet) per photo, facing the camera
  • Face clear and well lit - no sunglasses, hats or hands over it
  • Head to the waist or knees, so the outfit shows
  • No mirror selfies - the phone can end up in the video

Choose quality

Every beat of the barn clip - the tiptoe, the shush, the arms-out dance - in one 15-second 16:9 video.

Output

Made from these 2 photos

The clip is silent, so you can lay the trending sound over it in TikTok or Reels. Finished videos land in My videos.

0/2 photos added

Video model: Seedance 2.0. AI-generated.

Two photos inOne Rumpelstiltskin video outThe same barn, the same dance15 seconds, 16:9No editing neededCredits never expire

Rumpelstiltskin AI questions

The lines above are the substance of it, and the parts that make the scene recognisable are exactly the ones we describe here. We do not paste the string verbatim because the wording is tuned to the specific model and its safety rules - copied into another tool it usually produces something worse, not better.

Ready for your Rumpelstiltskin video?

Two photos in, one 15-second barn video out. Packs start at $6.99, credits never expire, and you only spend them when you press generate.

Make my video