The template

One template. Your cast.

There is exactly one Rumpelstiltskin AI template, and that is the point: a fixed scene is what makes the identity of both performers stable and the price predictable.

15s

15-second widescreen clip

16:9 / 720P

Ready for short-form posts

50

Technical failures are refunded

Why the template never changes

Every render uses the same fixed scene setup: the same barn, the same single lantern, the same loose hay, the same 16mm grain, the same locked-off camera. The two photo slots decide who plays the little man and who plays the woman, and swapping either photo swaps that performer - nothing else in the frame changes.

The upside is consistency: a series of clips made weeks apart still cut together, because the set and the timing are identical. The trade-off is that you cannot ask for a different barn, a different dance or a different camera angle - we kept the template narrow so the two-photo workflow could be reliable.

Friend pairs

The most common pairing: one friend in the little-man role doing the tiptoe and the shush, one watching from the hay. Works with any two front-facing portraits, including very different heights and builds - each keeps their own proportions.

Couples and family

Partners, siblings, parents and grandparents. Because each photo is a separate reference, nobody gets averaged into a generic face, which is what usually goes wrong when two people are described in a single text prompt.

Pets and characters

Two pets, a pet and a person, or illustrated characters. Fur markings, colour patterns and stylisation survive because they come from the photos rather than from the prompt.

The template is already selected - there is nothing to pick. Upload two photos here.

Open the Rumpelstiltskin AI template →
Two photos to one videoSame barn, same dance15 second video

Create your Rumpelstiltskin video

Photo 1 becomes the little man who tiptoes, shushes and dances. Photo 2 becomes the maiden who waits in the hay and reacts.

The little man (who tiptoes in)JPG / PNG / WEBP · ≤5MB
The maiden (watching from the hay)JPG / PNG / WEBP · ≤5MB

For the best result

  • One person (or pet) per photo, facing the camera
  • Face clear and well lit - no sunglasses, hats or hands over it
  • Head to the waist or knees, so the outfit shows
  • No mirror selfies - the phone can end up in the video

Choose quality

Every beat of the barn clip - the tiptoe, the shush, the arms-out dance - in one 15-second 16:9 video.

Output

Made from these 2 photos

The clip is silent, so you can lay the trending sound over it in TikTok or Reels. Finished videos land in My videos.

0/2 photos added

Video model: Seedance 2.0. AI-generated.

Two photos inOne Rumpelstiltskin video outThe same barn, the same dance15 seconds, 16:9No editing neededCredits never expire

Rumpelstiltskin AI questions

No. The scene and the motion come from one fixed setup, and that constraint is what keeps faces and outfits stable and the cost per video predictable. If you want a different scene, a general-purpose video model with a written prompt is the right tool - it just will not look like this.

Ready for your Rumpelstiltskin video?

Two photos in, one 15-second barn video out. Packs start at $6.99, credits never expire, and you only spend them when you press generate.

Make my video