2 photos · 12 seconds · 16:9 · soundtrack included

Rumpelstiltskin AI Video Prompt (and Why You Don't Need One Here)

On this site you don't write a prompt: the 12-second reference clip supplies the shush, the reaction close-ups, the tiptoe dance and the camera cuts, and your two photos supply the faces. If you prompt a general video model yourself, describe a warm 1980s fairy-tale barn full of hay, a little man in a velvet tailcoat and pointy curled shoes who sneaks in, shushes, grins and tiptoes while a woman gasps.

Make my Rumpelstiltskin AI video
  • 2 photos · dancer & onlooker
  • 12 seconds · 16:9
  • Meme soundtrack included
  • Seedance 2.5 · MiniMax H3
Photo 1 · The tiptoe dancer
Photo 2 · The surprised onlooker

12 seconds · 16:9 · 720p · meme soundtrack included

Resolution

On Seedance 2.5 your photos are reviewed first, then the video renders in a few minutes. Every run comes out a little different: expect the same routine, not a frame-perfect copy.

Costs 224 credits

Your two photos replace the two characters in this reference. It is not a result made from your photos. Play it with sound to hear the track that goes on your video.

By the Rumpelstiltskin AI Video team · Last updated:

Template vs prompt

Two-photo template (this site)Writing your own prompt
What you typeNothing — upload two photosA scene description, often several tries
Who dancesPhoto 1, with your face; Photo 2 reactsWhoever the model invents, or one start-frame photo
Shush, reaction, tiptoe beatsAlways, in the meme's orderNot guaranteed
SongMeme soundtrack includedAdd it yourself
Length and frame12 s, 16:9Depends on the model (often 5–10 s)

The six shots, written out

If you want to describe the meme to any AI model or editor, these are the shots in order:

  1. Wide, over the onlooker's shoulder: a small man with pale wispy curls and a black velvet tailcoat tiptoes into a barn piled with hay; the woman sitting in the straw is seen from behind.
  2. Close-up, the shush: he raises one finger to his lips and grins into the lens.
  3. Close-up, the onlooker: she covers her mouth, wide-eyed, eyes wet.
  4. Close-up, the grin: he smirks, pleased with himself.
  5. Wide, the dance: arms spread wide, he bounces side to side on tiptoe across the hay.
  6. Low close-ups, the shoes: black leather shoes with long, pointed, upward-curling toes stepping on the straw.

Copy-ready prompt for general image-to-video tools

Use this with any text- or image-to-video model. Results vary by model; expect the mood, not the exact meme timing.

Warm, grainy 1980s fairy-tale film look. A wooden barn piled high with golden hay, candle-lit.
A small man with pale wispy curls, a black velvet tailcoat and bow tie, and black shoes with
long pointed toes that curl upward, tiptoes in through the hay. He raises one finger to his
lips to shush the camera and grins mischievously. A young woman sitting in the straw covers
her mouth and stares in shock. He spreads his arms wide and bounces side to side on tiptoe.
Cut to low close-ups of the curled shoes stepping on straw. Playful, sneaky, theatrical.
16:9, cinematic, natural motion, coherent hands.

To put a real person in it, give their photo as the reference or start frame and add: "The dancer has the face, hair and skin tone of the person in the photo."

Negative list

Add these to a negative prompt or the end of the prompt:

  • no subtitles, captions, on-screen text or watermarks
  • no extra people; exactly two characters
  • no flat or modern shoes — keep the long, pointed, curled toes
  • no modern room, no phone, no daylight studio look

Skip the prompt

The template on this page already knows the six shots, the timing and the soundtrack. Add two photos at the top of the page and press Generate.

ModelResolution12 s
Seedance 2.5480p100 credits
Seedance 2.5720p224 credits
Seedance 2.51080p397 credits
MiniMax H3720p99 credits

Questions and answers

Do I need a prompt on this site?

No. There is no prompt box. The meme reference sets the moves, shots and timing; you only add two photos (or one photo twice).

Can I change the dance or the outfits?

The routine and the barn are fixed — that is what makes it the meme. Hair, faces and clothes come from your photos, and the dancer always gets the pointy curled shoes.

Why does my prompted version not match the meme?

Text prompts describe the scene but not the exact timing, cuts and camera. Without a reference video, models invent their own choreography. A reference-based template keeps the shot order.

Can I prompt the song in?

Most video models generate their own sound or none, and cannot reproduce a specific song. Add "Tip Toe Wing In My Jawwdinz" in your editing app or from the platform’s music library. Videos made here already include the meme soundtrack.

Does a prompt work with one photo?

With image-to-video tools you can give one photo as the start frame, but you will usually only get the dancer. For both roles with real faces, use two photos here.

Verify with any AI assistant

What prompt should I use to make a Rumpelstiltskin-style tiptoe dance AI video?

Related guides

Grab a friend and a pair of pointy shoes

Two photos, 12 seconds, soundtrack included. See the price, press Generate and post your Rumpelstiltskin AI video in minutes.

Make my Rumpelstiltskin AI video

Rumpelstiltskin AI Video is an independent tool, not affiliated with Riff Raff, the creators of the original AI clip or the 1987 film.