How we made a product demo animation in three prompts
A Fiverr animator stopped replying, so we built the mockup ourselves. One prompt turns Apple, Solana and Revolut ads into a style document with real numbers, one turns that document into a self-contained HTML spot, one slows it down. The exact prompts, and where the method stops.
The animator I found on Fiverr stopped replying. Not "no thanks", just silence, so instead of waiting I built the mockup myself, in an afternoon, out of three prompts and one HTML file.
The result is not a finished ad. It is roughly 70 percent of one: the rhythm is right, the copy is right, the colours are ours, and every beat lands where I want it to land. The remaining 30 percent is what you hire an animator for, and now I can hand them a running artifact instead of a paragraph of adjectives.
This is the method, with the actual prompts.

Why "make me a nice animation" produces nothing usable
Ask a model for a product animation with no further constraint and you get a generic result, because the request contains no information. "Modern", "clean", "dynamic" and "premium" are not specifications. They describe the feeling you want the viewer to have, not the thing you want built.
An animation is a stack of decisions that are all numbers: how long a card stays on screen, what curve the easing follows, whether text fades or cuts, how much silence sits before the logo. If you do not supply those numbers, the model invents them, and it invents them at the average of everything it has seen. Average motion graphics look like a template, because a template is exactly what an average is.
The trick is not to be a better prompter. The trick is to have the numbers. And if you do not have them, the first prompt is for getting them.
Prompt 1: borrow the taste of brands whose ads you already like
You almost certainly cannot describe the style you want. You can, however, name three companies whose advertising you would happily steal from. That is enough, because the work of turning "I like Apple ads" into "500 ms cubic ease-out, 1.2 s minimum hold, no dissolves" is exactly the kind of work a model is good at.
So the first prompt does not ask for an animation at all. It asks for a document.
Write a detailed markdown document describing the advertising grammar of Apple, Solana and Revolut. For each brand cover: the feeling, the image, the pace in seconds per shot, the transitions, how text enters and how it leaves the frame, the sound, the endframe, and the dominant form of the message. Finish with two tables: a one-page comparison of all three brands, and a table of transferable numbers I could hand to a motion designer. Cite sources for anything factual.
Pick brands that disagree with each other. Apple, Solana and Revolut were not chosen because I like all three equally. They were chosen because they solve the same problem in three incompatible ways, and the contrast is what makes the document useful. Three brands from the same category would have produced one blurred average, which is the thing we were trying to avoid.
The two closing tables are the point of the whole prompt. Everything above them is context; the tables are what the second prompt actually consumes.
What a usable style document looks like
The document that came back runs about 2 000 words. The part that mattered was 30 lines long. Here is the comparison table, abridged to the rows that changed decisions:
| Apple | Solana | Revolut | |
|---|---|---|---|
| Role of text | a title naming what you already see | the primary carrier of meaning, the rhythm of the edit | caption to the joke, plus the CTA |
| Text in | fade + scale 96→100%, 500 ms, ease-out | cut-in, 0 ms, glitch or typewriter | snap/spring, 150 ms, word by word |
| Text out | fade 250 ms, or knocked out by an object | hard cut or smear | cut on the beat |
| Pace | slow → fast → stop and silence | uniformly aggressive, build → drop | continuous flow, or very fast in performance cuts |
| Transitions | hard cut only, transformation happens inside the shot | glitch, whip, flash: the transition is the effect | morph or portal, invisible but continuous |
| Ending | product + logo, no CTA | logo slam + slogan + URL | line + logo + CTA |
| Failure mode | cold, does not convert | cringe, backlash | the joke buries the product |
That last row is worth the price of admission on its own. Knowing how a style fails tells you which parts of it you are allowed to copy.
And the second table, which is the one the animation was actually built from:
| Element | Spec |
|---|---|
| Title card in | fade + scale 96→100%, 500 ms, cubic ease-out |
| Title card hold | minimum 1.2 s, absolutely static |
| Title card out | fade 250 ms, or object wipe |
| Word slam section | one word per beat, hard cut in and out, 0 ms transition |
| Shot length, narrative | 0.8–1.5 s |
| Shot length, feature reel | 0.4–0.6 s |
| Camera | one axis per shot, exponential easing, never linear, never handheld |
| Pre-endframe silence | 1.5 s of empty frame before the logo |
| Endframe hold | 2.5–3 s |
Nothing there is a matter of opinion. Every row is a number a machine can execute, which is why the second prompt works at all.
One editorial decision got made at this stage and it is the only strategic one in the whole process: Apple never gives you a CTA, Revolut always does. For a developer tool the answer is obvious. Take Revolut's ending and build it with Apple's restraint: flat background, silence, long hold, one button.
Prompt 2: new session, your material, one instruction
Open a new session for this. The style document is a file now, so the entire research conversation that produced it is dead weight you would otherwise carry through every subsequent turn. Attaching the finished document costs a fraction of what re-reading its own derivation costs.
Read the attached ad-language-teardown.md. Build a self-contained HTML animation for bundle.social that follows it. Product: https://bundle.social Sitemap: https://bundle.social/sitemap.xml Brand colours: #c71343 pink, #45207b purple, #eb6700 orange Wordmark: bundle.social, with the dot in orange 16:9, one file, no external assets, plays on load, loops. Take the transferable-numbers table literally.
Three things in there are doing real work:
The sitemap. It is the cheapest possible way to hand over what your product does and what you call the things it does. The copy in the final spot ("Publish", "Schedule", "Analyze", "Webhooks", "Unlimited social accounts") is our own site's vocabulary, not invented marketing language, and I did not have to type any of it.
Hex codes, not colour names. "Purple" gets you somebody else's purple. #45207b gets you ours, and everything else in the file gets derived from it: the stage background in the final version is #140b26, which is our purple pulled down toward black rather than a generic dark grey.
"No external assets." This is what keeps the whole thing portable. One file, no CDN, no font download, no build step. You double-click it and it plays, on any machine, forever.
Prompt 3: say "25 percent slower" before you have to

The first pass will be too fast. Not sometimes, essentially always, and the reason is structural rather than a failure of the model: it has read the spec, it knows every beat is meant to be short, and it has no idea what its own output feels like at speed. Nothing in a text file tells you that six words in four seconds reads as a flicker.
So put it in the prompt before you see the result:
Everything is 25 percent slower than you think it should be. Long holds are the thing that makes it look expensive.
The rest of the third prompt is corrections, and this is where the "three prompts" framing is honest about being a shape rather than a count. It was one prompt in structure and several rounds in practice. What we changed:
| Correction | Why |
|---|---|
| Slow the whole thing down, one shared beat constant | So tempo is a single number to tune, not 40 scattered durations |
| Two whip lengths instead of one | The transition into the logo needs to be slower than the ones inside a section |
| Hold the endframe far longer | It was leaving before you had read the button |
| Clear the code scene before it becomes visible | On loop, the previous run's payload flashed for a frame |
| Kill the interpolation of copy | The model kept improving my product claims. Product truth is not its call |
That fourth row is the kind of thing you only find by watching it loop twenty times, and it is a one-line fix once you have. That is the actual argument for this format: the bug was visible, the cause was legible, and the fix was three lines moved above one other line.
What the 35 seconds actually does
The finished file runs about 35 seconds and is built from six scenes. The interesting thing is that each one is a different brand's grammar, used for the job that brand is best at:
| # | Scene | Borrowed from | What happens |
|---|---|---|---|
| 1 | Word slam | Solana | "Fifteen platforms." / "Fifteen integrations?" / "No." / "One request." then two full-frame cards: UNLIMITED SOCIAL ACCOUNTS |
| 2 | Feature beats | Revolut spring, Apple hold | Publish, Schedule, Analyze, Webhooks, each springing up over a still uppercase subtitle |
| 3 | The breath | Apple | Wordmark alone, 3.6 s slow zoom, the orange dot pops in at 0.9 s. No claim, no text, nothing to do |
| 4 | The hero shot | Revolut's portal, made literal | A terminal types a real POST /api/v1/post payload, 201 Created lands, then Instagram, TikTok, X and LinkedIn tick in one at a time, then +11 more platforms |
| 5 | Counters | Solana pace, Apple typography | ∞ social accounts / 15+ platforms / 1 API to maintain, counting up |
| 6 | Endframe | Revolut structure, Apple restraint | "One API. Every platform.", the button, the wordmark, a long hold |
Scene 4 is the one that justifies the exercise. The teardown describes Revolut's best trick as the phone screen scaling up to become reality, and the developer translation of that is not a screen recording of a dashboard: it is the payload typing itself and the platforms confirming one by one. That idea came out of the document, not out of me, and I would not have got there by asking for "something showing the API".
Scene 3 is the one everybody wants to cut, and it is the one that makes the rest look deliberate. Three and a half seconds of nothing but a wordmark is uncomfortable to leave in when you are the person who made it. Leave it in.
Why HTML and not a video model
The obvious alternative is to describe the ad to a video model and take what comes out. For this job HTML wins on four counts, and none of them are about visual quality:
| HTML + CSS | Generative video | |
|---|---|---|
| Editing one thing | Change one number, everything else is untouched | Re-roll, get a different everything |
| Your exact brand colours | #c71343, guaranteed, forever | Approximately, if you are lucky |
| Legible text | It is text | Historically the weakest part, and product copy is the whole ad |
| Handing it to an animator | A readable spec with real timings | A reference clip they cannot open |
Determinism is the whole argument. A product spot is 80 percent typography and 20 percent motion, and typography is the thing a browser has been optimised for since 1994. The failure mode of this approach is that you cannot have anything that is not made of rectangles, text and gradients. No 3D, no footage, no texture, no wool-knitted dreamscape. Which is fine, because the style document says Apple's mode A is a locked-off camera on a flat background, and that is a spec you can meet with a div.
Getting it out as a file
Two routes, and they trade off in the direction you would expect:
OBS. Window capture on the browser, record, trim. Takes five minutes, no setup beyond installing OBS, and it captures exactly what you see. The catch is that it is a real-time screen capture: a dropped frame during recording is in the file permanently.
Headless render. Ask for a script that drives the page frame by frame and stitches the output into an MP4. It costs more setup and it gives you a clean, frame-accurate render at whatever resolution you name.
Either way, remember that the browser chrome is not part of your ad. The file we ended up with hides the playback controls and drops the rounded corners on the stage for exactly this reason, so the capture is edge to edge.
There is no sound in any of this, which is the single biggest gap between this and a real spot. The teardown is blunt about it: at Apple the sound design is the star, at Solana the music carries the drop. A silent version of either is a mood board with timing. Budget for audio separately, or accept that the file is for autoplay-muted feeds and nothing else.
What this is worth, and what it is not

Honest accounting:
| Freelance animator | Three prompts | |
|---|---|---|
| Time to a first version | Days, and the clock starts when they reply | About 40 minutes |
| Iteration on a note | A new round, a new wait | Immediate, and you do it yourself |
| Sound | Yes | No |
| Anything not made of text and rectangles | Yes | No |
| Ceiling | High | Around 70 percent |
| What you end up owning | A video file | A video file and a spec |
That last row is why I would now do this even when the animator answers on time. Walking into a brief with a running 35-second artifact changes the conversation from "make me something modern" to "this, but the transitions carry weight and scene 4 needs real depth". You have moved the argument from taste to craft, which is where you actually want a professional's opinion.
Where the method genuinely stops: anything with people in it, anything with real footage, anything where the idea is a visual metaphor rather than a rhythm. Revolut's wool world is not reachable from a browser and no prompt changes that. What is reachable is discipline, and discipline is most of what separates a cheap ad from an expensive-looking one.
One asset is not one asset
The part that surprised me after the file was done: the animation was the short half of the job.
A 16:9 export is one platform's format. Reels, TikTok and Shorts want 9:16. A LinkedIn feed video and an X post have different duration ceilings, different maximum file sizes and different codec tolerances, and the aspect ratio that reads perfectly on a laptop crops the endframe button off a phone. Because the source is HTML, re-cutting for vertical is a change to one CSS rule rather than a new render, which turned out to be a bigger practical advantage than anything to do with the animation itself.
What each platform will actually accept is documented, inconsistently, in fifteen different places. We keep the consolidated version here: social media API media requirements covers the format, duration and size ceilings per platform, and the social media posting API publishes the finished file to all of them from one payload, which is the same call scene 4 spends four seconds typing.
bundle.social
You made one video. Fifteen platforms want it in their own format, through one API.
One contract-first API to publish, schedule and measure across 15 platforms.
Frequently asked questions
Can Claude actually export the animation as an MP4?
Not directly, but it can write the thing that does: a script that steps the page frame by frame in a headless browser and stitches the frames with ffmpeg. The fast alternative is an OBS window capture, which takes five minutes and is good enough for a mockup. Real-time capture can drop frames; the headless route cannot.
Why three separate prompts instead of one long one?
Because the first prompt's output is a reusable file and its conversation is not. Researching three brands' advertising grammar produces a lot of context you never need again, and carrying it into the build session means paying for it on every turn. Starting fresh with the finished document attached is cheaper and produces a better animation, because the model is reading a clean spec instead of the argument that led to it.
Is it really three prompts?
Three in shape. The third one repeated: the first pass is too fast, the loop has a visible seam, the copy drifts away from what your product does. Saying "25 percent slower than you think, long holds are what make it look expensive" in the second prompt removes the most predictable of those rounds before it happens.
Do I need to know CSS to change it?
To change the copy, colours and tempo, no: those are constants at the top of the file and you can ask for them to be grouped there. To change what a scene does, either you know CSS or you describe the change and let the model make it, which works because a single-file animation with named scenes is small enough to reason about in full.
Will this replace hiring an animator?
No, and it is a bad idea to pretend otherwise. It gets you to roughly 70 percent: rhythm, copy, colour and structure. Sound design, real depth, anything with footage or people, and the last pass of polish are what the remaining budget buys. What changes is that the animator starts from a running reference with real timings instead of a mood board.