Text-to-video models do their best work on one moment, not a story. This listing is built around that: it asks for a single scene and turns it into the kind of clip that holds up in a vertical feed.
What is built into every run:
- 9:16 at 768p, five seconds, the length most short-form clips are cut to.
- One continuous shot with one camera move, which is what keeps AI video from looking jittery.
- Action held in the centre, because apps cover the top and bottom of the screen with buttons and captions.
- No text in the picture. Models misspell words, and you will want your own captions anyway.
Write the scene the way you would describe a photograph: subject, place, action, light. For example: a barista pouring latte art in a sunlit cafe, steam rising, slow push-in.
Real, identifiable people and famous characters are not allowed, and the model may refuse or change them.