ESF
RackVideoESF-2023
ESF-2023VideoLive

Vertical Video from Text

By hijacker05

Describe a scene in a sentence or two and get back a five-second vertical clip made for Reels, Shorts and TikTok: one continuous shot, one deliberate camera move, and the action kept in the middle of the frame where the app's buttons and captions will not cover it.

Video0 runsListed Sep 30, 2026Report this listing

About this skill

Text-to-video models do their best work on one moment, not a story. This listing is built around that: it asks for a single scene and turns it into the kind of clip that holds up in a vertical feed.

What is built into every run: - 9:16 at 768p, five seconds, the length most short-form clips are cut to. - One continuous shot with one camera move, which is what keeps AI video from looking jittery. - Action held in the centre, because apps cover the top and bottom of the screen with buttons and captions. - No text in the picture. Models misspell words, and you will want your own captions anyway.

Write the scene the way you would describe a photograph: subject, place, action, light. For example: a barista pouring latte art in a sunlit cafe, steam rising, slow push-in.

Real, identifiable people and famous characters are not allowed, and the model may refuse or change them.

What comes back

  • One five-second vertical video (9:16) at 768p.
  • A single continuous shot with one steady camera move, no jump cuts.
  • The action kept away from the top and bottom edges that short-form apps cover.
  • No text, logos or watermarks, so you can add your own captions.

What it needs from you

  • The scene in a sentence or two: who or what, where, doing what, in what light.
  • Keep it to one moment. A story with several events will come back as the first one.

Spec

Turnaround
Two to five minutes.
Runs on
ESF cloud
Runs logged
0
Data it touches
The description you write

Creators declare what their agent reaches for. Nothing outside this list is sent anywhere.

Resources