What is a shot list?
The plan of a video, shot by shot: what each shot shows, how it is framed, what is said and how long it runs. For a Reel with an AI persona, each row becomes one image and one video clip.
Make a shot list for a Reel, clip by clip: the shot, what happens, the on-screen text, your spoken line and how many seconds each clip runs, against a target length. It is for creators who plan AI persona videos, and it exports a CSV or a JSON file you can import as a new project.
Free · No signup · Runs in your browser
Import the JSON as a new project and each row becomes a clip: an image prompt, a video prompt, your spoken line as dialogue and its length. The on-screen text stays in the CSV.
Pick a target length and your video model. Start from a structure or a blank clip.
For each clip, set the shot size, angle and camera move. Write what happens, the on-screen text and your spoken line.
Watch the running total and reorder or remove clips. Then download the CSV or the JSON.
As a rough guide, a 15-second video fits 3 or 4 clips and a 30-second video 5 to 7, with the hook shortest.
The JSON file is your plan as one project, one clip per row, in order. Import it as a new project and each row becomes a clip, up to 36 clips. Make each clip’s image first, then its video. The on-screen text stays in the CSV, for when you add text to the finished video.
| JSON field | Filled from |
|---|---|
image_prompt | Persona, place, what happens, shot size and angle |
video_prompt | What happens, shot size, angle and camera move |
dialogue | Your spoken line |
duration | The clip’s seconds, which must be a length the video model takes |
Turn each row into a clip of your AI persona: an image first, then a video, then an optional voice swap, on your own API keys. How making videos works →
The plan of a video, shot by shot: what each shot shows, how it is framed, what is said and how long it runs. For a Reel with an AI persona, each row becomes one image and one video clip.
No. You write what happens and what is said. The structures only name each part and hint at its job, such as “Hook” or “Call to action”, so you can fill it in.
Usually 5 to 7: a short hook, a few clips that show or explain, and a closing ask. Fewer, longer clips suit a talking video; more, shorter clips suit a demo.
It depends on the video model. Gemini Omni 1.1 Flash takes 3 to 10 seconds (4, 6, 8 or 10 on Kie) and Grok Imagine 1.5 takes 1 to 15. Pick your model at the top and the tool checks every clip.
The import takes four fields per clip: image prompt, video prompt, dialogue and duration. The on-screen text stays in your plan and the CSV, for when you add text to the finished video.
Yes. The CSV opens in any spreadsheet app, one row per clip: part, shot size, angle, camera move, what happens, on-screen text, spoken line, seconds and the start and end time. The plan is gone when you close the tab, so download the CSV or the JSON to keep it.