I wanted to publish videos without showing my face or filming anything. After a month of experimenting, here's the workflow that actually works — and the one thing most people get wrong.
The workflow
Every AI video starts with a script. The tool can't fix bad writing, so I spend most of my time there: short sentences, written for the ear, roughly 150 words per minute of video.
Then I pick a tool based on what the video needs:
A talking presenter? Use an AI avatar tool — you paste the script and a digital presenter reads it. (HeyGen clones your own face from a two-minute video; Synthesia offers a library of stock avatars.)
Original footage? Use a text-to-video tool — describe each scene and it generates the clip.
Repurposing an article? There are tools that turn a blog post into a video automatically.
The trick is matching the tool to the job. A presenter tool can't generate b-roll, and a text-to-video tool can't make someone talk to camera. Most "which tool is best" confusion online comes from comparing tools that don't compete.
What actually matters
After the video is generated, three things decide whether anyone watches: a hook in the first three seconds, captions (most short-form is watched muted), and a clean vertical export. The AI handles the visuals — the distribution is still on you.
I wrote up the full comparison and step-by-step setup here: https://aivideotest.com/













