Guide · Psychology
Make a Talking-Head Video People Finish — Scripted with AI, in Under an Hour
On TikTok, more than 70% of viewers decide whether to keep watching in the first three seconds. Drop below 60% retention in that window and the algorithm barely shows your video to anyone — those three seconds are the whole match. Here is the trap of doing this with AI: it will write you a full script in ten seconds, fluent and confident, and it cannot tell you whether a single stranger will watch to the end. This guide takes one idea you actually have and walks it to a published, under-90-second talking-head video. The tools do the cheap part. A 2007 book about why some ideas stick does the part that decides whether anyone watches.
Before you start
- One small topic you genuinely know and have an opinion about. Narrow beats broad: "why I stopped color-coding my calendar" beats "productivity tips." A specific claim is a video; a category is a yawn.
- An AI writing tool you'll prompt, not obey — ChatGPT, Claude, or similar. You stay the editor; it stays the intern.
- A way to capture voice and picture: your phone camera for a face-to-camera take, or a text-to-speech voice for a faceless version. CapCut's library runs to a few hundred voices; ElevenLabs is the other common pick.
- A short-video editor with auto-captions. CapCut (剪映) is free and does both AI captions in 100-plus languages and text-to-speech in one place.
- One account to publish on — TikTok, YouTube Shorts, Reels, 抖音, 小红书. Pick where your topic already lives.
Compress it to one sentence
Before you write a word of script, decide the one thing a viewer should carry away — because one thing is all they'll carry. The Heath brothers call this finding the core: strip an idea until what's left is a single sentence that could survive being repeated by someone who only half-listened. Most first attempts are three sentences wearing a trench coat. Force it down to one. A useful prompt: paste your messy paragraph and ask the AI, "compress this into one sentence a distracted person could repeat tomorrow, then list what I'd have to cut to make that the only message." The cut list is where the real decision lives.
Why this firstIf you can't say it in one sentence, no camera, voice, or edit will rescue it. A muddy video is almost always a muddy core wearing good production. Fix it here, where fixing costs a minute, not after you've filmed.
Win the first three seconds
This is where most videos die, so spend your best effort here. The hook's job is a curiosity gap: open a small question the viewer now needs closed, or contradict something they assumed was settled. Ask the AI for eight opening lines for your topic — "each should open a curiosity gap or break a common belief; reject any that begin with a greeting or self-introduction." Then pick the one that makes you slightly uncomfortable, the one that feels a touch too bold. Comfortable openings are the ones everyone scrolls past.
The number-one killer"Hi everyone, today I want to talk about…" That sentence is the curse of knowledge in costume: you feel the setup matters because you know what's coming, but the viewer doesn't, and they've already swiped. No greeting, no throat-clearing, no "so basically." Open in the middle of the action, on the most surprising true thing you've got.
Let AI draft, then run SUCCESs by hand
Now hand the AI a real brief: "Write a 60–80 second spoken script in second person, conversational, short sentences. Include one concrete example, one real number, and one tiny moment of story. No bullet points — this is meant to be said out loud." You'll get a clean draft in seconds. Then do the part it can't: walk the script line by line against the Heaths' checklist. Simple — does every line serve the one core, or did a second idea sneak in? Unexpected — does the hook still hold at second three? Concrete — is each abstract claim swapped for a picture, a number, a thing you can see? Credible — is there one detail a skeptic could check? Emotional — will they feel one thing, not nod at six? Story — is there a small arc, not a list?
Where AI always slipsThe default AI draft is abstract and list-shaped — a slide deck read aloud. Your single most valuable move is to trade each abstract sentence for one concrete image. "Notes pile up" becomes "I had 1,900 notes and could find none of them." Concreteness is the difference between a point made and a point remembered.
Beat the curse of knowledge: read it to an outsider
You know your topic too well to hear what's confusing in it — that's the curse, and it's the quiet reason most explanations fail. So borrow ears that don't have it. Read the script aloud to one person who doesn't know the subject, and watch where their face goes blank. No one handy? Tell the AI: "You're a sharp person who has never heard of this topic. Read this back in your own words, and flag every sentence where you got lost or had to guess." Each stumble is a spot the curse hid from you. Cut it, or make it concrete — and don't argue with the confusion. The viewer won't argue either; they'll just leave.
Record or voice it, then burn in captions
Two roads from here. Face to camera: look at the lens, not the screen; talk a hair faster than feels natural; and keep the first three seconds hook-only, with zero intro. Faceless: paste the script into a text-to-speech tool, pick a voice that isn't robotic, and lay it over simple B-roll or a clean caption board. Either way, burn in captions. Most people watch on mute, so on-screen text isn't an accessibility nicety — it's a second hook doing half your retention work. CapCut's auto-captions handle this in a click; just read them back for the word it heard wrong.
Captions aren't optionalAssume the sound is off until the viewer chooses to turn it on — and they only choose that if the captions already pulled them in. The first caption on screen should be your hook, big and unmissable, before a single word is spoken aloud.
Cut under 90 seconds, then change one thing next time
Trim every "um," every breath before the real sentence, every wind-up. Keep the shape the Heaths would want: hook, one point, then a final line that snaps back to the hook so the video closes a loop instead of just stopping. Export vertical 9:16, put a caption on the very first frame, write a title that restates the hook, add two or three honest topic tags, and publish. Then comes the discipline that compounds: on your next video, change exactly one variable — the hook, or the title, or the cover — and nothing else. One change at a time is the only way to learn what actually moved the numbers instead of guessing.
On lengthClips under about 15 seconds finish at very high rates, but a tight 60–90 second talking-head holds fine when the hook has earned the stay. Length isn't the enemy; a slow start is. Earn second four, and you've usually earned the rest.
Run the whole loop on a real example. You think most note-taking apps quietly make people worse at thinking. Step one, the core: "Your notes app is a junk drawer, not a second brain." Step two, the hook: "I deleted 1,900 notes last night and remembered more this morning" — a curiosity gap, no hello. Step three, the AI draft, then you swap "information overload" for the image of a search bar returning nothing you need. Step four, you read it to a friend who stops you at "wait, why did deleting help?" — that blank look is the curse, so you add one concrete line about retrieval. Step five, a 70-second face-to-camera take, captions auto-generated and fixed. Step six, you cut to 68 seconds, drop a first-frame caption, publish — and resolve that next time you'll change only the cover. The whole thing fits inside an hour. The part AI couldn't touch — deciding which one sentence was worth three seconds of a stranger's life — stayed entirely yours.
Check your work
- I can say in one plain sentence what the viewer should remember.
- The first three seconds have no greeting or self-intro — they open mid-action on a curiosity gap.
- Every abstract claim in the script is carried by one concrete image, number, or example.
- I read the script to an outsider (or an AI playing one) and fixed every place they got lost.
- The video has burned-in captions, is vertical 9:16, and runs under 90 seconds.
- It's published — and I changed exactly one reviewable variable from last time.
The one line to keep
AI writes the words for free. What stays scarce — and stays yours — is deciding which single sentence is worth three seconds of a stranger's life.
Framework drawn from Chip & Dan Heath's Made to Stick — the SUCCESs checklist (Simple, Unexpected, Concrete, Credible, Emotional, Story) and the Curse of Knowledge. Tools named here (ChatGPT, Claude, CapCut, ElevenLabs) are examples, not endorsements, and their features change. Retention figures — that over 70% of short-video viewers decide within the first three seconds, and that sub-60% three-second retention sharply limits algorithmic reach — are drawn from 2026 short-form benchmark reporting. A popular-science, how-to reading; intellectual property belongs to the original authors. © vlog.bluecatbot.com 2026.