How to make a viral AI short in 30 minutes (2026 edition)
A four-tool stack and a repeatable script that turns a single idea into a polished, scroll-stopping vertical video — without a camera, a face, or a studio.
In 2026, a faceless 45-second short can pull a million views in a week and cost you nothing but the time it takes to read this guide. The tools have quietly become good enough that the bottleneck isn't AI quality anymore, it's the script and the rhythm.
This is the exact stack we use to turn a one-line idea into a polished vertical video in under 30 minutes. No camera. No studio. No on-screen face required.
The four-tool stack
You need exactly four pieces. Don't add more, every extra tool eats time and never pays it back on a 45-second clip.
Total monthly cost if you publish 3 shorts a week: about $30. The same stack a year ago cost over $100 and produced visibly worse output.
- A voice, ElevenLabs v3 for the narration.
- A look, Midjourney for 9:16 B-roll stills (or Flux 2 if you want lower cost).
- An editor, CapCut on desktop, free, fast, beat-sync built in.
- A finisher, Submagic for captions, hooks, and AI B-roll suggestions.
Step 1: Write the 45-second script (5 minutes)
Open a blank document and write three things:
Read it aloud and time yourself. If it runs longer than 45 seconds at a natural pace, trim words. Vertical video lives or dies on pacing.
- The hook (first 1.5 seconds of audio). One sentence, contrarian or curious. "Most people lose money on Bitcoin because they buy what their friends buy." Not "In this video we'll talk about…", that gets scrolled past in 0.4 seconds.
- The middle (about 110-130 words). Three concrete points, each tied to a visual. If you can't picture the B-roll for a sentence, cut the sentence.
- The CTA (1 sentence). Follow for X. Save for later. Comment your take. One ask, never two.
Step 2: Generate the voiceover (3 minutes)
Open ElevenLabs and pick a voice from the v3 library. Test three voices on the first sentence before committing, voice fit matters more than voice quality. A bored "premium" voice will out-bomb an enthusiastic "okay" one every time.
Settings that work well for shorts:
Generate, listen once, and only re-roll if a specific word is mispronounced. Don't chase perfection, viewers don't.
- Stability: 35-45 (a touch of variation reads as human)
- Similarity: 75
- Style: 20-30 if your script has emphasis, 0 if it's calm explainer
Step 3: Generate B-roll stills (10 minutes)
This is where most creators waste an hour. The fix is to commit to a single visual style up front and reuse a Midjourney style reference (--sref) on every prompt.
Workflow:
You'll end up with a visually coherent set of vertical stills that look like they came from one art director. That's what separates a good AI short from one that screams "AI slop."
If Midjourney's price isn't justifiable for your output volume, Flux 2 does 80% of the same job at a fraction of the cost, slightly less stylized but plenty good for shorts.
- Generate one "hero" image you love. Copy its --sref value.
- For every line of your script, write one prompt. Append --ar 9:16 --style raw --sref <your-code> to all of them.
- Run them in a batch. Pick the best variant of each. Don't upscale yet.
- Only upscale the 8-12 you'll actually use.
Step 4: Edit in CapCut (8 minutes)
Drop the voiceover on the timeline first. Always voice first, visuals second, the audio dictates the cut, not the other way around.
Then:
Export at 1080×1920, 60fps. Don't ship 30fps in 2026.
- Place each B-roll image on the beat where its sentence starts.
- Use Auto-cut to beat for any background music. CapCut's beat detection is shockingly good in 2026.
- Add subtle motion to every still: 5-second slow zoom, alternating directions. Static images on a vertical screen feel dead within 2 seconds.
- For background music, use CapCut's library and pick something at 110-120 BPM with a low ducking volume so the voiceover stays on top.
Step 5: Caption and hook in Submagic (3 minutes)
Drop your CapCut export into Submagic. Three things to do:
Export. You're done.
- Pick a caption template that matches your channel. Bold, animated, single-word emphasis works best for educational content. Word-by-word reveal works best for storytelling.
- Use the AI hook generator on your first 1.5 seconds. It rewrites your hook into 3-5 alternatives, pick the punchiest. This single step is worth the entire Submagic subscription.
- Accept B-roll suggestions sparingly. The AI will sometimes propose a stock clip that genuinely improves a beat, accept those, ignore the rest.
What this looks like in practice
A typical session, end to end:
The first short you make will take 90 minutes because you're learning the tools. The fifth one takes 25.
Mistakes that kill reach
A few patterns we see kill otherwise-decent shorts:
- Inconsistent visual style. This is why --sref matters. Mixing photorealism and illustration in one short reads as amateur even when individual shots are great.
- Generic captions. If your captions are just transcribed text in a default font, you're losing 30% of completion rate vs. a designed caption template.
- Two CTAs. "Follow AND comment AND save" trains the algorithm to think your audience does nothing. One ask only.
- Loud music. If your voiceover isn't 12-15 dB above the music bed, viewers swipe.
- No hook variant. Submagic generates 5 hook alternatives in 10 seconds. The default script-first hook is almost never the best one.
Where to go next
Once you've shipped 10 shorts with this stack, the bottleneck shifts from "can I make a video" to "can I find ideas worth making." That's the harder problem, but it's the one worth solving.
For now, just ship the first one. Today.