AI Talking Head Editor
title
subtitle
upload_video_tip
processing_time_tip
caption_style_tip
Upload a talking-head video and get back a tight edit: filler words, retakes, repetitions, and silent gaps are removed, cuts land on sentence boundaries, and word-level captions are burned in. The speaker is automatically kept in frame, even when you export a different aspect ratio.
Why use the AI talking head editor?
Removes fillers and retakes
Filler words, repeated phrases, false starts, and long silences are detected from the transcript and cut automatically.
Cuts on sentence boundaries
Edits are made at natural sentence boundaries so the trimmed video still flows like a normal take — no mid-word jumps.
Captions your way
Pick word-by-word highlighted captions, plain subtitles, or no captions at all. Captions are burned into the exported video.
Speaker stays in frame
When you export a different aspect ratio, the frame is recomposed around the speaker's face instead of cropping blindly.
10 output aspect ratios
Deliver the same edit as 16:9, 9:16, 1:1, 4:5, and more — one upload covers YouTube, Shorts, Reels, and TikTok.
Up to 120 minutes
Process anything from a 15-second intro to a two-hour interview in one pass, billed at 1 credit per 5 seconds of source video.
How it works
Upload your video
Add an MP4, WebM, or MOV up to 120 minutes. The duration is read locally to estimate credits before you start.
Set captions, ratio, language
Choose a caption style, pick an output aspect ratio (or keep the original), and set the spoken language or leave it on auto.
Download the edited video
Processing runs in the cloud. When it finishes, preview the result, download it, or grab it later from your history.
Made for talking-head content
Creator talking-head videos
Turn a raw single-take recording into a tight, captioned video for YouTube, LinkedIn, or your course platform.
Tutorials and courses
Clean up lessons and walkthroughs: remove stumbles and pauses while keeping every instruction intact.
Interviews and podcasts
Cut dead air and retakes from interview footage and export vertical clips for short-form platforms.
Vertical short clips
Reframe landscape recordings to 9:16 or 4:5 with the speaker centered and highlighted captions for silent-autoplay feeds.
FAQ
1 credit per 5 seconds of source video. Duration is rounded up to the next second with a minimum billed length of 3 seconds, and a single video can be up to 7200 seconds (120 minutes, 1440 credits). Caption style, language, and aspect ratio do not change the price. If processing fails, the credits are refunded automatically.
Filler words and verbal stumbles, repeated lines and retakes, and silent pauses beyond a natural beat. Cuts are placed on sentence boundaries so the result still sounds like one continuous take.
Three caption styles: Highlight (word-by-word), Plain (static subtitles), or None. Language can be auto-detected or fixed to English, French, Spanish, Portuguese, Italian, Russian, Simplified or Traditional Chinese, Japanese, or Korean.
Yes. Choose any output aspect ratio — 9:16, 1:1, 4:5, and seven more — and the frame is recomposed around the speaker so nothing important is cropped out. Leave it on Original to keep the source framing.