The speaker drifts out of frame
A 16:9 shot cropped to 9:16 loses whoever moves. Face tracking drives the crop across the clip instead of freezing it in the centre.
A podcast is the hardest thing to clip by hand: an hour of two people talking, three moments worth posting, and no visual cue to find them. Zypherius reads the transcript instead of the waveform, so the cut lands on what was said, not on where somebody laughed loudly.
One job costs 10 credits — the same for a 12-minute solo episode and a 58-minute interview.
A 16:9 shot cropped to 9:16 loses whoever moves. Face tracking drives the crop across the clip instead of freezing it in the centre.
Pauses and filler words are cut before the render, so a 40-second answer becomes a 30-second clip that keeps its meaning.
Subtitles are burned in and highlighted word by word in the language spoken, so the clip works on mute in a feed.
Every clip carries a Hook Score from 0 to 99 with a one-line reason, so the order you publish in stops being a coin toss.
Upload the episode as it came out of your recorder — up to 60 minutes and 500 MB, in MP4, MOV, AVI, WEBM, MKV or M4V. Anything past 30 minutes is analysed in overlapping windows, so a strong moment at minute 51 is found as reliably as one at minute 4.
A typical hour-long interview returns two to four clips. That is not a limitation dressed up as a feature: candidates that overlap each other by more than half are dropped, and what is left is what the transcript supports. Two clips cost 10 credits, three or more cost 15.
When a clip is nearly right, open it: retype a misheard surname, tap a pause to cut it, rewrite the hook line at the top. Re-rendering costs nothing, so the third attempt is as free as the second.
This tool cuts video. If your episode is a video recording — a camera, a Riverside or Zoom session, a stream — everything above applies. If it is audio only, there is nothing to crop or track, and you would be posting a static frame with subtitles, which the pipeline is not built to fake.
Two speakers in one frame work well. Separate speaker tiles in a grid work too, though the crop will favour the face it can track most confidently rather than switching between tiles on every line.
Up to 60 minutes and 500 MB per upload. Longer episodes need to be split before uploading.
Yes. Face tracking picks the speaker it can follow most reliably and moves the 9:16 crop across the clip. It does not cut back and forth between two tiles line by line.
The one being spoken. The transcriber detects it, and hooks and captions are written in that same language rather than translated into English.
10 credits for a job returning one or two clips, 15 for three or more. Credit packs start at EUR 5 for 100 credits, and credits do not expire.
Yes — pick 16:9 or 1:1 instead of 9:16 before the job runs. Face tracking is applied on the vertical crop, where it matters.
Seven days. After that the file is deleted and the interface shows an expired state instead of a dead link.
Ten credits are waiting in a new account — one full clip job, no card, no watermark on what comes out.