Going vertical means throwing away two thirds of the frame
Your source is 16:9. TikTok, Reels and Shorts are 9:16. Everything between those two numbers is a decision about what to lose, and getting it wrong makes a clip feel broken no matter how good the moment or the caption is. This is the part almost nobody talks about, and it is the part that decides whether a clip looks made by a person or made by a script.
The two easy answers, and why both fail
Centre-crop takes the middle slice and fills the screen. Punchy on a close-up, a disaster on a wide shot: the ball swings out to the wing and the crop simply does not show it, so you are watching a reaction to an action you cannot see.
Letterbox shrinks the whole frame and adds bars. Nothing is lost, but everything is tiny, and a full clip of it feels like a video of a video.
The real mistake is treating reframing as one global decision. The right framing is not a property of the video. It changes from moment to moment, so the decision has to change with it, and then hold steady inside each shot so the frame never jitters or drifts.
What it picks, shot by shot
Scene cuts set the boundaries, and each stretch between them is framed on its own terms. Talking content adds a second axis: one person, two, or a full panel each want a different layout, and choosing without cropping someone out or squashing everyone into a strip is most of the work.
One person carrying the moment. The crop tracks the speaker and holds steady rather than drifting, because a frame that breathes in and out reads as a mistake even when it is centred correctly.
A wide shot cropped to the middle is the classic failure: on a football clip the ball swings to the wing and you are watching a reaction to an action you cannot see. Instead the frame stays wide enough to keep the action, positioned on the actual subject rather than the geometric centre.
Two people in dialogue. The frame divides and puts a close-up on each, which is what a back-and-forth needs. It reverts to a single crop the moment one of them carries the segment alone.
Three or four speakers, each in their own cell. A four-way interview framed as a grid is legible on a phone in a way a cropped conference shot never is.
Gameplay, screen recordings and streams where a small camera sits in a corner. The content takes the top of the frame and the camera gets its own strip, instead of the crop picking one and losing the other.
Title cards, full-screen graphics and slides. Some things only make sense entire, so they are shown entire rather than cropped into nonsense.
And you get the last word
Automatic framing that you cannot correct is worse than no automation, because the one shot it misjudges is the one you cannot ship. The editor lets you set the framing for a single scene without disturbing the rest of the clip, and drag the crop box where you want it. Those per-scene choices hold all the way through to the export, which sounds obvious and is the thing most previews quietly fail to guarantee.
The honest part
None of this is one clever model. It is computer vision and audio feeding decisions we tune against real footage, over and over, watching where it looks wrong and fixing that. Most of the progress comes from the videos that break it: a fast football match, a multi-host podcast, a screen recording with no faces in it at all.
Fast sport is still where it is weakest. The frame shows the play without following it the way a human operator would, and the honest fix is adapting inside a shot rather than only at its edges. Reframe quality is not a feature you finish. For a clipping tool it is the whole game, so it keeps getting ground down. If you want to know where it is today, put a video through it and watch how it frames each moment.
Questions
Because it works on close-ups and fails on everything else. Centre-cropping a wide shot removes the part of the frame the shot existed to show. It is the single most common reason an automatically clipped video feels broken even when the moment chosen was right.
Nothing is lost, but everything is tiny. A whole clip in bars reads as a video of a video. Letterboxing is the right answer for a title card and the wrong answer for a conversation, which is exactly why the decision cannot be made once for the whole clip.
Yes, per scene. The editor lets you set the framing for one shot without touching the others, and those per-scene choices survive through to the export rather than being a preview-only illusion.
That is the case that breaks naive detectors, so it is handled explicitly: a screen recording or a stretch of b-roll with nobody in it gets framed on its content rather than on a face that is not there.
Fast sport is the weakest case. The frame shows the play without following it the way a human operator would, and the real fix is adapting within a single shot rather than at its boundaries. That work is not finished, and we would rather say so than let you discover it on a match.
No. It is face and person detection, scene-cut detection, speaker diarisation and audio, feeding a set of rules tuned against footage that breaks them. Most of the progress comes from the failures, not the successes.
Watch it frame your footage
The interesting test is not a talking head. Give it something with a wide shot, a title card and more than one person, and see where the frame goes.