Every clip arrives with a number, and the reason for it
A ninety minute episode holds maybe eight moments worth posting. Finding them by scrubbing is the part that makes people give up after one clip. Katto scores every candidate segment from 0 to 100 across four axes, so you open a ranked shortlist instead of a timeline.
Four axes, one number
The maximum each axis can contribute is fixed, and the weights are printed here because they are the whole argument. Hook and Flow carry sixty points between them: a clip that opens badly or rambles cannot be rescued by vocabulary.
The first three seconds only. A question word in the opening, an impact word, a number or a percentage, and whether the clip starts on a complete sentence. A clip that opens on sign-off language loses the whole axis.
Whether the clip holds together. Speech density measured by voice activity detection, how many scene cuts fall inside the segment, and whether audio energy stays sustained rather than lurching.
Density of substance words and figures, plus autonomy: does it start on a real word rather than a conjunction, and does it end on a finished sentence. Segments packed with self-promotion are penalised here.
The smallest axis on purpose. A question, a statistic, high-arousal vocabulary, and peaks in audio energy. It nudges the ranking, it never decides it.
What a score actually means
A number out of 100 invites you to read it as a percentage, and that reading is wrong. The useful question is where a clip sits against the others. Across the clips Katto has produced, the median lands at 68. About one in four reaches 80, and about one in ten reaches 85. So a 76 is a solid clip that beats roughly two thirds of the field, and a 60 is not a failure, it is simply below the middle.
This matters most when a whole video scores low. That usually says something true about the source, a quiet recording or a conversation without a clear turn, rather than something broken in the tool.
The same clip always gets the same score
Nothing in the scoring path asks a language model what it thinks. The inputs are the word level transcript, the voice activity map, the detected scene boundaries and the audio waveform, and the arithmetic on top of them is fixed. Even the sentence that explains a score is assembled from the four numbers rather than written for you.
That choice costs us some subtlety. A model would catch a joke landing in a way a word list cannot. What it buys is a system you can argue with: when a moment you wanted did not make the cut, there is a signal to point at, and re-running the video does not quietly produce a different answer. Models come in later, to write titles and descriptions for the clips that were already selected.
What gets pushed down before you see it
Two patterns score well and perform badly, so both are handled directly. The opening of a long video is high energy and hook-shaped, and it is also where sponsor reads and channel promotion live, so on anything over two minutes a segment starting in the first 45 seconds carries a fixed penalty. Sign-off language works the same way in reverse: a segment thick with subscribe-and-follow phrasing loses points on Value and on Trend, and a clip that opens on it loses its Hook outright.
What the score does not claim
It is not a view prediction, and any tool telling you otherwise is selling you a number it cannot stand behind. Reach depends on your audience, your caption, when you post and what the platform is favouring that week, none of which is visible in your source file. The score ranks the candidates inside one video against each other, so you start from the strongest three instead of the first three. The editorial call, which moment deserves the feed, stays yours.
Questions
No. Nothing in the scoring path calls a model. The four axes are computed from the transcript, the voice activity map, the detected scene boundaries and the audio waveform. A model writes the title and the description afterwards, on clips that were already chosen.
Yes. Same source, same numbers. That is the point of keeping the scorer deterministic: when a clip you expected does not appear, the cause is a signal you can name, not a model that answered differently today.
The median clip we produce sits at 68. Roughly one clip in four is at 80 or above, and one in ten reaches 85. A 60 is not a bad clip, it is a below-median one.
No, and we will not claim it does. Views depend on your audience, your thumbnail, your posting time and the platform mood that day. The score ranks candidates inside one video against each other. That is a sorting job, not a forecast.
Yes. Topics and a custom prompt add a keyword bonus to matching segments, so the moments you asked for climb the ranking. It is a plain text match, so it adds no cost and no extra latency.
Because they score well for the wrong reason. A cold open is high energy and hook-shaped, and it is also where sponsor reads and channel promos live. On videos longer than two minutes, anything starting in the first 45 seconds carries a fixed penalty.
See the scores on your own video
Every clip shows its total and its four axes, so you can check the ranking against your own judgement rather than take it on faith.