More clips on the first screen, and a shorter wait
A twenty to forty minute video now opens on at least nine clips rather than five, a longer one on thirteen, and the jobs got faster at the same time.
The model already decided how many clips a video holds, and gave its reason for each one. That number was then floored at five, so a long source offering fifteen good moments opened on five and left the rest behind a button.
The floor now follows the length of the source: at least five under twenty minutes, at least nine from twenty to forty, at least thirteen beyond. The model can still ask for more than the floor, and on a dense source it does.
Voice activity detection used to wait its turn behind transcription and scene detection, although it depends on neither. It runs alongside them now, which took about seven minutes off a thirteen clip job.
Measured on real jobs the same day, on different sources, so read it as an order of magnitude and not a controlled test: five clips took eighteen minutes and forty eight seconds in the morning; after the change, thirteen clips took ten minutes and eighteen seconds and twenty two clips took eight minutes and thirty six seconds.
2 videos a month, no card. What is still being built is on the roadmap.