Every route splits the job the same way. First you decide which seconds are worth using, then you produce a vertical clip out of them. The second half is close to solved and cheap. The first half is where the hours go, and the route that fits you depends more on how many long videos you have than on how long any one of them runs.
Cutting by hand
You scrub the timeline, find the moment, set in and out points, reframe, add captions, export. For one video this is fine, and it often gives the best result, because you are the person who knows which line actually landed and which one only felt good in the room. It stops working when the input is a weekly show. An hour of footage takes one or two full passes to mine properly, and nobody does that fifty weeks in a row.
Automatic clippers
You paste a link or upload a file and get back a set of candidate clips, already reframed and captioned. These tools pick on generic signals: speech density, sentiment swings, a question followed by an answer. When the good parts of your video look like the loud parts, this works well. When the good part is a quiet aside forty minutes in, it gets missed, and you will not notice, because you only ever see what came back. The full version of that route is walked through in the hands-off version of this. Its real limit is scope: it works on one file that you hand it, which means you still have to know which file.
Working from the transcript
You get a transcript, read or search it, write down timestamps, then cut against those timestamps in your editor. It is much faster than watching, and it is the only route where you can prove you covered the whole video. What it catches is what was said, so a reaction, a demo, or a shot that works with the sound off will not show up in it. Turning transcript lines into cut points is a specific skill and worth learning if you talk for a living.
When the input is a catalog, not a video
The question changes shape once you have a back catalog. Podcast teams hit this hard: hundreds of episodes, transcripts that were often paid for per episode, and still no way to answer "where did we explain that objection well" without opening files one at a time and trying search tricks on the website. At that point the expensive step is not cutting, it is deciding which of three hundred hours deserves an editor's afternoon.
This is the case for putting a retrieval layer over the archive instead of a clipping tool over one file. Each video is indexed once by content, and after that you ask questions and get timestamps back across everything you own. Vivu follows that route for the episodes you upload to a project: each is indexed there, and a question about all of them becomes a single search instead of another pass over the catalog. It does not cut anything for you, and the output is a set of timestamped moments you take into your editor. What matching on content actually means is unpacked in how these systems find a moment inside a file.
When you don't need any of this
If you publish a few clips a month from videos you recorded yourself and still remember, none of this pays for itself. You already know where the moment is, and every tool here is a way of buying back knowledge you have not lost. The same goes for a single long video you are about to work on this week. Open it, scrub it, cut it, done.
There is also the case nobody likes to say out loud. If the long video was not good, the shorts will not be either. Clipping is a distribution move, and it multiplies whatever was in the source. A search layer over a weak archive returns weak moments, faster.
Which side you are on
Two questions settle it. Do you know where the moment is, and how many videos deep is it? If you know and it is one video, cut it by hand today. If you do not know but it is one video, a transcript will get you there in twenty minutes. If you do not know and it could be in any of the last two hundred, the tool you need is one that answers questions about the whole archive, and the clipping is the easy part that comes after.
FAQ
How many shorts can one long video realistically produce?
Fewer than the tools imply. A one-hour conversation usually holds two or three moments that stand on their own without setup, and maybe five or six more that work if you are willing to add context in the caption or the first two seconds. Auto-clippers will hand you fifteen candidates from the same hour, which is not the same as fifteen publishable clips. The count depends on format: a scripted or teaching video is dense and yields more, a loose interview yields less than its length suggests.
Is it worth going back and clipping old episodes?
Yes for evergreen material, no for anything tied to a news cycle, a product version, or a price. Old episodes have an advantage new ones do not: you already know which ones performed, so you are choosing from a ranked list instead of guessing. The catch is that going back is the case where you have the least memory of what is inside, so it is the case that most needs a way to search the archive rather than rewatch it.
My show is audio-first. Can I still make shorts from it?
You can, but you are making a different thing. Audio with a waveform and captions is a format that works on some platforms and dies on others, and it competes against clips with faces in them. Teams with an audio-first show usually either film the recording sessions going forward, or accept that clips are for search and discovery rather than reach. Deciding that before you buy tooling saves you from buying the wrong kind.
Can I clip footage I did not record myself?
That is a rights question rather than a tooling question, and the answer depends on your agreement with whoever owns the footage. Guest appearances, licensed music, client work, and event recordings each carry different terms, and no clipping tool checks any of them for you. The practical habit is to settle reuse rights at the point the footage is created, because tracking down permission afterwards costs more than the clip is worth.