There is no single open-source model that embeds video, because video is three signals in a trench coat. What exists is a set of families, and which one you want follows from what you intend to match. Image encoders like CLIP, OpenCLIP, and SigLIP embed individual frames, so you sample frames and treat video as a pile of pictures. Video encoders like VideoMAE, InternVideo, and V-JEPA take a clip and encode motion across it. ImageBind puts several modalities into one shared space. And Whisper does not produce video embeddings at all, but it produces text, which any open text embedding model will happily turn into vectors. Most working systems use two of these at once.
Frame encoders are the default starting point
Sampling frames and running an image encoder over them is the most common approach, and for a reason: the models are mature, the tooling is everywhere, and you can describe a shot in words and get back the frames that look like it. CLIP established the pattern of training image and text into one space; OpenCLIP is the openly trained lineage of it; SigLIP is a later model in the same shape.
What you lose is time. A frame does not know what happened just before it. Anything defined by motion, a door closing or a ball being caught, is only visible to a frame encoder if a single frame happens to catch it mid-event. For a lot of archives that is an acceptable loss, because the thing people search for is a subject rather than an action.
Clip encoders see the motion, and cost more to run
VideoMAE and InternVideo are trained on video rather than stills, so the unit they embed is a short span. V-JEPA comes from the same ambition. These give you representations where a sequence means something, which matters when your queries are about events.
The price is practical: more compute per hour of footage, more decisions about how long a span should be, and a smaller ecosystem of ready-made recipes. Teams often start with frames and move to clip encoders only after seeing which queries fail.
Speech is a separate pipeline, and usually the cheaper win
Whisper and the open models around it transcribe audio. You then embed the text with any open text embedding model, or skip embeddings entirely and use ordinary keyword search over the transcript. For interviews, lectures, and podcasts, this path answers more real questions than any visual model will, and it is far less work. The cost side of extracting embeddings and timestamps usually confirms this: words are cheap, pixels are not.
Is there a best open-source model for video embeddings?
No, and anyone who names one is answering a different question than yours. The models differ in what they encode, not in a ranking. If your queries are about what is on screen, a frame encoder is the first thing to try. If they are about what happened, you want a clip encoder. If they are about what someone said, you want transcription and text search, and the visual models are beside the point. The useful move is to write down twenty real searches your team has made, sort them into those three piles, and let the biggest pile pick the family. For the version of this question that includes commercial services, the open-source options against a video understanding API cover the same ground from the other side.
The model is maybe a fifth of the work
This is the part that surprises people. Once you have embeddings, you still need to decide how footage gets chunked, where vectors live, how a user's sentence becomes a query, how results get ranked and deduplicated, and what someone looks at when the results come back. That last one decides whether the project is used. A list of vector hits with scores is not something a producer can act on; a set of time ranges they can open and preview is.
Vivu is a hosted version of that whole stack, where footage uploaded to a project is indexed once at ingest and a plain-language question comes back as time ranges a person can open rather than vectors someone still has to render into something reviewable. That is the trade against building: you stop choosing the encoder, and you stop maintaining the five things around it.
When you do not need embeddings at all
If your archive is searchable by filename because someone has been disciplined about naming, embeddings will not beat that. If everything you look for was spoken, transcripts plus grep get you most of the way. If the library is a few dozen videos, a person who has watched them is the fastest index you will ever have. Embeddings start to pay when nobody on the team can hold the library in their head and the queries are about what things look like. Whether to build the pipeline or connect a hosted one is the question that follows.
How to decide
Pick by the shape of your queries, not by what is new. Sort your real searches into on-screen, in-motion, and spoken, and build for the biggest pile first. Then be honest about the second decision, which is whether you want to run this system for years. Open models are free to download and expensive to operate, and the operating cost lands on whoever maintains the ingest pipeline at 2am.
FAQ
Can I use CLIP for video search?
Yes, and it is the usual first build. You sample frames from each video, embed them with CLIP or OpenCLIP, store the vectors with their timestamps, and embed the user's text query into the same space to find matching frames. The result is a list of moments where the screen looked like the description. The known limit is that each frame is judged alone, so queries about motion or sequence do not land well, and you will need to decide how to collapse many matching frames from the same scene into one result.
How many frames should I sample per video?
Sample rate is the main knob, and it trades cost against misses. Dense sampling catches brief events and multiplies your storage and compute; sparse sampling is cheap and will walk straight past anything that was on screen for a moment. A common compromise is to sample more densely where the picture changes, using shot detection to put at least one frame in every shot, rather than sampling on a fixed clock. Test it against real queries before locking it in.
Where do I store video embeddings?
In a vector database or a vector index inside a database you already run. The part specific to video is that every vector needs a timestamp and a source file alongside it, because the answer a person wants is a position in a video, not a document ID. Plan for re-indexing too: when you change the model or the sample rate, every vector you have is stale, and the storage layer should make replacing them cheap.
Can open-source models find things people said as well as things on screen?
Not with the same model. Visual encoders embed pixels and have no access to speech; what people said comes from transcription, then either text search or text embeddings over the transcript. Systems that answer both kinds of question run both paths and merge the results. If you only have budget for one, start with speech, because it is cheaper to run and it answers the most common requests in interview and lecture footage.