---
name: thumbnail-frame-candidates
description: "Pull thumbnail frame candidates from your own footage with Vivu: find moments where the product is held up to the camera and reaction close-ups, check each, export sharp stills. Use after a shoot."
---Thumbnail frame candidates from your own footage
A creator hands over the footage of one episode they shot (1 to 3 local files, the A camera talking head and demo, raw or the finished cut) and names the kinds of picture they want behind the thumbnail: holding the product up to the camera, a big reaction close-up, the first look at the product. Claude indexes the files in a private Vivu project, runs one search per kind of picture, checks every returned window against the footage, pulls frames every 0.25 seconds inside the confirmed shots, ranks them for sharpness, drops closed eyes and motion blur, and hands back 3 to 5 full resolution PNGs per kind, a contact sheet, and a candidates.csv with an empty "use this" column for the creator to fill.
The value is in the picture itself. A transcript knows what the creator said, but not the second they lifted the box to the lens, that the editor cut to a punch-in close-up for less than a second, or that the laugh everyone remembers has the eyes squeezed shut in every frame. Vivu narrows 20 minutes of footage to a handful of windows; Claude then looks at the frames, because the windows are wider than the shot and a search result alone is not a thumbnail.
When to use
Use when a creator asks for thumbnail frames from an episode they shot: "find me thumbnail frames from this episode", "pull stills where I hold the product up", "grab my best reaction shots", "give me a few options for the thumbnail background".
Not for:
- Designing or editing the thumbnail. The skill stops at the raw frame; the creator builds the thumbnail in their own editor. Vivu does not generate or retouch images.
- Predicting which frame gets more clicks, rating looks, or identifying people. The skill does none of these.
- Footage the creator has no rights to, such as clips from other channels.
- One quick question about one video ("is there a moment where I hold up the box?"): search Vivu directly instead.
Working principles
- Report measured numbers, not estimates. When a number is an estimate, say so.
- Nothing is "verified" until it has been checked against the source video. A Vivu result is a candidate window until Claude has looked at frames from it.
- Stop and tell the user when a required capability or tool is missing. Do not guess around it.
- Ask the user before anything that is expensive to redo or that acts on their behalf: indexing (it uses index minutes) and the shot list (it decides the searches).
- Report the real count. When a kind of picture has fewer good frames than the target, hand back what exists and say so; never pad with blurry or near duplicate frames.
- Reaction close-ups are the least reliable search. Treat every reaction result as a candidate and keep the window sheet, so the creator can pick the expression themselves.
What you need before starting
Check each item at the start of the run and tell the user plainly what is missing before doing anything else.
| Requirement | Why | How to check |
|---|---|---|
| Vivu connector with write access | create a private project, open its upload page, search | vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the connection is read only; ask the user to reconnect Vivu with write access |
| A shell on the computer that holds the footage, with a folder it can write | measuring minutes, extracting frames and ranking them all run locally | create the working folder in Step 1 |
| ffmpeg and ffprobe | frames, contact sheets, sharpness scores | ffmpeg -version and ffprobe -version |
| The blurdetect filter in ffmpeg (optional) | a sharpness score per frame | ffmpeg -hide_banner -filters lists blurdetect; without it, rank by eye |
| An upload path | moving the footage into Vivu | vivu_open_upload_page plus the user's own browser, or a browser tool that can attach local files |
This skill runs in Claude Code on the user's computer (the terminal or the Code tab of Claude Desktop), because it reads local video files and runs ffmpeg. It needs no residential IP (nothing is downloaded), no other connector, and no scheduler.
Inputs to collect
Ask for anything missing, most important first.
- The footage: 1 to 3 local files for this episode. Required.
- The shot list: 3 to 5 kinds of picture in the creator's own words. Default: "holding the product up to the camera" and "big reaction close-up"; offer "first look at the product" as an optional third (its search is the least tested, see the worked example).
- Frames per kind. Default: 3 to 5.
- Whether anyone other than the creator appears on camera (family, a guest, people in the background), and whether any of them is a minor. Required; it decides the consent question in Compliance.
- Episode name for the folder. Default: the first file's name without extension.
Files and state
Keep everything in one working folder next to the footage:
thumbnail-picks/EPISODE/
config.json files, shot list, Vivu project id, frames per kind
state.json finished steps, Vivu video ids, job id per kind, verdict per window
windows/ per window: start, middle, end frames and a 0.25 s contact sheet
scores/ blurdetect output per confirmed shot
frames/KIND/ chosen full resolution PNGs, named SOURCE_MM-SS.s.png
sheets/KIND_sheet.png contact sheet of the chosen frames
candidates.csv one row per chosen frame
state.json records each finished step, the video ids, the job id of each search and the verdict on each window, so a rerun after an interruption starts at the first unfinished step and never searches or checks the same window twice.
Step 1: Check the Vivu connector and the local tools
Goal: confirm everything in What you need before touching the footage.
- Call vivu_get_account. If the tool is missing, tell the user to add the Vivu connector in Claude (https://mcp.vivu.ai/mcp) and stop. If can_create_projects is not true, or a later write fails with "has not granted vivu.write", ask the user to reconnect Vivu with write access.
- Run ffmpeg -version and ffprobe -version. Run ffmpeg -hide_banner -filters and look for blurdetect; note whether it is there.
- Create thumbnail-picks/EPISODE/ with the subfolders above and an empty state.json.
Done when vivu_get_account shows can_create_projects: true, ffmpeg and ffprobe run, and the working folder exists.
Step 2: Collect the footage and the shot list, then price it
Goal: an approved shot list and an approved cost before anything is uploaded.
- Collect the inputs above. Ask the consent questions in Compliance now, before any upload.
- Measure each file:
ffprobe -v error -show_entries format=duration -of csv=p=0 FILE
FILE is the path to one video. Sum the durations and divide by 60 for the index minutes. 3. Call vivu_get_usage for the plan and what remains this month. Show one table:
| This episode | Plan allowance | Remaining this month | |
|---|---|---|---|
| Index minutes | sum of durations | from vivu_get_usage | from vivu_get_usage |
| Search credits | kinds of picture × 5 (estimate) | from vivu_get_usage | from vivu_get_usage |
The plan facts: Free is $0 a month with 20 indexing minutes a month and 50 search credits a month; Premium is $30 a month with 180 indexing minutes and 500 search credits. A precise search costs 5 credits in total; a fast search costs 1 credit. Every search in this skill is precise, because only precise returns a time window. Label the credit line as an estimate, since a reworded search costs another 5. 4. When the episode does not fit, offer to index only the files that matter for the thumbnail first (for example the A camera file without the B roll, or the first half where the unboxing happens), then a larger plan as the last option. Do not drop files without asking. 5. Show the shot list as the exact search wording you will use (see Step 4) and one sample candidates.csv row, and wait for a yes.
Done when the user has approved the shot list, the sample row and the cost table.
Step 3: Create a private project and upload
Goal: every file indexed and ready.
- Call vivu_list_projects and reuse a project for this channel if one exists; otherwise call vivu_create_project with a name like "Thumbnails CHANNEL" and visibility "private". The default visibility is "organization", which shows the footage to everyone in the user's Vivu workspace, and raw takes are often not meant for that.
- Call vivu_open_upload_page with the project id right before uploading. The link expires in 180 seconds, so request it at the moment of use and never paste it into a file or message.
- Give the link to the user to open in their own browser and pick the files, or open it with a browser tool that can attach local files. Claude in Chrome accepts at most 10 MB per upload call, which most episode files exceed; in that case the user adds the files in the Vivu web app. Never split or recompress the footage to fit.
- Vivu replaces spaces and brackets in file names with underscores. Record each video id and Vivu file name in state.json next to the local path, so results map back to the right local file.
- Poll vivu_list_videos until every file shows ready. Files can sit in "queued" for several minutes before indexing starts; that is not a failure.
Our test upload used a different, internal route; the browser upload by the user was not exercised in our test run.
Done when vivu_list_videos shows every file ready and state.json maps each video id to a local file.
Step 4: Search once per kind of picture
Goal: one list of candidate windows per kind.
Write each search the way the creator would describe the frame, and add what does not count, because the first draft of a wording usually lets in look-alikes. These are the wordings that held up in our test run; adapt the nouns to the episode:
| Field | Query | Mode | maximum_results |
|---|---|---|---|
| product_to_camera | the creator lifts a product or its box up toward the camera to show it to the viewer, so the product is big and clearly visible in the shot; holding something low in the lap while using it, or hands working on a desk filmed from above, does not count | precise | 20 |
| big_reaction | the video cuts to a tight close-up of the creator's face, zoomed in so the face fills most of the frame, while the creator makes an exaggerated expression such as wide eyes, a wide open mouth, a big grin or laughing; the creator talking calmly in a normal medium shot does not count | precise | 10 |
| first_reveal (optional) | the first moment the product itself is taken out of its packaging and seen in full | precise | 10 |
The first_reveal search can return an accessory coming out of its own box; treat its windows as candidates like the others.
maximum_results is also the ceiling on how many windows come back. 20 leaves room for a long episode where the product is shown many times; 10 is enough for reactions and the reveal, which happen a few times per episode. Raise it when a kind clearly has more moments than came back.
These are visual searches, and finding a picture by what it shows is one of Vivu's weaker abilities: most windows it returns are right, but it misses some, it can return a look-alike, and an empty result does not prove the footage has no such shot. That is why Step 5 checks every window and why the creator sees the sheets. The chain is one precise search per kind, then frames: the search narrows the episode to windows, and Step 5 decides what is in them. There is no "said" to "shown" pairing here, because the thumbnail needs the picture; when a creator's kind of picture is tied to a line ("when I say wow"), a speech search can be added as a second signal, but that was not exercised in our test run.
Run each query with vivu_search_videos (project_id, query, mode, maximum_results). It returns a job id. Call vivu_get_search_results until complete is true; each call can wait up to 45 seconds, so a pending search is not a stalled one. Show the Vivu result page link in the reply if the user wants to browse; it expires after four hours, so it never goes into candidates.csv or any file. Save each job id in state.json.
Done when every kind has a completed search and its windows are listed in state.json.
Step 5: Check every window against the footage
Goal: a verdict on every window before any frame is chosen.
A returned window is wider than the shot inside it. A punch-in close-up can last well under a second inside a window several seconds long, with the medium shot or the B camera on either side, so three frames are not enough to judge it.
- Pull the start, middle and end frames:
ffmpeg -v error -ss SECONDS -i FILE -frames:v 1 -q:v 3 windows/KIND_N_POS.png
SECONDS is start_ms / 1000, the midpoint, or end_ms / 1000 minus 0.1; KIND is the field, N the window number, POS start, mid or end. 2. Make a 0.25 s contact sheet of the window:
ffmpeg -v error -ss START -to END -i FILE -vf "fps=4,scale=320:-1,tile=8x4" -frames:v 1 windows/KIND_N_sheet.png
START and END are the window in seconds. One sheet holds 8 seconds at 0.25 s; for a longer window use fps=2 or fps=1 and then make a 0.25 s sheet of the part that matters. Cell number times 0.25 plus START is the time of that cell. 3. Look at the sheet. The window is real when the kind of picture is visible in it as the creator described it. Write down the exact span where it is on screen (for example 58.0 to 58.75 s), or "false" with the reason. 4. Judge from the frames, not from the reason text Vivu returns. The reason is a description, not proof: in our test run it called a stylus pointed at the camera a product shown to the viewer. 5. Count the false windows and tell the user how many there were per kind.
Done when every window has a verdict and, for real ones, the on screen span is written in state.json.
Step 6: Pick sharp frames with eyes open
Goal: 3 to 5 good frames per kind, or the real number if there are fewer.
- For each confirmed span, score every 0.25 s frame (needs blurdetect; lower blur is sharper):
ffmpeg -v error -ss SPAN_START -to SPAN_END -i FILE -vf "fps=4,blurdetect=block_width=32:block_height=32,metadata=print:key=lavfi.blur:file=scores/KIND_N.txt" -f null -
Each frame's pts_time in scores/KIND_N.txt is seconds after SPAN_START. 2. Rank only inside one confirmed shot. The score is not comparable across a cut: a busy medium shot scores sharper than a close-up of a face, and a shallow focus close-up of a product scores worse than the cluttered frame before it. Keep SPAN_START and SPAN_END inside the shot from Step 5. 3. Look at the best few frames of each span and drop closed or half closed eyes, motion blur, a hand covering the face, and burned-in overlays the creator will not want (graphics, captions). Without blurdetect, pick from the 0.25 s sheet by eye. 4. Spread the picks across different moments; two frames a quarter second apart are one option, not two.
Done when each kind has its picks listed, or a note saying how many good frames exist and why there are not more.
Step 7: Export the stills and the candidates table
Goal: files the creator can open in their thumbnail editor.
- Export each pick at full resolution as PNG (lossless):
ffmpeg -v error -ss SECONDS -i FILE -frames:v 1 frames/KIND/SOURCE_MM-SS.s.png
SOURCE is the local file name without extension; MM-SS.s is the time, for example 03-59.0. 2. Build one sheet per kind:
ffmpeg -v error -pattern_type glob -i "frames/KIND/*.png" -vf "scale=480:-1,tile=4x2" -frames:v 1 sheets/KIND_sheet.png
- Write candidates.csv with one row per pick. Sample row:
shot_type,source_file,timecode,frame_path,blur_score,eyes,notes,use_this
product_to_camera,EPISODE_A_CAM.mp4,01:52.0,frames/product_to_camera/EPISODE_A_CAM_01-52.0.png,7.62,open,box held centered at chest; shopping bag at right edge,
| Field | Source | If unavailable |
|---|---|---|
| shot_type | the kind from the shot list | none |
| source_file, timecode | the user's local file and the frame time | none |
| frame_path | the exported PNG | none |
| blur_score | scores/KIND_N.txt | "visual pick" |
| eyes | Claude's look at the frame (open, half closed, not in frame) | UNCERTAIN |
| notes | what is in the frame, overlays, anything in the way | blank |
| use_this | left empty for the creator | blank |
- Show the sheets in the reply and list per kind: windows returned, windows that were false, frames kept, and anything missing.
Done when every kind has its PNGs, a sheet, and rows in candidates.csv, and the user has seen the sheets.
Worked example from our test run
In our test run we used 3 public creator unboxing videos (14.0 minutes in total) shared under a Creative Commons license: an expressive console unboxing with punch-in close-ups and a top-down B camera, a watch unboxing filmed in one medium close shot, and a calm tablet first impressions video. Indexing took 14.0 index minutes, and all 3 files were ready about 8 minutes after the upload started. Ground truth was written from contact sheets before any search; after the search we corrected it twice from the frames (one chest-height product hold that had been filed under first reveal was added, and one reveal time was moved), and the product counts below use the corrected set of 10 moments.
Product held up to the camera: the first wording returned 5 windows, 5 real, and missed 5 moments (it found only holds right at the lens). A second wording returned 8 windows with 2 false (a tablet held low while drawing, a small dock at chest height in a wide shot). The wording in the table above, after 2 rewordings on the same corpus, returned 11 windows: 10 real, 1 false (a stylus pointed at the camera), 1 missed (the console held over the head).
Reaction close-up: the first wording returned 1 window, 1 real, and missed 3 of the close-ups. The wording in the table, after 1 rewording on the same corpus, returned 3 windows, 3 real, 1 missed. All the reactions were in the one expressive video, and the laugh was audible, so sound may have helped. Each window was about 7 seconds while the close-up inside it was far shorter, and the picked laugh frame had the eyes nearly shut, as did every frame of that laugh; a further reaction frame was dropped for closed eyes.
First reveal: 3 windows, 2 real, 1 false (a charger coming out of its own small box). A search for a dog or cat, which the footage does not contain, returned 0.
In all, 8 precise searches were run, 40 credits (estimate). A normal run uses one search per kind of picture, about 15 credits per episode with the optional reveal included (estimate).
Compliance
- Use only footage the creator shot or has the rights to. Confirm this in Step 2, before upload.
- People other than the creator on camera (family, guests, people in the background) should agree before their face goes on a thumbnail. Ask in Step 2 and leave their frames out if the answer is no or unknown.
- Minors: pick frames showing a minor only with a parent or guardian's explicit consent; otherwise leave those frames out. Ask in Step 2.
- The footage stays in the user's Vivu project until they delete it. The skill creates the project as private; it never deletes projects or videos unless the user asks, and asks again before calling vivu_delete_video or vivu_delete_project.
- The skill does not recognize faces or identify anyone, does not rate looks, and does not predict click rates. Expressions are described by what the frame shows ("laughing with mouth open"), nothing more.
Known failure modes
| Symptom | Cause | Fix |
|---|---|---|
| (observed) most reaction close-ups missing from the first search | the first wording asked for a close-up with a big reaction but did not say the edit cuts to a tight close-up; it found 1 of 4 in our test run | name the cut to a tight close-up and exclude the calm medium shot, as in the Step 4 table; still treat results as candidates |
| (observed) a reaction window is mostly the medium shot or the B camera | the window is about 7 seconds while the close-up inside it is much shorter | make the contact sheet in Step 5 and pick only inside the close-up span |
| (observed) product search misses long holds at chest height or a tablet held up | the first wording only matched holds right at the lens (5 missed in our test run) | describe lifting the product up toward the camera so it is big and clearly visible, raise maximum_results, and check the sheet for long holds that come back as more than one window |
| (observed) a stylus pointed at the camera, or a tablet held low while drawing, returned as a product shown to the viewer | the search treats any held object near the camera as a match; the reason text repeats the claim | judge on the frames; add what does not count to the wording |
| (observed) the best laugh frame has the eyes shut | people squeeze their eyes when they laugh; every frame of that laugh was affected | keep it with eyes marked half closed in candidates.csv, or let the creator pick another moment from the sheet |
| (observed) the sharpest score belongs to a frame from the next shot | blurdetect scores busy wide frames as sharper than a face close-up, and shallow focus product shots as blurrier | rank only inside one confirmed shot (Step 6) and look at the frames |
| an empty result for a kind of picture | the footage may not have it, or the search missed it; an empty result does not prove the footage has no such shot | scan a 1 s sheet of the A camera file for that kind before telling the user there is none |
| "has not granted vivu.write" | Vivu connected read only | user reconnects Vivu with write access |
| upload page asks to sign in or shows an error | the one time upload link expires after 180 seconds | request a new link right before opening it |
| a file is rejected by the browser upload tool | file above the tool's size limit | the user adds that file in the Vivu web app; do not split or recompress |
| a result's file name does not match the local file | Vivu replaced spaces and brackets with underscores | match on the video id recorded in state.json |
| burned-in graphics on an otherwise good frame | the creator's own edit added overlays (captions, emoji, effects) | note it in candidates.csv; prefer the raw A camera file over the finished cut when both exist |