---
name: stream-vod-clip-shortlist
description: "Turn your own long stream recording into a timestamped clip shortlist with Vivu: game banners checked on frames, stories checked on captions. Use after a stream to find moments worth clipping."
---Stream VOD clip shortlist with Vivu
This skill turns a streamer's own recording of a long stream (the VOD file downloaded from Twitch or YouTube, or a local OBS recording) into a shortlist of moments worth cutting into short videos, each with its time in the original file. Claude uses the chat replay or the streamer's stream markers to pick a few busy windows, cuts only those windows out of the recording with ffmpeg, prices them against the user's Vivu plan, uploads and indexes them in a private Vivu project, runs two standing searches (game banners and notices on screen, and stories the streamer tells), checks every candidate against frames or captions, and writes clip_shortlist.csv plus a readable clip_shortlist.md. The streamer opens the recording at those times and cuts the clips; the skill does not edit, render or post anything.
The value is in what a title, a chat log or a transcript cannot show. A stream has no script, so its title says nothing about which minute is good, and the chat replay only shows when viewers were typing, not what was on screen. Game banners (a quest title, a reward notice, a level up, the name of a new area) are text drawn on the game screen, so they never reach the captions. Stories are in the speech, and the footage shows whether that minute is watchable. Vivu's banner search is used only to collect candidates, because in our test run most of what it returned was not a banner (name tags floating over other players, a book page, a chat line) and its reason text described banners that were not there. A banner reaches the shortlist only after Claude has seen it in the frames, and the banner text in the shortlist is read from a frame.
When to use
Use when a streamer asks to find clip moments in their own stream recording, to list timestamps worth clipping from last night's VOD, to find where a quest started or a reward popped up in a recording, or to find the stories they told on stream.
Not for these:
- Laughs, shouts and other loud reactions. The skill does not look for them, because in our test run the search could not tell them from normal talk (numbers below). Use stream markers placed live for those moments.
- Other people's streams. The skill works on recordings the user owns (see Compliance).
- A one off question about one moment ("when did I finish that quest?"): search Vivu directly.
- Cutting, captioning or posting the clips: this skill only lists times; use an editor for the rest.
In our test run a search for laughs and shouts returned 13 moments, 5 real and 8 false; Vivu's reason called quiet smiles bursting out laughing.
Working principles
- Report measured numbers, not estimates. When a number is an estimate, say so.
- Nothing is "verified" until it has been checked against the source video. Candidates stay candidates until triaged: rejected ones never reach the shortlist, and ones that could not be checked (a story with no captions or transcript, a frame Claude cannot place) are listed apart from it as unchecked.
- Stop and tell the user when a required capability or tool is missing. Do not guess around it.
- Ask the user before anything that is expensive to redo (the windows and their index minutes, before upload) and before writing the full shortlist (one sample row first). This skill posts and sends nothing; if the user later asks Claude to post or send something, ask again before each write.
- Vivu's banner results are candidates only. The text in the shortlist comes from a frame Claude looked at, never from Vivu's reason text, which invents banner text.
- Times in the shortlist are times in the user's original recording, so the user can jump straight there. A Vivu result page link expires after four hours, so it appears only in the live reply.
What you need before starting
Check each item at the start of the run and tell the user plainly what is missing before doing anything else.
| Requirement | Why | How to check |
|---|---|---|
| Vivu connector with write access | create a private project, open its upload page, search, read summaries | vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the user reconnects Vivu and allows write access |
| The stream recording as a file on the user's computer, or on the user's own YouTube channel | windows are cut from it and every time in the shortlist points into it | ffprobe -v error -show_entries format=duration -of csv=p=0 VOD_FILE prints its length in seconds; for a recording that is only on YouTube, the --print duration command in Step 1 |
| A shell on that computer with ffmpeg and ffprobe | cut windows, extract frames and contact sheets | ffmpeg -version, ffprobe -version |
| A way to pick windows: the chat replay, stream markers or the streamer's notes | indexing a whole stream costs more index minutes than a plan has | ask; for YouTube, Step 2 downloads the chat replay |
| yt-dlp and node, only for YouTube chat replay, captions or window downloads, on a residential IP | YouTube blocks cloud and datacenter IPs, and yt-dlp needs a JavaScript runtime | yt-dlp --version, node --version |
| python3, only for the chat counting script | count chat messages per bucket | python3 --version |
| An upload path | move the window files into Vivu | vivu_open_upload_page plus the user's own browser, or a browser tool that can attach local files |
| Captions or a transcript in the stream's language | stories are checked against what was said; without them stories come back unchecked | Step 1 lists the YouTube captions; for Twitch or a local recording, ask whether the user has a transcript |
This skill needs Claude Code on the user's computer (the terminal or the Code tab of Claude Desktop). Claude on the web and cloud sessions cannot reach the recording. A residential IP matters only for the YouTube downloads. Nothing recurs, so no scheduler is involved; run the skill once per stream.
Inputs to collect
Ask for anything missing, most important first.
- VOD_FILE: the path to the recording. Required, unless the recording exists only on the user's own YouTube channel; then there is no VOD_FILE and Step 4 downloads just the windows. VOD_NAME is VOD_FILE's name (or the YouTube VIDEO_ID) without the extension, spaces or brackets; it names the working folder and the window files. The shortlist keeps VOD_FILE's real file name.
- Where the stream ran (YouTube, Twitch or other) and, for YouTube, the video URL, used for the chat replay and captions. Required.
- The game, and what its banners look like. Default: the Step 6 query, which names places, quests, rewards and level ups and was tested on one game only. For match result screens, kill notices or another game, adapt its examples as Step 6 says and treat the first run as a test.
- Windows. Default: the 3 busiest stretches of chat after the opening, 15 minutes each. The user may name times they remember instead.
- Index minute budget. Default: what vivu_get_usage says remains this month.
- Clip padding. Default: suggested cuts start 5 seconds before a banner appears and end 5 seconds after it goes; stories run from the first to the last caption line of the story.
Files and state
Keep everything in one working folder:
clips-VOD_NAME/
config.json VOD_FILE or the YouTube URL, its length, OFFSET, CAPTION_LANG, the game, windows (start and end second in the recording), Vivu project id
chat.live_chat.json YouTube chat replay, when there is one
chat_buckets.tsv chat messages per bucket (Step 2)
src/ short YouTube downloads (clock check, windows)
captions/ the recording's captions, when there are any
windows/ the cut windows, VOD_NAME_wSTART.mp4 (START = second in the recording)
results/ raw search results, one JSON per query, without the result page link
sheets/ contact sheets and full size frames from Steps 2 and 7
candidates.csv every candidate, its judgment (real, rejected or unchecked) and why
clip_shortlist.csv checked moments only
clip_shortlist.md the same rows sorted by time, then the unchecked candidates under their own heading
state.json steps done, windows uploaded, Vivu video ids, candidate keys judged, moment keys reported
A rerun reads state.json first and skips finished steps. A window already marked uploaded is never uploaded again. A candidate key (video_id, start_ms and field) that is already judged is not judged again, and a moment key (VOD_NAME, field and second in the recording) already in the shortlist is not added twice. The next stream gets its own folder and its own Vivu project.
Step 1: Check the Vivu connector and the setup
Goal: confirm every requirement before spending anything.
- Call vivu_get_account. If the tool does not exist, tell the user to add the Vivu connector in Claude (
https://mcp.vivu.ai/mcp) and stop. If the account does not showcan_create_projects: true, or a later write call fails with "has not granted vivu.write", ask the user to reconnect Vivu and allow write access, then stop until they have. - Run
ffmpeg -versionandffprobe -version, and the ffprobe length command from the table on VOD_FILE. If the recording is only on YouTube, read its length withyt-dlp --js-runtimes node --skip-download --print duration "https://www.youtube.com/watch?v=VIDEO_ID"instead. - For a YouTube stream, run
yt-dlp --versionandnode --version, then list the captions withyt-dlp --js-runtimes node --skip-download --list-subs "https://www.youtube.com/watch?v=VIDEO_ID". The automatic caption whose code ends in -orig (en-orig, de-orig and so on) is the stream's own speech; the others are machine translations of it. Record that code in config.json as CAPTION_LANG. For Twitch or a local recording, ask whether the user has a transcript. - If there are neither captions nor a transcript, tell the user now: story candidates cannot be checked, so they will be listed apart as unchecked, and the story search can be skipped to save its credits. Record their choice.
- For the chat script, run
python3 --version.
In our test run the recording existed only on YouTube and its captions were downloaded directly in Step 7, so the length and caption listing in this step were not exercised in our test run.
Done when vivu_get_account shows can_create_projects: true, ffmpeg and ffprobe print their versions, the recording's length is known, and config.json records CAPTION_LANG (or that there is none, with the user's choice about the story search).
Step 2: Pick the windows worth indexing
Goal: 3 or 4 windows of 15 to 20 minutes where something probably happened, chosen without spending any Vivu allowance.
A whole stream of 5 to 7 hours is 300 to 420 index minutes (estimate), while the Premium plan has 180 index minutes a month. Chat activity is free to read and points at the busy stretches. Confirm the Compliance items with the user before this step.
First line up the clocks. A chat replay, stream markers and captions count seconds from the start of the platform's copy of the stream. If VOD_FILE was downloaded from that same video, or there is no VOD_FILE, OFFSET is 0. A local OBS recording can start earlier or later (recording began before going live, or the stream dropped and restarted), and a gap of more than half a minute would make Step 7 read the wrong captions. So for a local recording measure OFFSET, the second in VOD_FILE minus the second in the YouTube video for the same moment: take one frame of the YouTube video at second 600, then find it in the recording, first at one frame per 10 seconds, then at one per second.
yt-dlp --js-runtimes node --extractor-args "youtube:player_client=web_embedded" -f "bv*[height<=720]+ba/b[height<=720]" --merge-output-format mp4 --download-sections "*600-605" --force-keyframes-at-cuts -P "home:clips-VOD_NAME/src" -o "align_600.%(ext)s" "https://www.youtube.com/watch?v=VIDEO_ID"
ffmpeg -v error -y -i clips-VOD_NAME/src/align_600.mp4 -frames:v 1 clips-VOD_NAME/sheets/align_youtube_600.png
ffmpeg -v error -y -ss 0 -to 1200 -i VOD_FILE -vf "fps=1/10,scale=320:-2,tile=10x12" -frames:v 1 clips-VOD_NAME/sheets/align_coarse.png
ffmpeg -v error -y -ss COARSE_START -to COARSE_END -i VOD_FILE -vf "fps=1,scale=480:-2,tile=5x4" -frames:v 1 clips-VOD_NAME/sheets/align_fine.png
Cell c of the coarse sheet, counted from 0 left to right and top to bottom, is about second 10 times c of the recording. Find the cell with the same game scene, HUD and facecam as the YouTube frame; COARSE_START is that second minus 10 (not below 0) and COARSE_END that second plus 10. Cell j of the fine sheet is second COARSE_START plus j, and OFFSET is COARSE_START plus j minus 600. If no coarse cell matches, second 600 is too plain (a loading screen, a menu) or the gap is over ten minutes: try another YouTube second, or ask the user when they started recording. For Twitch, ask the user to find one moment in both the Twitch VOD and their recording and give both times. Write OFFSET into config.json; every chat, marker or caption time plus OFFSET is a time in the recording.
The clock check was not exercised in our test run, where the windows were cut from the YouTube video itself.
For a YouTube stream, download the chat replay and count text messages per 5 minute bucket. Save the script as clips-VOD_NAME/chat_buckets.py:
import json
import sys
from collections import Counter
path = sys.argv[1]
bucket = int(sys.argv[2]) if len(sys.argv) > 2 else 300
counts = Counter()
with open(path, encoding="utf-8") as f:
for line in f:
try:
replay = json.loads(line).get("replayChatItemAction", {})
except ValueError:
continue
offset = replay.get("videoOffsetTimeMsec")
items = [a.get("addChatItemAction", {}).get("item", {}) for a in replay.get("actions", [])]
if offset is None or not any("liveChatTextMessageRenderer" in i for i in items):
continue
counts[int(offset) // 1000 // bucket] += 1
for b in sorted(counts):
print("{}\t{}\t{}".format(b * bucket, (b + 1) * bucket, counts[b]))
yt-dlp --js-runtimes node --skip-download --write-subs --sub-langs live_chat -P "home:clips-VOD_NAME" -o "chat.%(ext)s" "https://www.youtube.com/watch?v=VIDEO_ID"
python3 clips-VOD_NAME/chat_buckets.py clips-VOD_NAME/chat.live_chat.json 300 > clips-VOD_NAME/chat_buckets.tsv
VIDEO_ID is the 11 character YouTube ID of the stream. Each output line is bucket start second, bucket end second and message count, in YouTube time; add OFFSET to get recording time.
In our test run the chat replay came through yt-dlp and the counts were taken with grep one bucket at a time; the script above was not exercised in our test run.
Pick windows:
- Skip the opening, which is mostly greetings. In our test run the busiest bucket of a stream was its opening.
- Take the busiest buckets after that and widen each into a window of 15 to 20 minutes. Merge windows that overlap.
- Show the user a table (window start and end in recording time, H:MM:SS, and messages) and let them swap in times they remember.
For Twitch, chat exports need a third party tool; use one only if the user already has an export with timestamps, and treat it like the YouTube file. Stream markers from Twitch's video producer, or the streamer's own notes, work as window starts. Neither Twitch path was exercised in our test run.
A busy chat is a cheap first filter, not proof. Whether chat-busy windows hold more clip moments than quiet ones was not tested. If a stream's chat is thin (the chat ran on another platform), use markers or notes.
Done when config.json records OFFSET and lists the windows the user approved, each with its start and end second in the recording.
Step 3: Price the windows and get approval
Goal: an indexed set the user's plan can pay for.
Sum the window lengths and call vivu_get_usage. Show one table:
| This stream | Plan allowance | Remaining this month | |
|---|---|---|---|
| Index minutes | sum of the window lengths | from vivu_get_usage | from vivu_get_usage |
| Search credits | 10 per stream, or 5 if the story search is skipped, plus 5 for each rewording or rerun (all estimates) | from vivu_get_usage | from vivu_get_usage |
Plan facts: Free is $0 a month with 20 index minutes a month and 50 search credits a month; Premium is $30 a month with 180 index minutes and 500 search credits. A precise search uses 5 credits in total, and this skill runs 2 precise searches per stream. Whether a larger maximum_results costs more credits is untested.
If the windows do not fit, offer levers in this order: fewer windows, shorter windows, a later stream for the rest of this one, and last, for a Free account, the Premium plan. Premium is the top plan, so past its allowance only the first three levers apply. Never drop a window the user asked for without saying so.
The internal account used in our test run returns no allowance figures, so the approval table and the user's yes were not exercised in our test run.
Done when the user approves the windows and the minute total, and config.json records both.
Step 4: Cut the windows from the recording
Goal: one file per window whose second 0 is a known second in the recording.
ffmpeg -v error -ss START -to END -i VOD_FILE -vf "scale=-2:'min(720,ih)'" -c:v libx264 -preset veryfast -crf 23 -c:a aac clips-VOD_NAME/windows/VOD_NAME_wSTART.mp4
ffprobe -v error -show_entries format=duration -of csv=p=0 clips-VOD_NAME/windows/VOD_NAME_wSTART.mp4
START and END are seconds in the recording. Re-encoding makes the cut exact, so every Vivu time plus START is a time in the recording. Copying streams (-c copy) snaps the cut to a keyframe and shifts every time after it. The scale filter brings a 1080p or 1440p recording down to 720p: the windows are only for indexing (the streamer cuts clips from the original), 720p keeps even a small notice at the bottom of the screen readable on a full size frame, and a 15 minute window at full resolution can run to gigabytes. Vivu's size limit for one uploaded file is unknown. Keep START in the file name: Vivu replaces spaces and brackets in file names with underscores, and the wSTART part is how results are matched back.
If there is no VOD_FILE (the recording exists only on the user's own YouTube channel), download just the windows, padded by 10 seconds, instead of the whole stream, then cut from the padded file with -ss 10. Times are then YouTube times and OFFSET is 0:
yt-dlp --js-runtimes node --restrict-filenames --extractor-args "youtube:player_client=web_embedded" -f "bv*[height<=720]+ba/b[height<=720]" --merge-output-format mp4 --download-sections "*PADDED_START-PADDED_END" --force-keyframes-at-cuts -P "home:clips-VOD_NAME/src" -o "%(id)s_section_PADDED_START.%(ext)s" "https://www.youtube.com/watch?v=VIDEO_ID"
ffmpeg -v error -ss 10 -to WINDOW_SECONDS_PLUS_10 -i clips-VOD_NAME/src/VIDEO_ID_section_PADDED_START.mp4 -c:v libx264 -preset veryfast -crf 23 -c:a aac clips-VOD_NAME/windows/VOD_NAME_wSTART.mp4
PADDED_START is START minus 10, PADDED_END is END plus 10, and WINDOW_SECONDS_PLUS_10 is END minus START plus 10. The youtube-competitor-watch skill (https://vivu.ai/skills/youtube-competitor-watch) covers channel verification and download troubleshooting in more depth.
In our test run both windows came out at 570.0 seconds after the cut. They were cut on the YouTube path; the local cut with the scale filter was not exercised in our test run.
Done when windows/ holds one file per approved window and each ffprobe length equals END minus START.
Step 5: Upload and index
Goal: every window indexed in a private Vivu project for this stream.
- Call vivu_list_projects and reuse a project named "Clips VOD_NAME" if one exists. Otherwise call vivu_create_project with that name and visibility "private". The default visibility is organization, which shows the project to everyone in the user's Vivu organization. Use one project per stream: a search always covers the whole project, and older windows would take up result slots.
- Call vivu_open_upload_page with the project ID immediately before uploading. The link expires in 180 seconds and works once, so never post or store it. Give it to the user to open in their own browser and select the files in windows/, or open it in a browser tool that can attach local files. Claude in Chrome accepts at most 10 MB per upload call, and window files are far larger, so they normally go through the user's own browser or the Vivu web app. Never split a window or recompress it again to fit.
- Poll vivu_list_videos until every window shows ready. Match each Vivu file name back to its window by the
wSTARTpart and record the video_id in state.json.
In our test run the 2 windows (19.0 minutes) were ready 561 seconds after the upload started. They went through the same upload page with an automated browser; opening the link in the user's own browser was not exercised in our test run.
Done when every window shows ready and state.json maps each window file to its Vivu video_id.
Step 6: Search the two moment types
Goal: candidate moments for the shortlist.
| Field | Query | Mode | maximum_results | Role |
|---|---|---|---|---|
| game_popup | a wide banner or notice box fixed to the top or bottom edge of the screen, on a dark strip or in a framed box, announces a place, a quest, a reward or a new level for a few seconds; text floating over a character's head does not count | precise | 100 | candidates only; Step 7 checks every one on frames |
| story | the streamer tells a story about something that happened to them before, from how it started to how it ended | precise | 20 | checked against captions or a transcript in Step 7; skipped if the user chose so in Step 1 |
The chain: game_popup finds what the game showed on screen; story flips to what the streamer said; Step 7 checks each against the frames or the captions before anything is listed. Both are precise because only precise returns a time window; fast returns whole files with an empty reason, and with a handful of windows there is nothing for fast to narrow.
maximum_results is also the recall ceiling.
In our test run both searches ran with maximum_results 20 on 19.0 minutes of windows, and game_popup returned 18 candidates.
A full set of windows is 3 to 4 times as long, so game_popup starts at 100, the upper limit, rather than risk a rerun; expect Step 7 to check several times as many candidates (estimate). If it still returns exactly 100, tell the user that some banners may be missing from the candidates and suggest fewer windows per project next time. story starts at 20 because stories are rarer; if it returns exactly 20, raise it to 100 and rerun before Step 7.
In our test run the story search returned 2 moments against its maximum_results of 20.
The game_popup wording was tuned and tested on one game, whose banners sit in fixed places at the top or bottom of the screen and name a place, a quest, a reward or a reputation level. Match result screens, kill notices, victory banners and other games were not in our test corpus. For another game, swap the query's examples for that game's own and tell the user the first run is a test: an adapted wording is untested until its first candidates have been through Step 7, and Step 7 is what keeps its mistakes out of the shortlist.
Run each query with vivu_search_videos (project_id, query, mode, maximum_results). It returns a job ID. Call vivu_get_search_results until complete is true; each status call can wait up to 45 seconds, so a pending search is not a stalled one. Save each completed result to results/FIELD.json with its result_page_url field removed (FIELD is game_popup or story). Show the Vivu result page link in the live reply only; it expires after four hours, so it never goes into a file or the shortlist.
Worked example from our test run
The test corpus was public: 2 videos (19.0 minutes), each a 570.0 second window cut from a public livestream archive of a multiplayer game, with the players' facecams and voice chat.
Each window was the busiest stretch of its stream's chat replay after the opening. The skill's windows are longer; these were 9.5 minutes because of our test budget. Banners and stories were marked by hand before any search.
- game_popup, with the wording in the table after 2 rewordings on the same corpus: 18 candidates, 6 real and 12 false, a false positive rate of 67%, and no banner marked by hand was missed.
- The two earlier game_popup wordings also missed no banner marked by hand and returned slightly fewer false candidates: 16 returned with 10 false, then 14 returned with 8 false. More than half of the candidates were false in every wording, so none of them can skip Step 7. The table keeps the last wording because the frame checks and the shortlist of this test run were made from its candidates; the second wording would have left 4 fewer candidates to check.
- Both searches ran with maximum_results 20, not the larger value in the table.
- The false game_popup candidates were name tags floating over other players, an in-game book page, a crew chat line and windows with no text at all. In each of the 12 false windows Vivu's reason described a banner, sometimes in the query's own words.
- Step 7's contact sheets showed the banner in each of the 6 real candidates and in none of the 12 false ones, so only banners seen in frames reached the shortlist. The text of all 6 was read from full size frames.
- The windows held no match result screens, kill notices or victory banners, so the result screen rule in Step 7 was not exercised in our test run.
- story, first wording with 0 rewordings: 2 moments, 2 real, 0 false, and 2 missed of the 4 stories marked in the captions. The missed ones were remarks of a sentence or so. A game character's voiced backstory and a host reading an in-game log aloud were not returned, which is right for this field.
- Laughs and shouts were tested and dropped: the first wording returned 13 moments with 5 real and 8 false, and a second wording returned 13 with 3 real, 10 false and 3 missed.
- Indexing took 561 seconds from the start of the upload until both windows were ready, and the shortlist had 8 rows.
Done when both searches are complete (or story was skipped by the user's choice) and saved in results/, and the user has been told if either returned exactly its maximum_results.
Step 7: Check every candidate
Goal: every candidate judged real, rejected or unchecked against the recording, with the frame or caption that shows it.
For each game_popup candidate, make one contact sheet of its window at one frame per second:
ffmpeg -v error -y -ss HIT_START -to HIT_END -i WINDOW_FILE -vf "fps=1,scale=480:-2,tile=4x5" -frames:v 1 clips-VOD_NAME/sheets/game_popup_hN_every_second.png
HIT_START and HIT_END are the result's start_ms and end_ms divided by 1000, seconds inside the window file. WINDOW_FILE is the window in windows/, and N numbers the candidate. Cell k, counted left to right and top to bottom from 0, is window second HIT_START plus k. For a hit longer than 20 seconds use tile=4x8, and for one longer than 32 seconds make a second sheet that starts at HIT_START plus 32. A single frame is not enough: the hit is wider than the banner, which can sit in only a few of its seconds.
Look at the sheet. A candidate is real when a frame shows a banner or notice drawn on the game's interface: it sits fixed at the top, bottom or centre of the screen, often in a frame or with an emblem, for a few seconds while the scene moves behind it. Reject a name tag or title floating over a character (it moves with the character), a book or letter the player is reading, a map, a menu or inventory, a chat or quick chat line, a button prompt, subtitles of game characters, and a window with no text.
A full screen result screen (victory or defeat, match over, a final scoreboard or placement) is real too, even though it can be still and fill the screen like a menu. What tells them apart: a result screen appears on its own when a match or round ends and names an outcome, a score or a placement; a menu, inventory or map is opened by the player and lists items, settings or places. When Claude cannot tell which one a frame shows, the candidate is unchecked, not real.
Our test corpus had no result screens, so the result screen rule was not exercised in our test run.
For a real banner or result screen, extract one full size frame at a second where it is on screen and read the text there:
ffmpeg -v error -y -ss SECONDS -i WINDOW_FILE -frames:v 1 -q:v 3 clips-VOD_NAME/sheets/game_popup_hN_read.png
SECONDS is the window second from the sheet. Small notices at the bottom of the screen are unreadable on the sheet and readable on the full frame, so every real banner gets one. Write the text as it appears. Never copy it from Vivu's reason, which names banners that are not on screen. If the text cannot be read, write UNREADABLE and keep the row. Note the first and last second the banner is visible; they set the suggested cut.
For each story candidate, check the words:
- With YouTube captions (CAPTION_LANG from Step 1), download them once and read the lines around the candidate, starting half a minute before its hit, because a hit can start after the story does:
yt-dlp --js-runtimes node --skip-download --write-auto-subs --sub-langs "CAPTION_LANG" --sub-format vtt -P "home:clips-VOD_NAME/captions" -o "%(id)s.%(ext)s" "https://www.youtube.com/watch?v=VIDEO_ID"
grep -n "^HH:MM:S" clips-VOD_NAME/captions/VIDEO_ID.CAPTION_LANG.vtt
HH:MM:S is the YouTube time of the hit (the window's START plus HIT_START, minus OFFSET) with two digit hours and without the last digit of the seconds, so 0:22:41 becomes 00:22:4. Read the file from the line grep prints. A story is real when the speaker tells something that happened before, with a start and an end. Reading game text aloud, explaining rules and describing what is happening now do not count. The suggested cut runs from the first to the last caption line of the story, converted back to recording time by adding OFFSET.
- With a transcript the user already has, or one from a local speech to text tool, read it the same way; that path was not exercised in our test run.
- With neither, the story cannot be checked: mark it unchecked. vivu_get_video_summary with include_segments true and start_ms and end_ms around the candidate can say what that stretch is about, but a summary that mentions a story does not prove it, and one that does not mention it does not rule it out. Unchecked stories never go into clip_shortlist.csv; Step 8 lists them apart so the streamer can listen first.
- Paraphrase what was said from the captions. Quotes come only from captions or a transcript, never from Vivu's reason, which paraphrases.
Write every candidate to candidates.csv with its judgment (real, rejected or unchecked), the reason and the sheet, frame or caption lines behind it. Rejected and unchecked candidates stay there and never reach clip_shortlist.csv; tell the user how many were rejected and how many are unchecked per field.
In our test run a story hit started after the story began, and the caption lines from before the hit gave the right start. One of the 2 summary segments around a real story did not mention it.
Done when every row of candidates.csv has a judgment of real, rejected or unchecked, and each real banner names a sheet and a readable full size frame file that exist.
Step 8: Build the shortlist
Goal: a shortlist the user can cut from, in their recording's own times.
Convert each real candidate: recording second = the window's START plus the window second from Step 7, written H:MM:SS. Suggested cuts use the padding from Inputs. Show the user one sample row and the field mapping, and wait for an OK before writing the rest:
vod_file,vod_time,field,screen_text,what_happens,suggested_in,suggested_out,checked_with,evidence
"2026-09-20 20-58-03.mkv",0:16:58,game_popup,"QUEST COMPLETE","banner at the top of the screen right after the crew hands in a quest",0:16:53,0:17:05,frames,sheets/game_popup_h15_read.png
| Column | Source | If unavailable |
|---|---|---|
| vod_file | VOD_FILE's name as it is on disk, or the YouTube URL when there is no VOD_FILE | none; every row has it |
| vod_time | window START plus the second the banner appears or the story starts | drop the row |
| screen_text | the full size frame from Step 7 (banners only) | UNREADABLE |
| what_happens | Claude's description of the frames, or a paraphrase of the captions (inferred) | NOT VERIFIED |
| suggested_in, suggested_out | banner seconds from the sheet plus padding, or first and last caption line | the window's start and end |
| checked_with | frames, captions or transcript | not in this file: unchecked candidates go under the md's own heading |
| evidence | the sheet, frame file or caption times | none |
what_happens is Claude's reading of the frames and captions. If it is wrong, the streamer opens the wrong moment, so every row keeps its time and evidence file to check against.
Then write clip_shortlist.md, sorted by time, one line per row, with the unchecked candidates after the shortlist under their own heading. DATE is the stream date, N and M the windows indexed and their minutes, K the candidates checked, R the rejected ones and U the unchecked ones:
# Clip shortlist: VOD_NAME (DATE)
Windows indexed: N, M minutes. Candidates checked: K, rejected: R, unchecked: U.
- 0:16:58 game_popup: "QUEST COMPLETE" banner at the top after the crew hands in a quest. Cut 0:16:53 to 0:17:05. (frames)
- 0:22:41 story: how the first ranked match of the season went wrong (paraphrase). Cut 0:22:41 to 0:23:44. (captions)
## Not checked: look or listen before cutting
- 1:04:05 to 1:04:40 story candidate: no captions or transcript for this stretch.
Record every reported moment key in state.json. If fewer moments were found than the user hoped, report the real count; never pad the list with rejected candidates or unchecked ones.
Done when clip_shortlist.csv and the shortlist part of clip_shortlist.md hold the same rows, unchecked candidates appear only in candidates.csv and under the md's "Not checked" heading, the user approved the sample row and the mapping, and state.json lists every reported moment key.
Compliance
- This skill is for the streamer's own recordings. Downloading someone else's stream, chat replay or captions may conflict with the platform's terms of service; do not do it. Have the user confirm the recording is theirs before Step 2.
- Friends in voice chat and guests on camera are in the recording too, and Vivu indexes their voices. Before a clip with them is posted anywhere, the streamer asks them. The skill does not identify anyone.
- If the streamer or anyone in the recording is a minor, a parent or guardian must agree first; otherwise stop.
- The game footage belongs to its publisher. Most publishers allow stream clips, but the rules differ, so the streamer checks them before posting.
- The windows stay in the user's Vivu project until the user deletes them. Create the project as private; the default is visible to the whole Vivu organization. Delete videos or the project only when the user asks, and confirm first.
- The skill writes only local files. It does not post, upload clips or message anyone, and it does no face, logo or identity recognition.
Known failure modes
| Symptom | Cause | Fix |
|---|---|---|
| (observed) most banner candidates are not banners: 18 were returned, 12 false and 6 real, and none of the known banners was missed | the search matches short text anywhere in the scene | treat every game_popup result as a candidate and check it on a sheet (Step 7) |
| (observed) a banner candidate is a name tag floating over another player | name tags look like a name with a title line | reject text that moves with a character; a banner stays fixed on screen |
| (observed) a banner candidate is an in-game book page, a map or a crew chat line | readable text of any kind matches | reject it on the sheet |
| (observed) the reason places a banner "at the bottom of the screen on a dark strip" that is not in any frame | the reason reuses the query's own words | read banner text only from frames |
| (observed) a story clip starts in the middle of the story | the hit began after the story started | read the captions from half a minute before the hit |
| (observed) the summary segment around a real story "does not mention the story" | segments are chapter level summaries | use captions or a transcript; without them list the story as unchecked, apart from the shortlist |
| (observed) a story told in a sentence or so is not returned | the query asks for a story with a start and an end | read the captions around each window if the user wants every mention |
| (observed) a search for laughs or shouts returns quiet smiles and normal talk | in our test run the search did not tell loud reactions from normal talk | this skill does not search for reactions; use stream markers |
| (observed) "HTTP error 403 Forbidden" and "ERROR: ffmpeg exited with code 8" when downloading a window from YouTube | YouTube refused the yt-dlp client | use the web_embedded player client as in Step 4 |
| a small notice is unreadable on the contact sheet | the sheet scales each frame down | extract a full size frame at that second |
| the chat replay shows no clear peaks | the chat ran on another platform | use stream markers or the streamer's notes |
| upload page asks to sign in or shows an error | the one time upload link expires after 180 seconds | request a new link right before opening it |
| a window file is rejected by the browser upload tool | the tool accepts at most 10 MB per call | the user adds the file in their own browser or the Vivu web app |
| "has not granted vivu.write" | Vivu connected read only | the user reconnects Vivu with write access |
| a result's file name does not match windows/ | Vivu replaced spaces or brackets with underscores | match on the wSTART part of the name |
| an empty game_popup result | the windows have no banners, or the wording does not fit this game | an empty result does not prove there are none; adapt the examples in the query and check the new wording's candidates |