---
name: event-recap-selects-sheet
description: "Turn a client's conference or awards night camera rushes into a recap selects sheet with Vivu: name reads, audience questions, audience cutaways and big screen text, each checked on frames."
---Event recap selects sheet from camera rushes
This skill turns a client event's camera files and its agenda into a selects sheet for the recap edit. Claude cuts only the agenda blocks a recap uses (opening, keynote, awards, Q&A) out of each camera file, prices them against the user's Vivu plan, uploads and indexes them in a private Vivu project, runs four standing searches (a host reading out a name, an audience member asking a question at a microphone, audience cutaways, big screen titles and numbers), checks every hit on extracted frames and the local transcript, and writes selects.csv: one row per moment with camera file, source timecode, what the picture shows, a frame, the words said (from the transcript), the screen text (read from the frame) and an empty "use" column for the editor.
The value is in what the transcript does not hold. A transcript says "our next awardee is NAME"; it does not say which camera caught the walk-up and the handshake, when the director cut to the room, or that a student stood at the floor mic to ask a question. Those moments are what make a recap feel like it was there. Every row on the sheet has been looked at on a frame before it is written, and audience cutaways, the weakest search, stay marked as candidates until Claude has checked them.
When to use
Use when an editor or producer says "pull selects from the event rushes", "find the award moments and audience shots for the recap", "log the Q&A questions with timecodes", or "make a selects sheet from the conference cameras". For a single question about one clip ("where does the CEO mention pricing?"), search Vivu directly instead. For per attendee recap emails from product sessions, this skill is the wrong fit; it produces an editor's sheet, not messages.
Working principles
- Report measured numbers, not estimates. When a number is an estimate, say so.
- Nothing goes on the sheet as a select until it has been checked against the source frames or the local transcript. Everything else stays a candidate, labeled as one.
- Stop and tell the user when a required capability or file is missing (no transcript for a clip, no ffmpeg, no write access in Vivu). Do not guess around it.
- Ask the user before anything that is expensive to redo or that acts on their behalf: the agenda cut list before cutting, the indexing cost before upload, the sample row before the full sheet.
- Names on the sheet come only from the agenda or from what the host says in the transcript. Claude never identifies people from their faces.
What you need before starting
Check each item at the start of the run and tell the user plainly what is missing.
| Requirement | Why | How to check |
|---|---|---|
| Vivu connector with write access | create a private project, open its upload page, search | vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the connection is read only; the user reconnects Vivu and allows write access |
| A shell with ffmpeg and ffprobe on the machine that holds the camera files | cut agenda blocks, measure minutes, extract frames, measure loudness | ffmpeg -version and ffprobe -version |
| The camera files on local disk, one folder per camera | the sheet points at the original files; Vivu has no export tool, so frames and clips come from the local files | ls CAMERA_FOLDER |
| The agenda with clock times, and the time of day each camera started recording | picks the blocks to index and maps agenda times to file offsets | the user pastes the agenda; ffprobe -v error -show_entries format_tags=creation_time -of csv=p=0 FILE, or the user gives a slate time |
| A transcript with timestamps (SRT or VTT) for each cut block | quotes and names on the sheet come from it, never from Vivu's reason text | ls BLOCK_FOLDER for one .srt per .mp4; the editing app's transcription export or any local speech to text tool can make it |
| A browser that can open the Vivu upload page | the connector has no direct upload tool | the user opens the link, or Claude in Chrome is connected |
No residential IP is needed (nothing is downloaded from YouTube) and nothing is scheduled.
Inputs to collect
Ask for anything missing, most important first.
- The camera folders and which camera is which (for example CAM_A stage, CAM_B wide, CAM_C audience). Required.
- The agenda with clock times. Required.
- Which agenda blocks the recap needs. Default: opening, keynote, awards, Q&A.
- Names the recap must include (award winners, speakers). Default: every name the host reads out in the chosen blocks.
- How many audience cutaways per block. Default: up to five, best framed first.
- The Vivu project name. Default: "EVENT_NAME recap selects", private.
Files and state
Keep everything in one working folder next to the rushes:
recap-selects/
config.json event name, camera map, agenda blocks with clock times, camera start times, vivu project id
blocks.csv one row per cut block: block file, camera file, source start seconds, duration
blocks/ cut blocks (.mp4) and their transcripts (.srt)
results/ raw search results, one JSON file per query
frames/ extracted frames and contact sheets
loudness/ ebur128 output per block
selects.csv the sheet
state.json uploaded block names and video ids, queries run with job ids, hits checked with verdicts, rows written
state.json is updated after every step. A rerun reads it first: blocks already uploaded are not uploaded again, queries already run are not rerun, and hits already checked keep their verdicts, so an interrupted run picks up where it stopped and no moment is written twice.
Step 1: Check the Vivu connector and the setup
Goal: every row of What you need is confirmed before anything is cut or uploaded.
- Call vivu_get_account. If the tool is missing, tell the user to add the Vivu connector in Claude (https://mcp.vivu.ai/mcp) and stop. If can_create_projects is not true, or a write call later fails with "has not granted vivu.write", ask the user to reconnect Vivu with write access and stop until they have.
- Run ffmpeg -version and ffprobe -version. List the camera folders.
- Ask for the agenda and the camera start times if they are missing.
Done when vivu_get_account shows can_create_projects: true, ffmpeg and ffprobe print versions, and config.json holds the camera map and the agenda.
Step 2: Cut the agenda blocks
Goal: one file per camera per agenda block, so only the minutes the recap needs get indexed.
- For each chosen block and each camera, compute START and END in seconds from the start of that camera file (agenda clock time minus the camera's start time; add a 30 second margin on each side). Show the cut list to the user and wait for a yes.
- Create the folders first; ffmpeg does not create an output folder and stops with "Error opening output files: No such file or directory". Then cut with a re-encode, so the block starts exactly at START and every timecode on the sheet maps back to the camera file:
mkdir -p recap-selects/blocks recap-selects/frames recap-selects/results recap-selects/loudness
ffmpeg -v error -ss START -to END -i CAMERA_FILE -c:v libx264 -crf 20 -c:a aac recap-selects/blocks/CAM_BLOCK.mp4
The rest of this skill runs from inside recap-selects/, so later paths start at blocks/.
CAMERA_FILE is the original file, CAM_BLOCK a short name such as CAMB_awards, START and END are seconds. Keep names to letters, digits and the "_" character; Vivu rewrites spaces and punctuation in file names, and plain names match back without guessing. 3. Write each block to blocks.csv with its source start seconds. A time on the sheet is source start plus the offset inside the block. 4. Put the transcript for each block next to it as blocks/CAM_BLOCK.srt. If the transcript covers the whole camera file, shift it by START.
In our test run the blocks came from public recordings rather than camera cards, and their transcripts were auto subtitles; the cut command above was run once on one block to confirm it, and cutting from camera cards with an editing app transcript was not exercised in our test run.
Done when blocks/ holds one .mp4 and one .srt per approved block and blocks.csv has one row per block.
Step 3: Price the indexing and get approval
Goal: the user sees what the blocks cost before anything is uploaded.
- Measure each block: ffprobe -v error -show_entries format=duration -of csv=p=0 blocks/CAM_BLOCK.mp4, and sum the minutes.
- Call vivu_get_usage for the plan and what remains this month.
- Estimate search credits: four standing queries per project, each precise. A precise search uses 5 credits in total and a fast search uses 1 credit. Label this as an estimate, since reruns after rewording add to it.
- Show one table and wait for approval:
| This event | Remaining on the plan | |
|---|---|---|
| Index minutes | measured sum | from vivu_get_usage |
| Search credits | estimate | from vivu_get_usage |
For reference, the Free plan has 20 index minutes a month and 50 search credits a month; Premium is $30 a month with 180 index minutes and 500 search credits. If the blocks do not fit, offer these levers in order: index one camera per block instead of all cameras (the wide or audience camera for cutaways, the stage camera for name reads), drop blocks the recap can live without, shorten the margins, and only then a larger plan. Never drop material the user asked for without saying so.
In our test run the account was an admin account whose vivu_get_usage shows no remaining allowance, so the comparison against a real plan and the approval were not exercised in our test run.
Done when the user has approved the minutes and the credit estimate.
Step 4: Upload and index
Goal: every approved block is ready in a private Vivu project.
- Call vivu_list_projects and reuse the project named in config.json if it exists. Otherwise call vivu_create_project with that name and visibility "private". Client footage belongs in a private project; the default visibility is the whole organization.
- Call vivu_open_upload_page with the project ID right before the upload. The link expires in 180 seconds and is a sign in link: do not paste it into any message or file. Give it to the user to open in their own browser and choose the files in blocks/, or attach the files with a browser tool that can upload local files. Claude in Chrome accepts at most 10 MB per upload call, and camera blocks are usually larger; the user adds those in the Vivu web app. Never split or recompress a block to fit a tool limit.
- Poll vivu_list_videos about every 30 seconds until every block shows ready. Match each video back to blocks.csv by file name; Vivu turns spaces and punctuation into underscores. Record the video IDs in state.json.
Done when vivu_list_videos shows every block in blocks.csv as ready.
Step 5: Run the standing searches and check every hit
Goal: four lists of checked moments, one per sheet category.
| Field | Query | Mode | maximum_results |
|---|---|---|---|
| name_read | the host at the lectern announces the next speaker or award winner by name and calls them to the stage | precise | 20 |
| audience_question | a person in the audience is talking into a floor microphone, asking the speaker a question out loud in their own voice | precise | 20 |
| audience_shot (candidates) | a shot of the audience: rows of seated people seen from the front or side, listening or clapping, not the stage | precise | 20 |
| screen_text | a large screen or full-screen graphic showing a title, a person's name or a big number in large letters | precise | 20 |
The audience_question wording asks for someone talking into the mic because the shorter wording ("stands at a microphone and asks the speaker a question") also returned a reaction shot of a questioner listening to the answer.
Why this chain: name_read gives the recap its spine, and each hit points at the walk-up and the handshake that follow it (Step 6 finds them). audience_question does the same for the Q&A block, from what is said to the person seen at the mic. audience_shot fills cutaways for the same blocks, and screen_text gives title cards and big numbers for lower thirds and openers. All four are precise because the sheet needs time ranges; fast only returns whole files. maximum_results 20 is above the number of name reads or questions in a typical block, so a list that stops at exactly 20 means the cap was hit: raise it and rerun that query.
Run each query with vivu_search_videos (project_id, query, mode, maximum_results). It returns a job ID. Call vivu_get_search_results until complete is true; each status call can wait up to 45 seconds, so a pending search is not a stalled one. Save each completed result as results/FIELD.json. Show the result page link in the live reply only; it expires after four hours, so it never goes into selects.csv or any file.
Check every hit before it goes anywhere:
- Extract three frames inside the window (start, middle, end):
ffmpeg -v error -ss SECONDS -i blocks/CAM_BLOCK.mp4 -frames:v 1 -q:v 3 frames/FIELD_hN_POS.png
SECONDS is start_ms / 1000, the midpoint, or end_ms / 1000 minus 0.5; N is the hit number; POS is start, mid or end. 2. For name_read and audience_question, read the transcript lines inside the window. The hit is real only if the transcript has the host announcing the next person or the award (name_read), or a person in the room asking a question (audience_question). The reason text is a paraphrase, not a transcript: quotes and names on the sheet come from the .srt. 3. For audience_shot, the hit is real only if the frames show seated audience members as the subject, seen from the front or side. A stage wide shot from the back of the room with heads in the foreground, a close-up of one honoree with a blurred crowd behind, and a single questioner at the mic are not audience shots. 4. For screen_text, the hit is real only if a projection screen, a video wall or a full frame graphic is in the picture; Vivu also returns lectern signs and printed backdrops. Read the words and numbers from the full frame, never from the reason. If a frame is too small to read, extract the full resolution frame again at the same time. 5. Record every verdict in state.json. Rejected hits stay in results/ and are counted, not written to the sheet. 6. Cross-check the name reads. For every name on the agenda without a checked name_read hit, look for a screen_text hit that shows that name on a name card, and read the transcript around it. Shows that play a name card while the host speaks off camera are where the name_read search is most likely to miss.
Vivu returns a time range that contains the moment, not an exact frame; in our test run the median window was 13 seconds for name reads and 38 seconds for audience questions. Before writing a row, narrow start and end to what the frames and transcript show.
audience_shot is the weakest search here. Vivu reads composition (who fills the picture, what is in focus) loosely: a stage wide shot taken from behind the seats, or a close-up of one honoree with the crowd blurred behind, often comes back described as a clear audience shot. Treat every audience_shot result as a candidate, look at its frames, and keep only windows where seated people are the subject. A window can hold a stage shot and an audience cutaway back to back; write the cutaway's own seconds on the sheet, not the whole window. An empty or short list does not prove the footage has no audience shot; when a block needs more cutaways, make a two frames a second contact sheet of the audience camera block and pick by eye.
Worked example from our test run
In our test run we used 4 public event blocks, 32 minutes in total: the opening and a run of award calls from an engineering awards dinner, and the opening and part of the audience Q&A from a university forum, both published by the institutions. Both are switched programs, so the director's audience cutaways stood in for an audience camera. From the first upload to the last block ready took 13 minutes.
The name_read query returned 7 moments, all 7 real on the transcript, and missed 3 of 10 name reads on our list written before searching. All 3 misses were in the awards block, where the host speaks over a full frame name card, and the screen_text query returned the name card for each of those 3.
The audience_question query returned 3 moments, 3 real, 0 false and 0 missed, with the wording we settled on after 1 rewording on the same corpus. The first wording returned 4 moments with 1 false: a questioner shown listening during the answer.
The audience_shot query returned 9 moments: 3 real and 6 false, and 0 missed against our list of cutaways. We wrote that list before searching from small contact sheets, then removed two entries from the awards dinner after the search, because full size frames showed an honoree close-up with the crowd out of focus; with the original list the count would have been 4 real, 5 false and 1 missed. All 6 false ones came from the ballroom event (stage wide shots from the back of the room, honoree close-ups with the crowd out of focus, one walk-up past the side tables).
The screen_text query returned 12 moments: 9 real (name cards and title slides) and 3 false (a lectern sign and a printed stage backdrop described as a screen). The corpus had no big numbers on screen, so numbers were not tested.
The upload in the test run did not go through a user's browser or Claude in Chrome; that path was not exercised in our test run.
Done when every hit of every standing query has a verdict in state.json and the user has seen the true and false counts per query.
Step 6: Find the handshake, the applause and the best frame locally
Goal: each checked name read has the walk-up or handshake frame and the applause stretch next to it.
- For each checked name_read hit, first look at the 20 seconds after the window starts, two frames a second, on every camera that has a block for that agenda time (the same moment on CAM_B is at the same clock time, so shift by the difference in source start from blocks.csv):
ffmpeg -v error -ss SECONDS -t 20 -i blocks/CAM_BLOCK.mp4 -vf "fps=2,scale=320:-2,tile=8x5" frames/handoff_hN_CAM.png
If there is no walk-up in those 20 seconds, the show probably plays a tribute video or a bio between the name and the handshake. Then cover the whole stretch up to the next name read, one frame a second (DURATION is the next name read's start minus SECONDS):
ffmpeg -v error -ss SECONDS -t DURATION -i blocks/CAM_BLOCK.mp4 -vf "fps=1,scale=240:-2,tile=10x8" frames/handoff_hN_CAM_long.png
Pick the frame where the person reaches the stage, shakes hands or takes the award, and extract it at full resolution for the sheet. Keep the quotes around the filter; some shells treat the comma list as a pattern otherwise. 2. Measure short term loudness per block, ten values a second:
ffmpeg -v error -i blocks/CAM_BLOCK.mp4 -vn -af "ebur128=metadata=1,ametadata=mode=print:key=lavfi.r128.S:file=loudness/CAM_BLOCK.txt" -f null -
- Find applause with the transcript, not with loudness alone. Save this as applause.py and run python3 applause.py loudness/CAM_BLOCK.txt blocks/CAM_BLOCK.srt NAME_READ_SECONDS:
import re, sys
loud_file, srt_file, t0 = sys.argv[1], sys.argv[2], float(sys.argv[3])
times, vals = [], []
for line in open(loud_file):
m = re.search(r"pts_time:([\d.]+)", line)
if m:
times.append(float(m.group(1)))
m = re.search(r"lavfi\.r128\.S=(-?[\d.]+)", line)
if m:
vals.append(float(m.group(1)))
def secs(ts):
h, m, s = ts.replace(",", ".").split(":")
return int(h) * 3600 + int(m) * 60 + float(s)
speech = []
for m in re.finditer(r"(\d+:\d+:\d+[.,]\d+) --> (\d+:\d+:\d+[.,]\d+)\s*\n(.+)", open(srt_file).read()):
text = m.group(3).strip()
if text and not text.startswith("["):
speech.append((secs(m.group(1)), secs(m.group(2))))
def talking(t):
return any(a <= t <= b for a, b in speech)
quiet = [(t, v) for t, v in zip(times, vals) if t0 <= t <= t0 + 25 and not talking(t)]
if not quiet:
print("no gap without speech in the 25 s after the name read")
sys.exit()
start = end = quiet[0][0]
best = (0, start, end)
for (a, _), (b, _) in zip(quiet, quiet[1:]):
if b - a > 0.15:
start = b
end = b
if end - start > best[0]:
best = (end - start, start, end)
seg = [v for t, v in quiet if best[1] <= t <= best[2]]
speech_vals = [v for t, v in zip(times, vals) if t0 - 10 <= t < t0 and talking(t)]
print("longest stretch without speech: %.1f-%.1f s, loudness %.1f LUFS vs speech %.1f LUFS"
% (best[1], best[2], sum(seg) / max(len(seg), 1), sum(speech_vals) / max(len(speech_vals), 1)))
The stretch with no transcript words right after a name read is where the applause is. Loudness tells you how the room sounds in that stretch: on a feed from the room microphones applause is louder than speech, on a desk mix from the lectern microphone it can be quieter. Look at the handoff contact sheet for the same seconds (hands clapping, people standing) before writing the applause times on the sheet. 4. Record in state.json the handoff frame and the applause start and end for each name read.
In our test run the handoff contact sheets and the loudness measurement were run; applause.py itself was not exercised in our test run (the same comparison was done by reading the loudness output by hand). In the forum the applause after the speaker introduction was the loudest stretch of the block; at the awards dinner the applause after each call was quieter than the host's voice, and the handshake came after a tribute video, well past the short contact sheet.
Done when every checked name_read row has a handshake frame or the note "no handoff on this camera", and an applause start and end or "not measured".
Step 7: Assemble selects.csv and confirm the sample
Goal: the editor gets one sheet in recap order.
- Build one row per checked moment, ordered by agenda block and then time:
segment,camera_file,start_mmss,end_mmss,moment_type,picture,frame,said_local,screen_text,use
Awards,CAMB_awards.mp4,01:35,01:46,name_read,host at lectern then name card on the screen,frames/name_read_h3_mid.png,"Our second inductee to the academy this year is NAME",,
- Field mapping:
| Field | Source | If unavailable |
|---|---|---|
| segment | agenda block from blocks.csv | none |
| camera_file | the original camera file, not the block | none |
| start_mmss, end_mmss | block source start plus the checked start and end inside the window (the transcript line for name_read and audience_question, the cutaway's own seconds for audience_shot), as MM:SS in the camera file | none |
| moment_type | the query that found it | none |
| picture | what Claude saw on the frames | UNCERTAIN |
| frame | path of the checked frame | none |
| said_local | transcript lines inside the window | blank for audience_shot and screen_text |
| screen_text | read from the full frame | blank |
| use | left empty for the editor | blank |
- Show the user one sample row per moment type and the mapping, and wait for a yes before writing the full sheet. Say which fields are inferred (picture) and what changes if Claude read the frame wrong: the editor sees the wrong shot at that time and skips it.
- Write selects.csv (in our test run the sheet was built from the checked hits; the user approval of sample rows was not exercised in our test run). Add a short note under the sheet with the counts: checked selects per type, candidates rejected per type, and blocks where a type returned nothing. An empty result for audience shots does not prove the footage has no audience shot; say so for those blocks.
Done when selects.csv exists with one row per checked moment and the user has approved the sample rows.
Compliance
- Use only footage the client has authorized for this recap. Confirm this before Step 2.
- People on camera: the event's filming notice (registration terms, signs at the door) covers attendees in the recap only as the client set it up. Ask the user before Step 7 whether anyone asked not to appear, and leave their shots out.
- Minors: if the event had minors in the audience, use their shots only with guardian consent; otherwise leave those rows out.
- Names come from the agenda or the host's words in the transcript. The skill does no face recognition and does not identify anyone.
- The blocks stay in the user's Vivu project until the user deletes them. The skill never calls vivu_delete_video or vivu_delete_project unless the user asks, and confirms first.
- Nothing is posted or sent. selects.csv stays in the working folder.
Known failure modes
| Symptom | Cause | Fix |
|---|---|---|
| (observed) most audience_shot results are stage wide shots from the back of the room or honoree close-ups; in our test run 6 false of 9 returned | composition is read loosely; the reason still calls them "A clear shot of the audience seated at tables, clapping and listening, filmed from the side/front." | treat audience_shot as candidates, keep only frames where seated people are the subject, pick more cutaways from a contact sheet of the audience camera |
| (observed) a questioner listening to the answer comes back as a question; the first wording in our test run returned 1 false of 4 | the picture (person at a mic) matched, the speech did not | use the wording that asks for someone talking into the mic, and check the transcript for a question in the window |
| (observed) no walk-up or handshake on the short contact sheet right after a name read | the show plays a tribute video or bio between the name and the handshake | make the long contact sheet up to the next name read and pick the handshake there |
| (observed) a name on the agenda has no name_read hit; in our test run 3 missed | the host speaks off camera over a full frame name card | cross-check screen_text name cards and the transcript for every name without a hit |
| (observed) a lectern sign or a printed stage backdrop comes back as screen text; in our test run 3 false | the reason calls printed signage a screen or a graphic overlay | keep only windows where the frame shows a projection screen, a video wall or a full frame graphic; read the words from the frame |
| (observed) ffmpeg stops with "Error opening output files: No such file or directory" | the output folder does not exist yet | run mkdir -p for blocks, frames, results and loudness before cutting |
| (observed) the applause stretch is quieter than the host's speech | the block's audio is a desk mix from the lectern microphone; room applause is mixed low | find applause as the gap without transcript words after a name read, and confirm it on the contact sheet |
| "has not granted vivu.write" | Vivu connected read only | the user reconnects Vivu with write access |
| upload page asks to sign in or shows an error | the one time upload link expires after 180 seconds | request a new link right before opening it |
| a block is rejected by the browser upload tool | Claude in Chrome takes at most 10 MB per upload call | the user adds that block in the Vivu web app; never split or recompress it |
| a result's file name does not match blocks.csv | Vivu replaced spaces or punctuation in the name | name blocks with letters, digits and "_" only, and match on the block name |
| a quote on the sheet does not match what was said | it was copied from the reason text, which paraphrases | take every quote from the block's transcript |
| a list stops at exactly maximum_results | the cap cut the list short | raise maximum_results and rerun that query |