---
name: wedding-speech-selects-sheet
description: "Turn a couple's toast and vow recordings into a selects sheet with Vivu: where each moment the editor wants was said, the words as spoken, and what the picture shows. Use when cutting speeches."
---Selects sheet from wedding toasts and vows
This skill takes the ceremony and reception camera files from one wedding, the day's run sheet, and a short list of moments the editor wants to use (the best man's story of how the couple met, the father's thank you to both families, one particular vow promise). It returns one selects sheet, a CSV with a row for each place a moment was found: the moment, which speech it is in, who was speaking according to the run sheet, the camera file, start and end as MM:SS, the words as spoken from a local transcript, what the picture shows at that point (the speaker, the couple reacting, the room raising glasses), a frame path, and an empty column for the editor's call. Claude cuts the vows and each toast out of the long camera files using the run sheet times, prices the indexing against the user's Vivu plan, indexes only those segments in a private Vivu project, runs one search per moment, checks every result against the transcript and against frames, and writes the CSV. The editor does the cutting.
The value is in what the file names and a word search cannot tell the editor. Camera files are named by card and clip number, and toasts are unscripted, so the only record of who told which story where is the audio itself. A word search through a transcript finds "met" or "promise" in every speech; a toast holds three or four short stories that share names and words, and the editor wants one of them. Vivu finds the passage by what it is about, the transcript confirms it is the right story and gives the exact words, and frames from the footage show whether the speaker or the couple is on screen when the line lands, which decides whether it can go in the highlight film.
When to use
Use when a wedding filmmaker or editor says "find where the best man tells how they met", "pull the father's thank you and the vow about coffee", "I need selects from the toasts for the highlight film", or "which toast has the story about the snowstorm". For one question about one short clip, search Vivu directly. This skill does not cut, color or render anything, does not judge which take is better, and does not identify anyone by face or voice: speakers come from the run sheet or from how they introduce themselves.
Working principles
- Report measured numbers, not estimates. When a number is an estimate, say so.
- Vivu's results are candidates. A moment is on the sheet as verified only after the transcript inside the window shows it is that story, and the picture column is filled only from frames. Rejected windows are counted and kept in rejected.csv with the reason.
- Stop and tell the user when a required capability or file is missing (no local files, no run sheet times, no transcript, no ffmpeg, no write access in Vivu). Do not guess around it.
- Ask the user before anything that is expensive to redo or that acts on their behalf: the indexing cost before upload, and the sample rows before the full sheet.
- Words on the sheet come from the transcript, never from Vivu's reason text, which is a paraphrase.
What you need before starting
Check each item at the start of the run and tell the user plainly what is missing.
| Requirement | Why | How to check |
|---|---|---|
| Vivu connector with write access | create a private project, open its upload page, search | vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the connection is read only; the user reconnects Vivu and allows write access |
| A shell with ffmpeg and ffprobe on the machine that holds the footage | cut the speech segments, measure minutes, make contact sheets and frames | ffmpeg -version and ffprobe -version |
| The camera files as local files (the studio's edit drive or NAS) | segments and frames come from the local files; Vivu has no export tool | ls FOOTAGE_FOLDER |
| The run sheet with times and speaking order (text, from the planner or the couple) | which file and time range holds the vows and each toast, and who spoke | read RUN_SHEET; every speech has a start time and a speaker |
| A transcript with timestamps for each speech segment (an .srt or .vtt from the editing app's transcription, or a local speech to text tool) | the words as spoken and the check that a window holds the right story | ls speech-selects/segments for one .srt or .vtt per segment after Step 3 |
| A browser that can open the Vivu upload page | the connector has no direct upload tool | the user opens the link, or Claude in Chrome is connected |
The skill needs a shell and local files, so it runs in Claude Code on the user's computer (terminal or the Code tab of Claude Desktop). It downloads nothing from YouTube, so no residential IP is needed. Nothing is scheduled and nothing is written outside the working folder.
Inputs to collect
Ask for anything missing, most important first.
- FOOTAGE_FOLDER: the folder with the wedding's camera files. Required.
- RUN_SHEET: the run sheet, with the time each speech started and who gave it. Required. If times are missing, ask the editor for rough times; do not upload whole camera files to find them.
- The moments, five to eight of them, each in the editor's words with one detail only that story has (a place, an object, what happened). Required. Example: "best man: how they met (laundromat, wrong basket)".
- Which camera to use for each speech. Default: the one with the cleanest sound, usually the camera fed from the mixer or a lavalier.
- The Vivu project name. Default: "Speeches COUPLE DATE", private.
Files and state
Keep everything in one working folder next to the footage:
speech-selects/
segments/ one cut file per speech, plus its .srt or .vtt
results/ raw search results, one JSON file per moment
frames/ contact sheets and single frames per row
candidates.csv every window Vivu returned, as returned
selects.csv one row per verified place a moment was said
rejected.csv windows that failed the transcript check, with the reason
state.json segments cut, uploaded files and video ids, searches with job ids, every window with its verdict, rows written
state.json is updated after every step. A rerun reads it first: segments already cut are not cut again, files already uploaded are not uploaded again, searches already run are not rerun, and windows already checked keep their verdicts, so an interrupted run resumes where it stopped and no moment gets two rows for the same window.
Step 1: Check the Vivu connector and the setup
Goal: every row of What you need is confirmed before anything is cut or uploaded.
- Call vivu_get_account. If the tool is missing, tell the user to add the Vivu connector in Claude (https://mcp.vivu.ai/mcp) and stop. If can_create_projects is not true, or a write call later fails with "has not granted vivu.write", ask the user to reconnect Vivu with write access and stop until they have.
- Run ffmpeg -version and ffprobe -version. List FOOTAGE_FOLDER and read RUN_SHEET.
- Create the working folders; ffmpeg does not create an output folder:
mkdir -p speech-selects/segments speech-selects/results speech-selects/frames
Done when vivu_get_account shows can_create_projects: true, ffmpeg and ffprobe print versions, and the run sheet has a start time and speaker for every speech.
Step 2: Confirm rights and collect the moments
Goal: the user has confirmed the footage may be processed this way, and Claude has the list of moments.
- Ask Compliance items 1 to 3. Stop if the contract does not allow cloud processing.
- Collect the moments. For each one, ask for the detail only that story has. Searches by meaning mix up neighboring stories in the same toast when the wording is generic (see Step 6), and the detail is what keeps them apart.
In our test run the moments came from the synthetic speeches' design, so this step was not exercised in our test run.
Done when the user has confirmed Compliance items 1 to 3 and the moments, each with its detail, are written into state.json.
Step 3: Cut the speech segments from the camera files
Goal: one short file per speech, so only the speeches are indexed.
- For each speech on the run sheet, find the camera file that covers it and the offset inside that file (camera files carry their own start time in the metadata; check with ffprobe -v error -show_entries format_tags=creation_time FOOTAGE_FOLDER/CAMERA_FILE, or ask the editor). Cut from a little before the speech starts to a little after it ends, so a late start or a long laugh is not lost (CAMERA_FILE is the source file, START and END are seconds inside it, SEGMENT is a name like toast1_bestman_camB):
ffmpeg -v error -ss START -to END -i FOOTAGE_FOLDER/CAMERA_FILE -c:v libx264 -c:a aac speech-selects/segments/SEGMENT.mp4
Re-encoding makes the cut land where asked; a stream copy snaps to the nearest keyframe and can drop the first seconds of a toast. 2. Name each segment with the speech and the speaker from the run sheet. The name is the only place the sheet records who spoke, because the skill never guesses a speaker from a voice. 3. Make or export a transcript for each segment into the same folder, named SEGMENT.srt or SEGMENT.vtt.
In our test run the segments were synthetic files already cut to one speech each; the cut command ran once on one of them to check it, and finding offsets in real multi hour camera files was not exercised in our test run.
Done when segments/ holds one file and one transcript per speech and state.json lists them with their speaker.
Step 4: Price the indexing and get approval
Goal: the user sees the cost before anything is uploaded.
- Measure each segment (ffprobe takes one file per call) and sum the minutes:
ffprobe -v error -show_entries format=duration -of csv=p=0 speech-selects/segments/SEGMENT.mp4
- Call vivu_get_usage for the plan and what remains this month.
- Search credits: one precise search per moment at 5 credits each, plus 5 for each rewording. Label the total as an estimate.
- Show one table and wait for approval:
| This wedding | Remaining on the plan | |
|---|---|---|
| Index minutes | measured sum | from vivu_get_usage |
| Search credits | estimate | from vivu_get_usage |
For reference, the Free plan has 20 index minutes a month and 50 search credits a month; Premium is $30 a month with 180 index minutes and 500 search credits. As an estimate, a 15 minute ceremony plus four 10 minute toasts is about 55 index minutes, more than Free covers and about a third of Premium, and seven moments are about 35 credits. Uploading whole camera files instead (three cameras over six hours is over a thousand minutes) is far beyond any plan, which is why Step 3 cuts first. If the segments still do not fit, offer these levers in order: drop the speeches with no moments on the list, trim long silences at the start and end of each segment, and only then a larger plan.
In our test run the account was an admin account whose vivu_get_usage shows no remaining allowance, so the plan comparison and the approval were not exercised in our test run.
Done when the user has approved the minutes and the credit estimate.
Step 5: Upload and index
Goal: every segment is ready in a private Vivu project.
- Call vivu_list_projects and reuse the project named in the inputs if it exists. Otherwise call vivu_create_project with that name and visibility "private". Wedding footage is the couple's private material; the default visibility is the whole organization.
- Call vivu_open_upload_page with the project ID right before the upload. The link expires in 180 seconds and is a sign in link, so never paste it into a message or a file. Give it to the user to open in their own browser and choose the files in segments/, or attach them with a browser tool that can upload local files. Claude in Chrome accepts at most 10 MB per upload call; a 10 minute segment from a 1080p camera is usually larger, so the user adds those files in the Vivu web app. Never split or recompress a segment to fit.
- Poll vivu_list_videos about every 30 seconds until every segment shows ready. Match each video back to the local file by name; Vivu turns spaces and punctuation in names into "_". Record the video IDs in state.json.
In our test run the files went up through a script, not through a user's browser or Claude in Chrome; that upload path was not exercised in our test run.
Done when vivu_list_videos shows every segment as ready.
Step 6: Search each moment
Goal: one raw result file per moment.
Run one search per moment with vivu_search_videos (project_id, query, mode "precise", maximum_results 5). Each call returns a job ID. Call vivu_get_search_results until complete is true; each status call can wait up to 45 seconds, so a pending search is not a stalled one. Save each completed result as results/MOMENT.json (MOMENT is the short field name, such as how_they_met) and copy every window into candidates.csv. Show the result page link in the live reply only; it expires after four hours, so it never goes into the sheet or any file.
Write each query as what the speaker talks about, in the editor's words, with the one detail that story has. Do not ask for a speaker by voice or name; the segment file already says whose speech it is. These are the queries from our test run, as examples of the shape:
| Field | Query | Mode | maximum_results |
|---|---|---|---|
| how_they_met | a wedding toast where the speaker tells the story of how the couple first met | precise | 5 |
| first_date_call | the speaker describes what the groom said right after his first date with the bride | precise | 5 |
| sick_care | a story about the groom taking care of the bride when she was sick | precise | 5 |
| thanks_families | the speaker thanks the groom's parents and welcomes the other family | precise | 5 |
| marriage_advice | the speaker gives the newlyweds a piece of advice about marriage | precise | 5 |
| coffee_vow | in the wedding vows, a promise about making coffee every morning | precise | 5 |
The thanks_families wording was reached after 2 rewordings on the same corpus. The first wording, "the speaker thanks both families, the bride's and the groom's", returned nothing, because the father thanked only the groom's parents and talked about two families becoming one; naming who is thanked found it. A moment that returns nothing gets one rewording with a more concrete detail, then the user is told; an empty result does not prove the moment was never said.
All searches are precise because the sheet needs time ranges; a fast search returns only whole files with no reason. A moment is usually said once in a wedding, so 5 leaves room for a story told twice (in a toast and again in the vows) without flooding the check. maximum_results is also the ceiling on what comes back: a list that stops at exactly 5 means the cap was hit, so raise it and rerun.
Why this chain: the search narrows each moment to one stretch of one speech; the transcript (Step 7) confirms it is the right story and gives the exact words and the real start and end; the frames (Step 8) say what is on screen while it is said, which only the footage can answer. There is no search for the picture itself: Vivu reads shot composition (who is the subject, who is in focus) poorly, so the picture column comes from frames, not from a search.
Done when results/ holds one result file per moment and state.json records the job IDs.
Step 7: Check each window against the transcript
Goal: every window has a verdict, and every verified moment has its words and its real start and end.
Vivu returns a time range that contains the passage, not the exact sentence. Windows run past the story into the neighboring line, a reaction shot or the start of the next story, and in a toast the neighbor is often a story that shares names and words.
- For each window, read the transcript lines inside it, padded a few seconds on each side.
- It is verified only if the lines tell the moment the editor asked for. A story about how the speaker met the groom is not "how the couple met"; a passing mention of the same event in the vows is not the toast story. Rejected windows go to rejected.csv with the reason, and the count goes into the report.
- Set start and end to the first and last transcript line of the story, not the window edges. Copy the story's words into line_as_spoken; shorten long stories with "..." but never change words.
- vivu_get_video_summary with include_segments, start_ms and end_ms around the window is a second opinion when the transcript is unclear. Its segments follow topics loosely and can merge two neighboring stories into one, so it never decides a verdict alone.
In our test run the transcript was the synthetic speeches' own script with measured line times; real speech to text on noisy reception audio was not exercised in our test run.
Done when every window in candidates.csv has a verdict and every verified moment has start, end and line_as_spoken in state.json.
Step 8: Read the picture from frames
Goal: the picture column of every verified moment is read from frames.
- Make a sheet of one frame every half second across the moment (START and DURATION are the moment's start and length in seconds from Step 7; KEY is a short name like how_they_met_toast1_0039; a 10x6 grid holds 30 seconds, so use a taller grid for longer stories):
ffmpeg -v error -ss START -t DURATION -i speech-selects/segments/SEGMENT.mp4 -vf "fps=2,scale=320:-2,tile=10x6" -frames:v 1 speech-selects/frames/KEY_sheet.png
- Look at the sheet and extract one full frame where the key line lands (SECONDS is that line's time from the transcript):
ffmpeg -v error -ss SECONDS -i speech-selects/segments/SEGMENT.mp4 -frames:v 1 -q:v 3 speech-selects/frames/KEY_line.png
- Fill picture from what the frames show, with times: the speaker at the mic, the couple reacting, guests laughing or raising glasses, the officiant, or UNCERTAIN (camera moving, someone blocking the shot). Say when a reaction shot sits just before or after the moment; editors often extend a cut to catch it.
In our test run the frames were flat cartoon scenes with clean cuts between the speaker, the couple and the room; real handheld footage, a second camera angle and dim reception light were not exercised in our test run.
Worked example from our test run
In our test run we used 4 synthetic speech files, 6.8 minutes in total, made for the test: the vows, and toasts from a best woman, a maid of honor and the father of the bride, each toast a run of short stories that share names and words, all in synthetic voices. Under the speech we mixed a low noise bed, a soft music bed, bursts of noise standing in for laughter and applause, and a feedback squeal. The picture is cartoons with no text, so only the audio tells the stories apart. Before searching we wrote down each moment, the neighboring stories most likely to be confused with it, and a moment nobody tells (the proposal). The files were ready about 5.5 minutes after the upload started.
The moment searches found six moments, with 0 false and 0 missed, and each search that found its moment returned a single window. How they met and the first date call are neighboring stories in the same toast and came back as separate windows. The story about how the speaker met the groom, the vow line that mentions the laundromat and the vow line about the snowstorm were not returned. The proposal search returned nothing. We ran nine precise searches in total, including the two rewordings, about 45 credits as an estimate. The thanks_families wording that worked came after 2 rewordings on the same corpus (the first one returned nothing). Windows ran from 14 to 48 seconds and every one spilled a few seconds into a neighboring line or a reaction shot, so start and end on the sheet come from the transcript. The video summary for the first toast put how they met, the first date call and part of the camping story in one segment and left the call out of its summary.
These were clean synthetic voices in short toasts, with synthetic noise instead of a real room. Real voices, accents, a speaker who wanders off mic, full length toasts and real speech to text were not exercised in our test run, so the transcript check stays in place for every row.
Done when every verified moment has a picture value and a frame path.
Step 9: Assemble the selects sheet
Goal: one CSV the editor can cut from.
- Write the rows in the order of the day (ceremony, then toasts in run sheet order), then by time:
content,speech,speaker_from_run_sheet,camera_file,start,end,line_as_spoken,picture,frame,status,use
how they met (laundromat),toast 1,best man (run sheet),toast1_bestman_camB.mp4,00:39,01:08,"Now, the story of these two. Four years ago...",speaker at mic; couple laughing 01:12-01:16 right after,frames/how_they_met_toast1_0039_sheet.png,verified,
- Field mapping:
| Field | Source | If unavailable |
|---|---|---|
| content | the editor's moment | none |
| speech, speaker_from_run_sheet | the run sheet, through the segment's file name | UNCERTAIN; never guessed from the voice |
| camera_file | the local segment file name | none |
| start, end | first and last transcript line of the story, as MM:SS in the segment | the window edges, marked UNCERTAIN |
| line_as_spoken | the transcript, word for word | NOT VERIFIED |
| picture | the frames from Step 8 | UNCERTAIN |
| frame | the sheet or frame the picture was read from | blank |
| status | verified, or candidate when the transcript could not be checked | none |
| use | left empty for the editor | blank |
- To find the time in the original camera file, add the segment's start offset from Step 3; say so in the reply so the editor can relink to the camera originals.
- Show two sample rows and the mapping and wait for a yes before writing the full sheet. Then write selects.csv and rejected.csv and report, per moment, found or not found (a moment not found was not found, not proven unsaid) and the number of rejected windows.
In our test run the sheet was written from the checked windows; the user's approval of the sample rows was not exercised in our test run.
Done when selects.csv exists with a row for every verified moment and the user has approved the sample rows.
Compliance
- Use only footage the contract allows to be processed by a third party cloud service. Many wedding contracts limit the footage to delivery to the couple; check before Step 3, and ask the couple if the contract is silent.
- The couple signed the contract; guests did not. Guests on camera, including children, are used only in the films delivered to the couple. Using a guest or a child in the studio's own promotion needs that person's (or a guardian's) consent, which this skill does not collect. Ask before Step 3.
- The segments stay in the user's Vivu project until the user deletes them. Create the project as private. The skill never calls vivu_delete_video or vivu_delete_project unless the user asks, and confirms first.
- No face recognition and no telling people apart by voice: the speaker column comes from the run sheet or from how the speaker introduces themselves. Vivu only finds passages; it does not cut or render anything, and the skill posts nothing anywhere.
Known failure modes
| Symptom | Cause | Fix |
|---|---|---|
| (observed) a moment search returns no results although the speech has it | the wording asked for something general (thanks to both families) and the speaker said something specific (thanks to the groom's parents) | reword once with the concrete detail of that story, then tell the user if it is still empty |
| (observed) every window runs a few seconds past the story into the neighboring line or a reaction shot | windows are wider than the passage | take start and end from the first and last transcript line of the story |
| (observed) the windows for two neighboring stories in the same toast overlap | the stories sit back to back and windows spill over | give each row only its own transcript lines; never merge the two |
| (observed) the advice window starts on the last line of a different lesson in the same toast | the window edge falls on the previous story | read the transcript inside the window before copying lines |
| (observed) the video summary puts two neighboring stories in one segment and leaves one of them out | segments follow topics loosely | use segments only as a second opinion; the transcript decides |
| a window holds a neighboring story that shares names or words with the moment | searches by meaning can match a related story in the same speech | reject it in Step 7, count it, and reword the query with the story's own detail |
| the reason text quotes the speaker | the reason is a paraphrase and can change words | copy words from the transcript only |
| the transcript garbles a line under laughter, music or feedback | speech to text on reception audio; recall under real banquet noise is untested | mark line_as_spoken NOT VERIFIED and let the editor listen |
| a list stops at exactly maximum_results | the cap cut the list short | raise maximum_results and rerun that search |
| "has not granted vivu.write" | Vivu connected read only | the user reconnects Vivu with write access |
| upload page asks to sign in or shows an error | the one time upload link expires after 180 seconds | request a new link right before opening it |
| a file is rejected by the browser upload tool | Claude in Chrome takes at most 10 MB per upload call | the user adds that file in the Vivu web app; never split or recompress it |
| a result's file name does not match the local segment | Vivu turns spaces and punctuation into "_" | match on the segment name with "_" in place of spaces and punctuation |