---
name: creator-opening-take-sheet
description: "Sort creator raw clips into an opening take sheet with Vivu: the takes found for each opening type, their times, the line as spoken and whether the product is in frame. Use when recutting hooks."
---Opening take sheet from creator raw clips
This skill takes the raw talking clips that creators delivered for a campaign (a folder of files named like creatorname_01.mp4, several retakes per file) and two or three opening types the editor wants to cut new versions from, such as "starts with the viewer's problem", "starts by introducing the product" and "starts with the price". It returns one take sheet: a row for each take the searches and the take split find for each opening type (a type with fewer takes than expected was not found, not proven absent), with the creator file, the take's start and end as MM:SS, the line as spoken from the transcript, whether the take ran to the end or stopped midway, whether the product is in the picture (held up, visible but not held, or not in frame), the frame that shows it, and an empty column for the editor's pick. Claude prices the indexing against the user's Vivu plan, indexes the files in a private Vivu project, runs one search per opening type, splits each result into single takes using the silences between them and the transcript, checks every take against the transcript and against frames, and writes the CSV. The editor does the cutting.
The value is in what the file names and the transcript cannot tell the editor. File names only say which clip it was. A transcript finds the words, but a creator says the same opening three times in a row and only the picture shows which take had the product held up next to the face, which had it sitting on the desk, and which had no product in frame at all, and a hook that says "this is the FoldCup" over an empty hand cannot be used. The search narrows where each opening type was spoken; the transcript decides which type each take really is, so a take that opens with the problem and names the product later is not filed as a product opening; frames decide the product column.
When to use
Use when a brand's in-house editor or creative strategist says "pull every take where the creators open with the price", "which takes of the problem hook have the product in hand", "sort the raw clips from this campaign by how they open", or "I need all the alternate openings to cut new hook variants". For one question about one clip, search Vivu directly. This skill does not judge which opening performs better, does not cut or render variants, and does not identify the creators or any brand in the picture.
Working principles
- Report measured numbers, not estimates. When a number is an estimate, say so.
- Vivu's results are candidates. A take is on the sheet only after the transcript shows that it opens with that type, and the product column is filled only from frames. Rejected candidates are counted and kept in rejected.csv with the reason.
- Stop and tell the user when a required capability or file is missing (no local files, no transcript, no ffmpeg, no write access in Vivu). Do not guess around it.
- Ask the user before anything that is expensive to redo or that acts on their behalf: the indexing cost before upload, and the sample rows before the full sheet.
- Lines on the sheet come from the transcript, never from Vivu's reason text, which is a paraphrase.
What you need before starting
Check each item at the start of the run and tell the user plainly what is missing.
| Requirement | Why | How to check |
|---|---|---|
| Vivu connector with write access | create a private project, open its upload page, search | vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the connection is read only; the user reconnects Vivu and allows write access |
| A shell with ffmpeg and ffprobe on the machine that holds the clips | measure minutes, find the pauses between takes, extract frames | ffmpeg -version and ffprobe -version |
| The creators' raw clips as local files (downloaded from Frame.io, Drive, Dropbox or the creator portal) | frames come from the local files; Vivu has no export tool | ls CLIPS_FOLDER |
| A transcript with timestamps for each clip (an .srt or .vtt exported from the editing app's transcription, or the output of a local speech to text tool) | the line as spoken and the opening type come from it | ls CLIPS_FOLDER for one .srt or .vtt per clip |
| A browser that can open the Vivu upload page | the connector has no direct upload tool | the user opens the link, or Claude in Chrome is connected |
The skill needs a shell and local files, so it runs in Claude Code on the user's computer (terminal or the Code tab of Claude Desktop). It downloads nothing from YouTube, so no residential IP is needed. Nothing is scheduled and nothing is written outside the working folder.
Inputs to collect
Ask for anything missing, most important first.
- CLIPS_FOLDER: the folder with the creators' raw clips and their transcripts. Required.
- The opening types, each with one example line in the editor's words. Default: pain ("opens with the viewer's problem before naming the product"), product ("opens by introducing the product by name, like 'this is' or 'meet'"), price ("opens with the price or a discount").
- PRODUCT_LOOK: a plain description of what the product looks like, for the product column (for example "a teal collapsible coffee cup with a dark lid"). Required; no brand names or logos are needed.
- Which creators or files to include. Default: every clip in CLIPS_FOLDER, up to what the plan covers (Step 3).
- The Vivu project name. Default: "Raw takes CAMPAIGN", private.
Files and state
Keep everything in one working folder next to the clips:
opening-takes/
results/ raw search results, one JSON file per opening type
frames/ contact sheets and full frames per take
candidates.csv every window Vivu returned, as returned
opening_takes.csv one row per take on the sheet
rejected.csv windows or takes that failed a check, with the reason
state.json uploaded files and video ids, searches with job ids, every take with its verdict, rows written
state.json is updated after every step. A rerun reads it first: files already uploaded are not uploaded again, searches already run are not rerun, and takes already checked keep their verdicts, so an interrupted run resumes where it stopped and no take gets two rows.
Step 1: Check the Vivu connector and the setup
Goal: every row of What you need is confirmed before anything is uploaded.
- Call vivu_get_account. If the tool is missing, tell the user to add the Vivu connector in Claude (https://mcp.vivu.ai/mcp) and stop. If can_create_projects is not true, or a write call later fails with "has not granted vivu.write", ask the user to reconnect Vivu with write access and stop until they have.
- Run ffmpeg -version and ffprobe -version. List CLIPS_FOLDER and check that every clip has a transcript file. A clip without one can still be searched, but its rows stay candidates with line_as_spoken NOT VERIFIED.
- Create the working folders; ffmpeg does not create an output folder:
mkdir -p opening-takes/results opening-takes/frames
Done when vivu_get_account shows can_create_projects: true, ffmpeg and ffprobe print versions, and the transcript check is reported per clip.
Step 2: Confirm rights and collect the inputs
Goal: the user has confirmed the clips may be used this way, and Claude has the opening types and PRODUCT_LOOK.
- Ask Compliance items 1 to 3. Stop for any clip that fails them.
- Collect the opening types with one example line each, and PRODUCT_LOOK.
In our test run the opening types and the product description came from the synthetic clips' design, so this step was not exercised in our test run.
Done when the user has confirmed Compliance items 1 to 3 and the opening types and PRODUCT_LOOK are written into state.json.
Step 3: Price the indexing and get approval
Goal: the user sees the cost before anything is uploaded.
- Measure each clip (ffprobe takes one file per call; FILE is one clip's file name in CLIPS_FOLDER, and the same FILE is used in the later commands) and sum the minutes:
ffprobe -v error -show_entries format=duration -of csv=p=0 CLIPS_FOLDER/FILE
- Call vivu_get_usage for the plan and what remains this month.
- Search credits: one precise search per opening type at 5 credits each, plus 5 for each rewording if the first page is not clean. Label the total as an estimate.
- Show one table and wait for approval:
| This campaign | Remaining on the plan | |
|---|---|---|
| Index minutes | measured sum | from vivu_get_usage |
| Search credits | estimate | from vivu_get_usage |
For reference, the Free plan has 20 index minutes a month and 50 search credits a month; Premium is $30 a month with 180 index minutes and 500 search credits. As an estimate, 15 clips of 2 minutes are about 30 index minutes, more than Free covers, and three opening types with one rewording each are about 30 credits. If the clips do not fit, offer these levers in order: index the creators the editor needs first (say which are left out), trim long clips to the first minutes where the openings are (only if the user confirms the openings are there), and only then a larger plan.
In our test run the account was an admin account whose vivu_get_usage shows no remaining allowance, so the plan comparison and the approval were not exercised in our test run.
Done when the user has approved the minutes and the credit estimate.
Step 4: Upload and index
Goal: every clip is ready in a private Vivu project.
- Call vivu_list_projects and reuse the project named in the inputs if it exists. Otherwise call vivu_create_project with that name and visibility "private". Raw takes are unreleased material; the default visibility is the whole organization.
- Call vivu_open_upload_page with the project ID right before the upload. The link expires in 180 seconds and is a sign in link, so never paste it into a message or a file. Give it to the user to open in their own browser and choose the clips, or attach them with a browser tool that can upload local files. Claude in Chrome accepts at most 10 MB per upload call; phone footage at 1080p or 4K is usually larger, so the user adds those files in the Vivu web app. Never split or recompress a clip to fit.
- Poll vivu_list_videos about every 30 seconds until every clip shows ready. Match each video back to the local file by name; Vivu turns spaces and punctuation in names into "_". Record the video IDs in state.json.
In our test run the files went up through a script, not through a user's browser or Claude in Chrome; that upload path was not exercised in our test run.
Done when vivu_list_videos shows every clip as ready.
Step 5: Search each opening type
Goal: one raw result file per opening type.
Run one search per opening type with vivu_search_videos (project_id, query, mode "precise", maximum_results 30). Each call returns a job ID. Call vivu_get_search_results until complete is true; each status call can wait up to 45 seconds, so a pending search is not a stalled one. Save each completed result as results/OPENING_TYPE.json and copy every window into candidates.csv. Show the result page link in the live reply only; it expires after four hours, so it never goes into the sheet or any file.
| Field | Query | Mode | maximum_results |
|---|---|---|---|
| pain | the creator starts a take by describing a problem or frustration the viewer has, before naming any product | precise | 30 |
| product | the creator starts a take by introducing the product by name, with words like "this is" or "meet" | precise | 30 |
| price | the creator starts a take with the price or a discount as the very first words, such as a dollar amount or a percent off, before saying anything else about the product; a price mentioned later in a sentence about something else, or a discount code in a closing call to action, does not count | precise | 30 |
| product held up (optional check) | a person holds PRODUCT_LOOK up next to their face toward the camera | precise | 30 |
The price wording names what does not count because the first, shorter wording ("the creator starts a take by saying the price or a discount first") returned a body line where the creator says what the product cost; see the worked example. The last row flips from what is said to what is shown. It is a quick way to see which clips have the product held up at all, but its windows start and end on frames where the product is not held, so the product column on the sheet always comes from the frames of each take (Step 7), never from this search.
Why this chain: the opening searches narrow where each type was spoken across the whole campaign; the pauses and the transcript turn each window into single takes and confirm the type; the frames of each take fill the product column, which only the picture can answer.
All searches are precise because the sheet needs time ranges; a fast search returns only whole files with no reason. maximum_results 30 leaves room for several takes per file across a whole campaign, and it is the ceiling on what comes back: a list that stops at exactly 30 means the cap was hit, so raise it and rerun that search. Write each query in the editor's own words for the type and add what does not count, as the wording above does.
Done when results/ holds one result file per opening type and state.json records the job IDs.
Step 6: Split windows into single takes and check the opening
Goal: every take inside every window has its own start, end, status and verdict.
Vivu returns a time range that contains the moment, not the exact take. When a creator says the same opening several times in a row, one window can cover two or three takes, the chatter between them and the start of the next line. Speech searches by meaning also bring back neighboring lines that share words with the type (a product opening that mentions the price second, a closing line with a discount code).
- For each window, find the pauses between takes inside it, padded 2 seconds on each side (START is start_ms / 1000 minus 2, DURATION is (end_ms - start_ms) / 1000 plus 4):
ffmpeg -hide_banner -nostats -ss START -t DURATION -i CLIPS_FOLDER/FILE -af silencedetect=noise=-35dB:d=0.8 -f null -
Each stretch between a silence_end and the next silence_start is one spoken line; add START to get times in the clip. A line still running at the end of the padded window runs to the next pause in the full clip. In a noisy room raise the noise threshold (for example -30dB); if the creator barely pauses, split by the transcript's line times instead. 2. For each line, read the transcript inside it. It is a take of the opening type only if its first sentence opens that way. A take that opens with the problem and names the product later is a pain take, not a product take; a product opening that mentions the price second is a product take, not a price take; chatter ("is this recording?"), body lines and a closing call to action are not takes. Copy the take's words into line_as_spoken. 3. A line inside a window that opens with a different type is still a real take: file it under its own type instead of dropping it. Searches can miss a take that sits between windows, and this is how it gets back onto the sheet. A take that sits outside every window is never looked at, so the sheet can miss it; in our test run 1 of the 8 product takes was missed by its own search and found only because it sat inside the pain window. 4. Mark status: full, or false_start when the creator stops midway ("sorry, let me start that again"). Keep false starts on the sheet; editors sometimes use the first half. 5. Windows with no take of the type go to rejected.csv with the reason. Count them.
In our test run the pauses between takes were clean (synthetic voices, no room noise), and the line text came from the fixture script instead of a transcript tool.
Done when every window in candidates.csv has a verdict and every take has a start, end and status in state.json.
Step 7: Check whether the product is in frame
Goal: the product column of every take is read from frames.
- Make a sheet of one frame every half second across the take (START and DURATION are the take's start and length in seconds; TAKE_KEY is a short name for the take such as pain_creatorA_01_0027; a 10x4 grid holds 20 seconds):
ffmpeg -v error -ss START -t DURATION -i CLIPS_FOLDER/FILE -vf "fps=2,scale=320:-2,tile=10x4" -frames:v 1 opening-takes/frames/TAKE_KEY_sheet.png
- Look at the sheet and extract one full frame from the middle of the take, where the line is being said (SECONDS is the take's start plus half its length):
ffmpeg -v error -ss SECONDS -i CLIPS_FOLDER/FILE -frames:v 1 -q:v 3 opening-takes/frames/TAKE_KEY_mid.png
- Fill product_in_frame from what the frames show against PRODUCT_LOOK: "in hand" (held up toward the camera), "visible, not held" (on the desk or in the background), "not in frame", or UNCERTAIN (blurred, half out of frame, a lookalike object). When the product appears only for part of the take, write when.
In our test run the frames were flat cartoon scenes; real footage with motion blur, a half visible product or a lookalike object was not exercised in our test run.
Worked example from our test run
In our test run we used 8 synthetic clips, 4.5 minutes in total, made for the test: cartoon creators at a desk, each creator with its own synthetic voice, and a teal collapsible cup that is held up next to the face, standing on the desk, or out of the picture. Before searching we wrote down 23 opening takes (9 pain, 8 product, 6 price), 4 of them false starts, plus chatter between takes, body lines, a closing line with a discount code, a product take that says the price second and a pain take that names the product second. Nothing is written on screen, and the opening a take uses is crossed with where the cup is, so the picture cannot give the opening away. The clips were ready about 4.9 minutes after the upload started.
The pain search returned 5 windows: 5 real, 0 false and 0 takes missed. The product search returned 5 windows: 5 real, 0 false and 1 take missed, so it covered 7 of the 8 product takes; the missed take's first words sat at the end of the pain window for the same clip, and the take split filed it under product, so the sheet still had it. The first price wording returned 4 windows: 3 real and 1 false, a body line where the creator says what the cup cost. After 1 rewording on the same corpus, which names what does not count, the price search returned 3 windows: 3 real, 0 false and 0 takes missed. Median windows were 14 seconds for pain, 10 for product and 13.5 for price, against takes of a few seconds; several windows held more than one take of the same opening, and one price window was the whole clip. The pauses split every line, and the product column read from frames matched our notes on 23 of 23 takes. The optional search for the cup held up returned 9 windows, all 9 real, but window edges fell on frames where the cup was on the desk or out of frame.
These were short clips with clean synthetic voices and pauses between takes, and the transcript was the clips' own script. Transcription of real creator audio, background music, a noisy room and full length raw clips were not exercised in our test run.
Done when every take has a product_in_frame value and a frame path.
Step 8: Assemble the take sheet
Goal: one CSV the editor can work from.
- Write the rows grouped by opening type, then by creator file and time:
opening_type,creator_file,take_start,take_end,status,line_as_spoken,product_in_frame,frame,editor_pick
pain,creatorA_01.mp4,00:27,00:34,full,"Ever had your coffee go cold before you even made it to the office? I finally found a fix that fits in my pocket.",in hand,frames/take_A01-pain-3_mid.png,
- Field mapping:
| Field | Source | If unavailable |
|---|---|---|
| opening_type | the search that found it, confirmed by the transcript | none |
| creator_file | the local file name | none |
| take_start, take_end | the spoken line from Step 6, as MM:SS | the window edges, marked UNCERTAIN |
| status | full or false_start, from the transcript | UNCERTAIN |
| line_as_spoken | the transcript, word for word | NOT VERIFIED |
| product_in_frame | the frames from Step 7 | UNCERTAIN |
| frame | the full frame the product column was read from | blank |
| editor_pick | left empty for the editor | blank |
- Show two or three sample rows and the mapping and wait for a yes before writing the full sheet. Say which parts are Claude's reading: the take boundaries come from pauses and the transcript, and a type with fewer takes than expected was not found, not proven absent.
- Write opening_takes.csv and rejected.csv and give the user the take counts per type and per creator, including false starts.
In our test run the sheet was written from the checked takes; the user's approval of the sample rows was not exercised in our test run.
Done when opening_takes.csv exists with a row for every take found and the user has approved the sample rows.
Compliance
- Use only clips the creators have licensed to the brand, and check that the license covers uploading them to a third party indexing service. Ask before Step 3.
- The creators on camera have signed for the shoot, but some contracts limit who may see raw takes. If a creator is a minor, the skill needs a guardian's consent or it does not process those clips. Ask before Step 3.
- The clips stay in the user's Vivu project until the user deletes them. Create the project as private, because raw takes are unreleased. The skill never calls vivu_delete_video or vivu_delete_project unless the user asks, and confirms first.
- The skill does not judge which opening is better, does not cut or publish variants, and posts nothing anywhere. No face or logo recognition: creators are known only by their file names and the product only by PRODUCT_LOOK.
Known failure modes
| Symptom | Cause | Fix |
|---|---|---|
| (observed) the price search returns a window that holds only a body line where the creator says what the product cost | a search by meaning matches any line that says a price | name what does not count in the query, and let the transcript check reject windows with no take of the type |
| (observed) one window covers several takes of the same opening and the chatter between them | windows are wider than takes, and repeated lines sound alike | split the window at the pauses and give each take its own row |
| (observed) a take is missing from its own type's search, but its first words sit at the end of another type's window | the window for that clip covered only the next take | file every line found inside any window under the type its first sentence has |
| (observed) a price window spans the whole clip | Vivu returned the clip as one range | split at the pauses as usual; it is one candidate, not one take |
| (observed) the product held up search returns windows that start on frames where the product is on the desk or out of frame | window edges run past the moment | fill product_in_frame from frames inside each take, never from that search's window |
| the video summary puts several takes of the same opening in one segment | segments follow topics, not takes | use segments only as a second signal; takes come from pauses and the transcript |
| the reason text quotes a line | the reason is a paraphrase and can change words | copy lines from the transcript only |
| silencedetect finds no pauses, or splits mid sentence | room noise, music, or a creator who never pauses | raise the noise threshold, or split by the transcript's line times |
| a type returns no windows | the creators never opened that way, or the wording missed | an empty result does not prove the opening is absent; reword once with an example line, then tell the user |
| a list stops at exactly maximum_results | the cap cut the list short | raise maximum_results and rerun that search |
| "has not granted vivu.write" | Vivu connected read only | the user reconnects Vivu with write access |
| upload page asks to sign in or shows an error | the one time upload link expires after 180 seconds | request a new link right before opening it |
| a file is rejected by the browser upload tool | Claude in Chrome takes at most 10 MB per upload call | the user adds that file in the Vivu web app; never split or recompress it |