---
name: podcast-remembered-moments-table
description: "Find moments you remember from your own video podcast recordings with Vivu and get a table of source file, MM:SS, checked frame and transcript line. Use when you can't find a remembered moment."
---Remembered moments from your video podcast recordings
This skill takes the moments a video podcaster remembers ("the guest laughed when we got to pricing", "I held up the product", "we turned to look at the screen") and the recordings on their disk, and returns a table they can edit from. Claude uses the local transcript to cut out only the stretches of each episode where the topic comes up, prices those minutes against the user's Vivu plan, uploads them to a private Vivu project, runs one precise search per remembered moment, checks every result against the transcript, then looks inside each checked window for the cue the user remembered: a loudness peak for laughter, a two frames a second contact sheet for something held up, someone walking in, or a turn to the screen. The output is moments.csv with the moment, source file, start and end MM:SS in the original recording, a frame, the transcript line and what Claude saw.
The value is in the cue, not the topic. A topic that matters to a show comes up again and again, so the transcript gives a list of places, not the place. The one the host remembers is the one where something happened in the room: a laugh, an object held to the camera, a person stepping into the shot. Transcripts rarely mark laughter and never mark gestures. Vivu narrows each topic to a few time ranges, and the skill finds the cue inside them on the host's own files. Vivu only finds moments here; the host opens the original and cuts it.
When to use
Use when someone says "I remember the guest laughing about pricing but I can't find it", "find the part where I held up the product", "where in these episodes did we look at the screen", or "give me timestamps for these five moments I remember". For a single moment whose words the user remembers exactly, searching the transcript in their editor is faster. For turning an episode into short clips or captions, this skill is the wrong fit; it only finds and lists moments.
Working principles
- Report measured numbers, not estimates. When a number is an estimate, say so.
- A moment is "found" only after its topic has been checked in the transcript and its cue has been seen on frames (a loudness peak alone is only a lead). Everything else is a candidate or "not found", labeled as such.
- Stop and tell the user when a required capability or file is missing (no transcript, no ffmpeg, no write access in Vivu). Do not guess around it.
- Ask the user before anything that is expensive to redo or that acts on their behalf: the cut list before cutting, the indexing cost before upload, the sample row before the full table.
- Vivu searches for what was said. Laughter and gestures are found locally, because Vivu's search for laughter and for hand movements is not reliable enough to trust on its own.
What you need before starting
Check each item at the start of the run and tell the user plainly what is missing.
| Requirement | Why | How to check |
|---|---|---|
| Vivu connector with write access | create a private project, open its upload page, search | vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the connection is read only; the user reconnects Vivu and allows write access |
| A shell with ffmpeg and ffprobe on the machine that holds the recordings | cut the topic stretches, measure minutes, loudness and frames | ffmpeg -version and ffprobe -version |
| Python 3 on the same machine | runs the two small helper scripts in Steps 2 and 6 | python3 --version |
| The episode recordings on local disk | the table points at the original files; Vivu has no export tool | ls EPISODE_FOLDER |
| A transcript with timestamps (SRT or VTT) for each episode | the topic filter in Step 2 and the quote on each row come from it, never from Vivu's reason text | one .srt or .vtt per episode; the editing app's transcription export or a local speech to text tool makes it |
| A browser that can open the Vivu upload page | the connector has no direct upload tool | the user opens the link, or Claude in Chrome is connected |
No residential IP is needed (nothing is downloaded), nothing is scheduled, and nothing is posted anywhere.
Inputs to collect
Ask for anything missing, most important first.
- The moments, one sentence each: the topic ("when we got to pricing") plus the cue ("the guest laughed", "I held up the jar", "Sam walked into the shot", "we turned to the screen"). Three to six works well. Required.
- The episode files and their transcripts. Required.
- Which episodes each moment might be in. Default: all given episodes.
- Margin around each topic stretch when cutting. Default: 60 seconds on each side.
- The Vivu project name. Default: "SHOW_NAME remembered moments", private.
Files and state
Keep everything in one working folder next to the recordings:
moments/
moments.json the user's moments: id, topic words, cue, episodes
cuts.csv one row per cut stretch: cut file, episode file, source start seconds, duration
cuts/ cut stretches (.mp4) and their transcripts (.srt)
results/ raw search results, one JSON file per moment
frames/ contact sheets and chosen frames
loudness/ ebur128 output per cut
moments.csv the table
state.json uploaded cut names and video ids, searches run with job ids, windows checked with verdicts, rows written
state.json is updated after every step. A rerun reads it first: cuts already uploaded are not uploaded again, searches already run are not rerun, and windows already checked keep their verdicts, so an interrupted run picks up where it stopped and no moment is written twice.
Step 1: Check the Vivu connector and the setup
Goal: every row of What you need is confirmed before anything is cut or uploaded.
- Call vivu_get_account. If the tool is missing, tell the user to add the Vivu connector in Claude (https://mcp.vivu.ai/mcp) and stop. If can_create_projects is not true, or a write call later fails with "has not granted vivu.write", ask the user to reconnect Vivu with write access and stop until they have.
- Run ffmpeg -version, ffprobe -version and python3 --version. List the episode folder and check that every episode has a transcript.
- Write the user's moments to moments.json, each with two to five topic words that would appear in the transcript (for "when we got to pricing": price, pricing, charge, money).
Done when vivu_get_account shows can_create_projects: true, the three tools print versions, and moments.json holds every moment with its topic words.
Step 2: Cut only the stretches where the topics come up
Goal: index minutes go to the parts of each episode that can hold a remembered moment, not to whole episodes.
- Find the topic stretches in each transcript. Save this as topics.py and run python3 topics.py TRANSCRIPT MARGIN WORD1 WORD2 ... (MARGIN in seconds, default 60):
import re, sys
path, margin, words = sys.argv[1], float(sys.argv[2]), [w.lower() for w in sys.argv[3:]]
def secs(ts):
parts = ts.replace(",", ".").split(":")
return sum(float(p) * 60 ** i for i, p in enumerate(reversed(parts)))
text = open(path, encoding="utf-8").read()
hits = []
for m in re.finditer(r"([\d:.,]+) --> ([\d:.,]+)[^\n]*\n(.+?)(?:\n\n|\Z)", text, re.S):
line = re.sub(r"<[^>]+>", "", m.group(3)).lower()
if any(w in line for w in words):
hits.append(secs(m.group(1)))
spans = []
for t in sorted(hits):
a, b = max(0, t - margin), t + margin
if spans and a <= spans[-1][1]:
spans[-1][1] = b
else:
spans.append([a, b])
for a, b in spans:
print("%d %d" % (a, b))
Each output line is a START and END in seconds. Merge the lines of all moments for the same episode, show the cut list with the minutes per episode, and wait for a yes. 2. Create the folders first (ffmpeg does not create an output folder), then cut with a re-encode so each cut starts exactly at START and every time on the table maps back to the episode file:
mkdir -p moments/cuts moments/frames moments/results moments/loudness
ffmpeg -v error -ss START -to END -i EPISODE_FILE -c:v libx264 -crf 20 -c:a aac moments/cuts/CUT_NAME.mp4
EPISODE_FILE is the original recording; CUT_NAME is a short name such as EP70_0372. Use only letters, digits and the "_" character; Vivu rewrites spaces and punctuation in file names, and plain names match back without guessing. The rest of this skill runs from inside moments/. 3. Write each cut to cuts.csv with its source start seconds; a time on the table is source start plus the offset inside the cut. Put the matching part of the transcript next to each cut as cuts/CUT_NAME.srt (shift the times by START).
In our test run the stretches came from public episodes, picked by searching their auto subtitles for the topic words by hand and downloaded as sections; the cut command above was run once on a stretch to confirm it, and topics.py itself was not exercised in our test run.
Done when cuts/ holds one .mp4 and one .srt per approved stretch and cuts.csv has one row per cut.
Step 3: Price the indexing and get approval
Goal: the user sees what the cuts cost before anything is uploaded.
- Measure each cut: ffprobe -v error -show_entries format=duration -of csv=p=0 cuts/CUT_NAME.mp4, and sum the minutes.
- Call vivu_get_usage for the plan and what remains this month.
- Estimate search credits: one precise search per moment, plus one more for any moment whose wording is changed and rerun. A precise search uses 5 credits in total. Label this as an estimate.
- Show one table and wait for approval:
| These moments | Remaining on the plan | |
|---|---|---|
| Index minutes | measured sum | from vivu_get_usage |
| Search credits | estimate | from vivu_get_usage |
For reference, the Free plan has 20 index minutes a month and 50 search credits a month; Premium is $30 a month with 180 index minutes and 500 search credits. A full 90 minute episode does not fit the Free plan, which is why Step 2 cuts. If the cuts still do not fit, offer these levers in order: tighter topic words, a smaller margin, fewer episodes per moment (the ones the user thinks most likely), and only then a larger plan. Never drop a moment or an episode the user asked for without saying so.
In our test run the account was an admin account whose vivu_get_usage shows no remaining allowance, so the comparison against a real plan and the approval (Step 3) were not exercised in our test run.
Done when the user has approved the minutes and the credit estimate.
Step 4: Upload and index
Goal: every approved cut is ready in a private Vivu project.
- Call vivu_list_projects and reuse the project named in moments.json if it exists. Otherwise call vivu_create_project with that name and visibility "private". Guest conversations belong in a private project; the default visibility is the whole organization.
- Call vivu_open_upload_page with the project ID right before the upload. The link expires in 180 seconds and is a sign in link: do not paste it into any message or file. Give it to the user to open in their own browser and choose the files in cuts/, or attach them with a browser tool that can upload local files. Claude in Chrome accepts at most 10 MB per upload call, and a ten minute cut is usually larger; the user adds those in the Vivu web app. Never split or recompress a cut to fit a tool limit.
- Poll vivu_list_videos about every 30 seconds until every cut shows ready. Match each video to cuts.csv by file name; Vivu turns spaces and punctuation into underscores. Record the video IDs in state.json.
In our test run the second cut waited in the queue until the first was ready. The upload did not go through a user's browser or Claude in Chrome; that path was not exercised in our test run.
Done when vivu_list_videos shows every cut in cuts.csv as ready.
Step 5: Search each moment's topic and check every window
Goal: for each remembered moment, the list of time ranges where its topic is really being discussed.
- Write one query per moment that describes what was being talked about, in plain words, and leave the cue out. "The guest laughed when we got to pricing" becomes a query about how stylists can make more money and what they can charge; the laugh is found in Step 6. Vivu is good at finding what was said by its meaning; its search for laughter and for gestures picked up quiet smiles and invented actions in earlier tests, so the skill does not ask Vivu for them.
- Run each query with vivu_search_videos (project_id, query, mode "precise", maximum_results 10). Precise is the mode that returns time ranges; fast only returns whole files and has no place here. Ten is above the number of times one topic usually comes up in the cut stretches; a list that stops at exactly 10 hit the cap, so raise it and rerun that query. It returns a job ID. Call vivu_get_search_results until complete is true; each status call can wait up to 45 seconds, so a pending search is not a stalled one. Save each completed result as results/MOMENT_ID.json. Show the result page link in the live reply only; it expires after four hours, so it never goes into moments.csv or any file.
- Check every window against the transcript: read the .srt lines between start_ms and end_ms. The window is real when those lines are about the moment's topic. The reason text is a paraphrase, not a transcript, so it never settles a verdict and never becomes a quote. Record every verdict in state.json; rejected windows stay in results/ and are counted.
These are the queries from our test run, on two episodes of a salon industry video podcast. They show the pattern: say what was being talked about, in the words the host would use, and nothing about the cue.
| Moment the user remembered | Query (topic only) | Mode | maximum_results |
|---|---|---|---|
| the guest holding up the jar of his own product | the guest explains the hair styling product he makes and sells himself, and why he made it | precise | 10 |
| a man walking into the shot as the "no product" example | someone is shown as an example of a hairstyle done with water only and no product, for the guest's magazine | precise | 10 |
| the guest flipping his magazine open when barbers came up | the guest says barbers insult him and send him threats over his YouTube video | precise | 10 |
| the close-up on the host in the red cap during the money talk | the hosts talk about money in a stylist's job: earning more, commission pay, how much they can charge, and bosses who want their staff to bring in more money | precise | 10 |
| the photos of guys with dyed beards | the hosts look at photos of men with mermaid colored hair and beards dyed green, blue or purple | precise | 10 |
| the big laugh during the salon stress talk | the hosts talk about how stressful working in a salon can be and what the owner or the team can do about it | precise | 10 |
Why this chain: the topic search narrows a whole cut to a few time ranges; the transcript check throws out ranges that only brush the topic; Step 6 then looks for the cue in what is left, from what was said to what was seen or heard. The money wording names commission pay and bosses because the first, shorter wording ("how stylists can make more money, grow their income and what they can charge") missed most of the places the topic came up; that wording was tuned after 1 rewording on the same corpus.
Worked example from our test run
In our test run we used 2 public episodes of a Creative Commons video podcast for hairstylists, a 15 minute stretch from each, 30 minutes in total: a trade show booth interview with a guest stylist, and a group of hosts at a table talking about salon stress, money and hair trends. From the start of the upload to both stretches ready took 14 minutes.
Across the six final queries Vivu returned 11 windows: 11 real on the transcript, 0 false, and 3 topic passages missed against our list written before searching. An entry on that list had its time corrected after the search (an arithmetic slip); left uncorrected, a real stress window would have counted as false. The misses were brief mentions inside other discussions: a passing line about barbers before the long barber passage, and 2 short money lines inside the stress talk. The first money wording returned 1 window and missed 3 of the money passages; the final wording returned 3 windows and missed 2, after 1 rewording on the same corpus. The guest's product search returned 2 windows, and the second was a teaser clip before the interview that repeats his answer word for word. A color product news item with the same color theme as the dyed beards was not returned.
Windows ran from 15.5 to 97 seconds. That width matters for the next step: a window is a place to look, not the moment.
Done when every window of every moment's search has a verdict in state.json and the user has seen the real and rejected counts per moment.
Step 6: Find the remembered cue inside the checked windows
Goal: pick, for each moment, the one checked window where the remembered cue happens, and pin its seconds.
- For a sound cue (laughter, a shout, applause), measure short term loudness once per cut, ten values a second:
ffmpeg -v error -i cuts/CUT_NAME.mp4 -vn -af "ebur128=metadata=1,ametadata=mode=print:key=lavfi.r128.S:file=loudness/CUT_NAME.txt" -f null -
Then list the loudest seconds inside each checked window. Save this as peaks.py and run python3 peaks.py loudness/CUT_NAME.txt START_S END_S:
import re, sys
path, a, b = sys.argv[1], float(sys.argv[2]), float(sys.argv[3])
t, rows = None, []
for line in open(path):
m = re.search(r"pts_time:([\d.]+)", line)
if m:
t = float(m.group(1))
m = re.search(r"lavfi\.r128\.S=(-?[\d.]+)", line)
if m and t is not None and a <= t <= b:
rows.append((t, float(m.group(1))))
if not rows:
sys.exit("no loudness values in that range")
vals = sorted(v for _, v in rows)
median = vals[len(vals) // 2]
best = {}
for t, v in rows:
s = int(t)
best[s] = max(best.get(s, -200), v)
top = sorted(best.items(), key=lambda x: -x[1])[:5]
print("median %.1f LUFS" % median)
for s, v in sorted(top):
print("%d s %.1f LUFS (+%.1f over median)" % (s, v, v - median))
A laugh from the whole table shows up as a stretch several LU above the window's median, often where the transcript has a gap or a [Laughter] tag. Loudness only proposes seconds; confirm each one on frames in the next action. On a recording where every voice is compressed to the same level, peaks can be flat, and the contact sheet is the only check. 2. For a picture cue (something held up, someone walking in, a cut to a close-up, a photo on screen), and to confirm a sound cue, make a two frames a second contact sheet across the window:
ffmpeg -v error -ss START_S -t DURATION -i cuts/CUT_NAME.mp4 -vf "fps=2,scale=320:-2,tile=8x5" frames/MOMENT_ID_wN_%02d.png
START_S is start_ms / 1000, DURATION the window length in seconds, N the window number (each sheet covers 20 seconds; a longer window writes _01, _02 and so on, and Claude looks at every sheet for that window). Keep the quotes around the filter. Look at the sheets and find the frames where the cue happens: the object in a hand toward the camera, a second person entering the frame, faces breaking into laughter, the studio shot replaced by a photo. 3. Pick the window and the seconds where the cue is. When the cue shows up in more than one checked window (the same close-up camera, a second laugh), list both and let the user choose; do not pick silently. Extract one full size frame at the cue for the table:
ffmpeg -v error -ss SECONDS -i cuts/CUT_NAME.mp4 -frames:v 1 -q:v 3 frames/MOMENT_ID_best.png
- When the cue is not in any window, widen before giving up. Start the contact sheet 10 seconds before the window, and read the transcript on both sides of each window: if the same topic keeps going past the window edge (or fills the gap between two windows), make a sheet over that stretch too. Windows mark where the topic is densest, and the remembered cue can sit just outside. Use the same command with a new start, length and name, for example a stretch from WIDE_START_S lasting WIDE_DURATION seconds:
ffmpeg -v error -ss WIDE_START_S -t WIDE_DURATION -i cuts/CUT_NAME.mp4 -vf "fps=2,scale=320:-2,tile=8x5" frames/MOMENT_ID_wide_%02d.png
WIDE_START_S is 10 seconds before the window start (or the end of the earlier window, for a gap), and WIDE_DURATION runs to the end of the stretch. Again, look at every numbered sheet it writes. 5. Mark the moment found (topic checked in the transcript and cue seen), candidate (topic checked, cue not seen even after widening: give the checked windows so the user can scrub them), or not found (no checked window for the topic).
In our test run the remembered cue turned up for all of our moments, but not always inside a window. The jar, the man walking in, the dyed beard photos and the laugh were inside a returned window. The magazine flip was just before its window started and showed up on the sheet started a few seconds early. The red cap close-up was in the gap between two money windows; the transcript showed the money talk running straight through the gap, and a sheet over the gap found it (a close-up of a different host inside one of the windows was the same kind of shot, the wrong person). The jar shot also appeared in the teaser clip window, so that moment had two candidates for the user to pick from. For the laugh, the loudest stretch of the whole cut was the table laughing, but each of the other stress windows also had its own loudest seconds, and there the sheet showed a host with a closer microphone talking. In the booth interview, recorded on close microphones in a loud hall, loudness hardly moved during speech, so loudness alone could not have found a laugh there. peaks.py itself was not exercised in our test run; the loudness output was read by hand.
Done when every moment is marked found, candidate or not found, and every found moment has a frame at the cue.
Step 7: Write moments.csv and confirm the sample
Goal: the user gets one table they can edit from.
- Build one row per moment:
moment,source_file,start_mmss,end_mmss,frame,said_local,loudness_peak_s,cue_seen,status
"the big laugh during the salon stress talk",EP87.mp4,22:01,22:10,frames/m6_best.png,(no words in the transcript during the laugh),544,whole table laughing; one host points across the table,found
- Field mapping:
| Field | Source | If unavailable |
|---|---|---|
| moment | the user's sentence | none |
| source_file | the original episode, from cuts.csv | none |
| start_mmss, end_mmss | cut source start plus the cue seconds from Step 6 (a few seconds before and after), as MM:SS in the original episode | the checked window, marked candidate |
| frame | frames/MOMENT_ID_best.png | blank |
| said_local | transcript lines at the cue | blank |
| loudness_peak_s | the peak second from peaks.py, for sound cues | blank |
| cue_seen | what Claude saw on the frames | UNCERTAIN |
| status | found, candidate or not found | none |
- Show the user one sample row and the mapping, and wait for a yes before writing the rest. Say which field is inferred (cue_seen) and what changes if Claude read the frames wrong: the time points at the right topic but the wrong instance, and the user scrubs the other checked windows listed for that moment.
- Write moments.csv. Under it, list for each moment the other checked windows with their times, so a wrong pick costs a scrub, not a new search. Report the real number of moments found; never fill a row with a guess. In our test run the table was built from the checked windows; the user approval of a sample row was not exercised in our test run.
Done when moments.csv exists with one row per moment and the user has approved the sample row.
Compliance
- Use only recordings of the user's own show. The skill does not download or import other people's videos. Confirm this before Step 2.
- Guests on camera: the user's guest release or recording agreement covers indexing their episode in the user's own Vivu account for editing. Ask before Step 4 whether any guest asked for their episode to stay off third party services, and leave that episode out.
- Minors: if a guest is a minor, go ahead only with a parent or guardian's consent for the recording; otherwise leave that episode out.
- Who laughed or who walked in comes from the transcript and what the user tells Claude. The skill does no face recognition and does not identify anyone.
- The cuts stay in the user's Vivu project until the user deletes them. The skill never calls vivu_delete_video or vivu_delete_project unless the user asks, and confirms first.
- Nothing is posted or sent. moments.csv and the frames stay in the working folder.
Known failure modes
| Symptom | Cause | Fix |
|---|---|---|
| (observed) the remembered cue is a few seconds before the window starts | a window marks where the topic is densest, not where the moment is | start the contact sheet a little before the window and look again |
| (observed) the cue is in neither of two windows for the same topic | the topic runs straight through a gap between two returned windows | read the transcript between the windows; if the topic continues, make a sheet over the gap |
| (observed) the cue shows up in two windows | the show opens with a teaser clip that repeats a later shot | list both on the table and let the user pick; do not choose silently |
| (observed) the loudest seconds in a window are someone talking, not a laugh | one host sits closer to the microphone | confirm every loudness peak on a contact sheet; a laugh shows faces breaking up |
| (observed) loudness barely moves during speech, so no laugh stands out | close microphones in a loud room, or heavy compression on the mix | find the laugh on contact sheets across the checked windows |
| (observed) the first wording misses most places a topic came up; in our test run it returned 1 window and missed 3 | the query described the topic too narrowly | name the sub-topics the user remembers (for money: commission, what to charge, bosses) and rerun once |
| (observed) a passing mention of the topic is not returned | Vivu favors longer passages about a topic | if the cue is not found, search the transcript for the topic words and check those seconds too |
| asking Vivu for "laughing" returns quiet smiles and invented reactions | Vivu's search reads what is said more reliably than how loud or how people react | search the topic, find the laugh locally (Step 6) |
| "has not granted vivu.write" | Vivu connected read only | the user reconnects Vivu with write access |
| upload page asks to sign in or shows an error | the one time upload link expires after 180 seconds | request a new link right before opening it |
| a cut is rejected by the browser upload tool | Claude in Chrome takes at most 10 MB per upload call | the user adds that cut in the Vivu web app; never split or recompress it |
| a result's file name does not match cuts.csv | Vivu replaced spaces or punctuation in the name | name cuts with letters, digits and "_" only, and match on the cut name |
| ffmpeg stops with "Error opening output files: No such file or directory" | the output folder does not exist yet | run the mkdir -p line in Step 2 first |
| a quote on the table does not match what was said | it was copied from the reason text, which paraphrases | take every quote from the transcript |
| a list stops at exactly maximum_results | the cap cut the list short | raise maximum_results and rerun that query |