---
name: inspection-findings-list
description: "Turn narrated home inspection walk videos into a numbered findings list with Vivu: time, your words, room, and a frame of each spot. Use after an inspection, before writing the report."
---Inspection findings list from narrated walk videos
This skill takes the videos a home inspector or facility technician recorded while walking a property and talking through what they saw (two to four files, each 10 to 20 minutes), plus a transcript of each file, and returns a findings list: one row per problem the inspector called out, with the file, the MM:SS, the inspector's own words copied from the transcript, the room or part as the inspector named it, a frame showing the spot, and empty "in report" and "severity" columns for the inspector. Claude measures the files, prices them against the user's Vivu plan, uploads them to a private Vivu project, runs one broad search for spoken findings and three by category (water, structure, electrical), reads the transcript inside every returned window to split it into single findings and drop the windows where nothing was wrong, and picks a frame for each finding from a half second contact sheet.
The value is in the frame. A narration says "right here", "see this", "you can tell it's not secure": what the problem looks like and which corner of which wall only the picture shows, and the report needs that picture. A transcript alone tells you what was said, not what the camera was pointed at; the picture alone does not tell you which patches the inspector called a defect and which ones they called cosmetic and left off the report. The hard part is that findings come in runs (a whole kitchen's worth in a minute or so), and Vivu often returns one wide window around several of them, so the transcript does the splitting and the list says plainly how many findings came from search and which the inspector should look for by hand.
When to use
Use when someone says "go through my inspection videos and list everything I flagged", "pull the findings and a photo for each from today's walkthrough", "I narrated the whole house, make me a defect list with timestamps", or "which problems did I mention in these move out walk videos?".
Not for these:
- A question about one moment in one video ("where did I talk about the water heater?"): search Vivu directly and read the transcript around the hit.
- Deciding severity, cost, or who is responsible. The list leaves those columns empty for the inspector.
- Footage nobody narrated. Without spoken findings this skill has nothing to find; Vivu's visual search for building conditions was unreliable in our test runs of other skills, so do not substitute it.
- Comparing the site against a schedule or a punch list; that is a different job.
Working principles
- Report measured numbers, not estimates. When a number is an estimate, say so.
- Nothing is a finding until the transcript inside the window shows the inspector saying something is wrong. Vivu windows are candidates; quotes come from the transcript, never from Vivu's reason text, which paraphrases.
- Stop and tell the user when a required capability or tool is missing. Do not guess around it.
- Ask the user before anything that is expensive to redo (the file list and its index minutes, before upload) and before writing the full list (one sample row first). The skill sends nothing to anyone.
- Never judge severity or responsibility. A frame Claude labels "part in view" is a suggestion for the inspector to confirm, not a verified photo.
What you need before starting
Check each item at the start of the run and tell the user plainly what is missing before doing anything else.
| Requirement | Why | How to check |
|---|---|---|
| Vivu connector with write access | create a private project, open its upload page, search | vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the user reconnects Vivu and allows write access |
| A shell on the user's computer with ffmpeg and ffprobe | measure the files, cut contact sheets and frames | ffmpeg -version, ffprobe -version |
| The inspection videos as local files | everything is read from the user's own files | ls INSPECTION_DIR shows the files |
| A transcript of each video with timestamps | the transcript decides what is a finding and is the source of every quote | a caption file (.vtt or .srt) next to each video, or a local speech to text tool such as whisper.cpp (whisper-cli --help); if neither exists, stop and ask |
| An upload path | move the files into Vivu | vivu_open_upload_page plus the user's own browser, or a browser tool that can attach local files |
This skill needs Claude Code on the user's computer (the terminal or the Code tab of Claude Desktop), because it reads local video files and runs ffmpeg. Claude on the web and Cowork's cloud sessions cannot reach the files. It downloads nothing, so no residential IP is needed, and nothing recurs, so no scheduler is involved; run it again for the next property.
Inputs to collect
Ask for anything missing, most important first.
- INSPECTION_DIR: the folder with the walk videos for one property. Required.
- The rooms or areas the inspector expects in the list, if any. Default: none; rooms come from what was said.
- Extra categories beyond water, structure and electrical (for example "roof" or "HVAC"). Default: none.
- WORK_NAME: a short name for the working folder and the Vivu project. Default: the property address without the house number, plus the date, such as maple-st-2026-10-02.
Files and state
Keep everything in one working folder:
findings-WORK_NAME/
config.json queries, Vivu project id, inspection folder
videos.csv file, duration_s, video_id, transcript_file
transcripts/ one .vtt or .srt per video
results/ one JSON per query, without the result page link
sheets/ contact sheets around each finding
frames/ the full frame picked for each finding
candidates.csv every window Vivu returned, with the transcript verdict
findings.csv one row per finding
state.json steps done, files uploaded, queries searched with job ids, windows checked
Run the commands from inside findings-WORK_NAME/ so the relative paths resolve.
A rerun reads state.json first and skips finished steps: a file marked uploaded is not uploaded again, a query with a saved result is not searched again, and a window with a recorded verdict is not checked again.
Step 1: Check the Vivu connector and the setup
Goal: confirm every requirement before spending anything.
- Call vivu_get_account. If the tool does not exist, tell the user to add the Vivu connector in Claude (
https://mcp.vivu.ai/mcp) and stop. If the account does not showcan_create_projects: true, or a later write call fails with "has not granted vivu.write", ask the user to reconnect Vivu and allow write access, then stop until they have. - Run
ffmpeg -versionandffprobe -version. - Ask for INSPECTION_DIR and list it. Check that each video has a transcript (Step 2).
- Confirm the Compliance items with the user: the footage is theirs, the owner or occupant agreed to the recording, and people, papers and house numbers seen on camera stay out of the report.
In our test run vivu_get_account returned can_create_projects: true and both tools were installed. The footage was public test material, so the user's Compliance confirmation was not exercised in our test run.
Done when vivu_get_account shows can_create_projects: true, both tools print their versions, the user has confirmed the Compliance items, and INSPECTION_DIR lists the videos.
Step 2: Transcribe and preview the findings
Goal: a timestamped transcript for every video, and a free preview of how many findings to expect, before anything is indexed.
- If a caption file already sits next to a video (some cameras and phone apps write one), copy it into transcripts/. Otherwise transcribe locally. With whisper.cpp, convert the audio to 16 kHz mono first, then write a .vtt:
ffmpeg -v error -y -i INSPECTION_DIR/FILE -ar 16000 -ac 1 transcripts/NAME.wav
whisper-cli -m MODEL_PATH -f transcripts/NAME.wav -ovtt -of transcripts/NAME
FILE is one video, NAME its file name without the extension, and MODEL_PATH the user's whisper model (for example ggml-base.en.bin). Transcription runs on the user's computer, so the audio does not leave it.
- Count the lines that sound like findings, as a preview:
grep -ciE "crack|leak|loose|stain|rot|corro|rust|damage|missing|replace|repair|recommend|not working|moisture" transcripts/NAME.vtt
A count near 0 for a file means the inspector called out little there; tell the user before paying to index it.
This step was not exercised in our test run: local transcription and the preview count were not run. The test videos were public, so YouTube's automatic captions stood in for the transcript in Steps 5 to 7. Those captions garble some words (caulking came out as "cocking", insulation as "halation"), which is why a quote can be marked NOT VERIFIED.
Done when transcripts/ has one transcript per video and the user has seen the preview counts.
Step 3: Price, upload and index
Goal: every video indexed in a private Vivu project, with the user's approval of the cost.
- Measure each file:
ffprobe -v error -show_entries format=duration -of csv=p=0 INSPECTION_DIR/FILE. Write duration_s to videos.csv. - Call vivu_get_usage and show one table:
| This property | Plan allowance | Remaining this month | |
|---|---|---|---|
| Index minutes | sum of duration_s / 60 | from vivu_get_usage | from vivu_get_usage |
| Search credits | 4 precise searches at 5 credits = 20 (estimate), plus 5 for each rewording | from vivu_get_usage | from vivu_get_usage |
Plan facts: Free is $0 a month with 20 index minutes a month and 50 search credits a month; Premium is $30 a month with 180 index minutes and 500 search credits. A precise search uses 5 credits in total. Three 15 minute walks are about 45 index minutes (estimate): Premium covers four such properties a month, Free covers one 15 minute file.
- If the files do not fit, offer levers in this order: index only the files where the Step 2 preview found findings, index the interior walk first and the rest later, and only then move to a larger plan. Never drop a file the user asked for without saying which one, and never trim or recompress a file to save minutes.
- Call vivu_list_projects and reuse a project named "Inspection WORK_NAME" if one exists; otherwise call vivu_create_project with that name and visibility "private". The default visibility is organization, which shows the project to everyone in the user's Vivu organization; footage of someone's home belongs in a private project.
- Call vivu_open_upload_page with the project ID immediately before uploading. The link expires in 180 seconds and works once, so never post or store it. Give it to the user to open in their own browser and select the files, or open it in a browser tool that can attach local files. Claude in Chrome accepts at most 10 MB per upload call and a 15 minute walk is far larger, so walk videos normally go through the user's own browser or the Vivu web app. Never split or recompress a file to fit.
- Poll vivu_list_videos until every file shows ready. Vivu replaces spaces and punctuation in file names with underscores; match each listed file_name back to videos.csv by that rule, or by duration_ms against duration_s when two names collide. Record each video_id in videos.csv and state.json.
In our test run the 3 walk videos measured 45 minutes and were all ready 13 minutes after the upload started (they waited in a queue behind other uploads first); each duration_ms equaled the ffprobe length. The test account returns no allowance figures, so the plan comparison and the user's approval were not exercised in our test run, and the files went through the upload page with an automated browser, so the user's own browser was not exercised either.
Done when the user has approved the index minutes, and every file in videos.csv shows ready and has a video_id.
Step 4: Search for spoken findings
Goal: a saved list of candidate windows for each query.
Every search is precise: only precise returns a time window, and fast returns whole files with an empty reason, which tells you nothing videos.csv does not. The chain: the broad query suggests windows where the inspector says something is wrong; the three category queries are a second net for findings the broad query merged into a neighbour or skipped; the transcript inside each window then decides what is a finding and splits a window into single findings (Step 5); and a frame from the same moment shows the part the inspector was talking about (Step 6). That last step is the flip from what was said to what was shown.
| Field | Query | Mode | maximum_results |
|---|---|---|---|
| finding_any | each time the inspector calls out a problem while walking the house: something damaged, loose, leaking, missing, not working, not secured or not to code, or a recommendation to fix, replace, add or have a specialist check it | precise | 50 |
| finding_water | the inspector points out a water problem: a leak, a drip, a water stain, water damage, rot, corrosion or a drain issue | precise | 20 |
| finding_structure | the inspector points out a crack, rot or damage in a wall, ceiling, floor, foundation, hearth or framing | precise | 20 |
| finding_electrical | the inspector points out an electrical problem: a loose outlet, missing GFCI protection, exposed, spliced or unprotected wiring, or a device that does not work | precise | 20 |
The finding_any wording above came after 1 rewording on the same corpus; the first wording asked for "something is wrong and needs to be repaired, replaced, adjusted or looked at by a specialist" and returned fewer windows. The phrases "each time" and "not secured or not to code, or a recommendation to ... add" pulled in safety items and missing devices.
Run vivu_search_videos (project_id, query, mode "precise", maximum_results as in the table) for each row. vivu_search_videos returns a job ID; call vivu_get_search_results until complete is true. Each status call can wait up to 45 seconds, so a pending search is not a stalled one. Save the results to results/FIELD.json without the result page link, and show the result page link only in the live reply: it expires after four hours, so it never goes into findings.csv or candidates.csv.
maximum_results is also the recall ceiling. The broad query uses 50 because a dense walk can have a finding every minute or so, and a value below the real count silently cuts findings; raise it toward 100 for a larger property. Use the inspector's own words for categories they care about (add a "roof" or "HVAC" row the same way, at 5 credits each).
In our test run every search was complete on the first call to vivu_get_search_results. The standing searches come to 20 credits per property (estimate); our test run used 25 credits including the rewording (estimate).
Done when every query has a saved result file and its job id is in state.json.
Step 5: Check every window against the transcript and split it into findings
Goal: each window labeled "finding" or "no finding", and every finding as its own row, before anything reaches the list.
- For each window, convert start_ms and end_ms to seconds and read the transcript lines whose times fall inside it, plus 5 seconds on each side (windows often start a moment after the sentence begins). In a .vtt file the times are on the line above the text, so
grep -n "^00:MM:" transcripts/NAME.vttfinds the lines for minute MM. - Label the window:
- finding: the inspector says something is wrong, missing, loose, leaking, damaged or not working, or recommends a repair, replacement or a specialist;
- no finding: the inspector checks something and says it is fine, calls it cosmetic or not reportable ("would not go on a report"), says a fix is not worth it, or is explaining a method or giving advice.
- Split a "finding" window into one finding per problem. A window often holds several findings said within a minute or two of each other; give each its own sentence time (the start of the sentence that names it). The same sentence returned by several queries is one finding, not several.
- Write every window into candidates.csv with its query, its label, the transcript line that decided it, and the finding numbers it produced. Count the "no finding" windows as false positives for that query. Vivu's reason text never decides the label.
- Read the transcript of each file once from start to end and list the findings no window covered. These are findings the search missed; put them in the list too, marked source "transcript only", so the list does not depend on search alone. A full read can still miss a remark said under background noise, so tell the user the list covers what the transcript shows.
Worked example from our test run
The test corpus was 3 public narrated walk videos, 45 minutes in all: a vacant condo with a handheld camera and live narration, an occupied older house where the inspector calls out a problem every minute or so, and a house about a century old walked by an inspection school, available only at low resolution. Before searching we wrote down from the captions every problem the inspectors called out, plus the moments where an inspector looked at something and said it was fine or not worth reporting.
The first broad wording returned 13 windows, 13 real, and missed 15 findings. After 1 rewording on the same corpus, the broad query returned 16 windows: 15 real and 1 false, with 12 findings missed. The category queries returned 22 windows: 21 real and 1 false. The water query missed 2 water findings; the structure and electrical queries missed none in their categories, and between them they recovered findings the broad query had skipped (laundry valves, a floor drain, an S trap, rusted pipe supports, loose outlets, a taped knob and tube splice).
Both false windows are moments an inspector dismissed out loud: a water heater drain pan he said was not worth adding, and a wall patch he called cosmetic and not reportable. The reason text described both as findings, which is why the transcript, not the reason, labels each window.
Windows often hold several findings. One window ran from 10:47 to 12:13 in the occupied house and held every kitchen finding from the dishwasher mounting to the oven's missing anti-tip device, while its reason named only the dishwasher. Splitting by the transcript turned it back into separate rows.
Still missed after all the searches: a window lock that was missing, a toilet seat that needed adjusting, missing closet doors, the water heater draft test, the furnace return air, and old copper with possible lead solder. Short remarks about small parts and long technical explanations were the ones search did not return; the full transcript read in item 5 found all of them. A vivu_get_video_summary call on the kitchen returned one long segment that left out the oven's anti-tip device and the gap around the microwave duct, so it is not a substitute for the transcript.
Done when every window in results/ has a label and a transcript line in candidates.csv, every finding has a number and a sentence time, and the transcript only findings are listed.
Step 6: Pick a frame for each finding
Goal: for each finding, one frame where the camera is on the part the inspector is talking about, or a plain note that there is none.
- Cut a contact sheet from 2 seconds before the sentence time to 10 seconds after it, one frame every half second, because the camera is often still moving onto the part while the sentence starts:
ffmpeg -v error -y -ss SHEET_START -t 12 -i INSPECTION_DIR/FILE -vf "fps=2,scale=240:-1,tile=6x4" -frames:v 1 sheets/fNN_NAME.png
SHEET_START is the sentence time minus 2, in seconds; NN is the finding number. Tile k (counting from 0 at the top left) is at SHEET_START plus k/2 seconds. 2. Pick the tile where the part is largest and sharpest, and extract it as a full frame:
ffmpeg -v error -y -ss SECONDS -i INSPECTION_DIR/FILE -frames:v 1 -q:v 3 frames/fNN_NAME.png
SECONDS is SHEET_START plus k/2 for the tile you picked. Look at the full frame; a tile can show a blur that the full frame does not. 3. Label frame_check: "part in view" when the frame shows the thing named in the sentence, "part not in view" when the camera points elsewhere for the whole sheet, "can't tell" when the part is there but the defect itself (a hairline crack, a missing screw) cannot be made out, and "nothing to show" for recommendations about something absent (no CO detector, no drain pan). A "part not in view" finding can still go in the report; the inspector then uses a photo from their camera.
In our test run the half second sheet starting just before the sentence showed the named part for most findings. The exceptions were useful to know: the inspector said the toilet seat needed adjusting while the camera faced the wall; the hearth crack in the low resolution video, a crack around a junction box, a taped splice in a dark basement and a rusted support could not be made out; and the CO detector and pressure reducer recommendations had nothing to show.
Done when every finding has a frame path and a frame_check value, or "nothing to show".
Step 7: Write the findings list
Goal: a list the inspector can write the report from.
- Show the user one sample row and the field mapping, and wait for a yes before writing the rest:
no,file,mm_ss,quote,room_or_part,category,frame,frame_check,source,in_report,severity
13,walkB_kitchen.mp4,11:04,"I know there's water damage down there",under the kitchen sink,water,frames/f13_walkB_kitchen.png,part in view,finding_any; finding_water,,
| Field | Source | If unavailable |
|---|---|---|
| no | finding number, in file and time order | none |
| file, mm_ss | the file and the sentence time | none |
| quote | copied from the transcript | NOT VERIFIED when the transcript line is garbled |
| room_or_part | the room or part the inspector names in or just before the sentence | UNCERTAIN; never guessed from the picture |
| category | water, structure, electrical or other, from what was said | other |
| frame, frame_check | Step 6 | "nothing to show" |
| source | search (the query fields whose windows covered it) or "transcript only" | none |
| in_report, severity | left empty for the inspector | stay empty |
- Write findings.csv. Keep a room name only when the inspector said it; a continuous walk often names few rooms, and a wrong room in a report is worse than a blank.
- Tell the user the counts: findings from search, findings found only by reading the transcript, windows set aside as "no finding", and findings with "part not in view". Say plainly that the search misses some findings that are said in passing, and that the transcript read in Step 5 is what catches them.
In our test run findings.csv had one row per finding called out in the three walks, each with a caption quote and a frame or a frame_check note; a few rooms are UNCERTAIN because the inspector never named them. The user's approval of the sample row was not exercised in our test run.
Done when findings.csv has one row per finding, the user has approved the sample row, and the counts are reported.
Compliance
- Use only footage the inspector or technician recorded themselves, with the owner's or occupant's agreement to recording. Confirm this before Step 3.
- People, children, documents, mail, screens and house numbers seen on camera stay out of the report. When a picked frame shows any of these, pick another tile or tell the inspector to crop or blur it.
- The videos go to a private Vivu project in the user's account and stay there until the user deletes them. Delete a project or video only when the user asks, and confirm first.
- Transcription runs on the user's computer; the skill sends nothing to anyone and writes only local files.
- The skill does no face, logo or identity recognition, and it does not judge severity, cost or responsibility. Those columns stay with the inspector, who signs the report.
Known failure modes
| Symptom | Cause | Fix |
|---|---|---|
| (observed) a window the reason calls a finding is something the inspector dismissed | the inspector said "I don't see a point in putting a drain pan on it right now"; the reason said the water heater has no drain pan | label from the transcript in Step 5 and count it as a false positive; never accept a window because of its reason |
| (observed) a cosmetic remark comes back from the structure query | the inspector called a wall patch cosmetic: "something like this would not go on a home inspection report" | label it as no finding; the list follows what the inspector decided to report |
| (observed) one window holds several findings and the reason names one | findings said within a minute of each other come back as one wide window | split the window by the transcript in Step 5, one row per problem |
| (observed) a finding said in one short sentence is not returned by any query | short remarks about small parts (a window lock, a toilet seat, closet doors) gave the search little to match | read every transcript once in Step 5 and add transcript only rows; an empty result does not prove nothing was said |
| (observed) the first page is clean but short | the first broad wording missed 15 findings | use the category queries and the final broad wording from Step 4; for a big property raise maximum_results |
| (observed) the frame for a finding shows a wall, not the part | the inspector talked about the toilet seat while the camera pointed elsewhere | label it part not in view and ask the inspector for a photo from their own camera |
| (observed) the defect cannot be seen in any frame | a hairline crack in a low resolution video, or a dark basement | label it can't tell; do not describe the defect from the narration as if it were visible |
| (observed) a summary segment leaves out findings | vivu_get_video_summary returned one long kitchen segment that skipped the anti-tip device and the duct gap | use segments only as a second signal; the transcript decides |
| a room name is missing or wrong | the inspector never said the room, or said it from the doorway of the next room | write UNCERTAIN; never infer the room from the picture |
| a quote has wrong words | automatic captions or a small speech model mishear trade terms | mark the quote NOT VERIFIED and listen to that second of the video |
| "has not granted vivu.write" | Vivu connected read only | the user reconnects Vivu with write access |
| upload page asks to sign in or shows an error | the one time upload link expires after 180 seconds | request a new link right before opening it |
| a walk video is rejected by the browser upload tool | Claude in Chrome accepts at most 10 MB per upload call | the user opens the upload link in their own browser or adds the file in the Vivu web app |
| a result's file name does not match videos.csv | Vivu replaced spaces or punctuation with underscores | match by the underscore rule or by duration_ms |