SkillsSkill for Claude

Road test scenario shortlists

This skill turns a batch of front camera road test videos, exported from your drive logs as mp4 files, into one CSV per scenario: pedestrians crossing in front of the car, vehicles cutting in close ahead, and cyclists on the road. Claude measures the batch, prices it against your Vivu plan, uploads it to a private Vivu project, runs one precise search per scenario, cuts a half second contact sheet for every window Vivu returns, labels each one looks right, looks wrong or can't tell, and writes every candidate into the CSV with an empty column for your reviewer's verdict. File names and log fields do not say what happened on the road, and exported camera streams have no audio to search, so the answer is only in the pictures. Vivu's search only shortlists candidates: in our test run most cut in candidates and many pedestrian candidates were wrong while Vivu's text described them as the real thing, so every label comes from the frames and a short list never proves the batch had no such moment.

Maintained by Vivu. Updated 2026-09-28.

Download

road-test-scenario-shortlists.zip

12 KB. Unzips to road-test-scenario-shortlists/SKILL.md. Upload the zip as it is in the Claude app, or unzip it into your skills folder for Claude Code.

SHA-256 35221de5111fc987bd29858405c1721165e1da0f26f9bddb46c2f7550485dd1d

At a glance

What the Road test scenario shortlists skill does, where it runs, what it needs, and when it asks
Looks forPedestrians crossing in front of the car, vehicles cutting in close ahead and cyclists on the road, in exported front camera video. File names and log fields do not record them and the camera streams have no audio. Every candidate is checked against the original video on a half second contact sheet, with a zoomed frame for far subjects, before it is labeled, and a reviewer makes the final call.
Runs onClaude Code on your computer (the terminal or the Code tab of Claude Desktop), with a shell, ffmpeg and the road test videos as local files. No residential IP is needed because nothing is downloaded. It uses the Vivu connector and runs once per batch, with no scheduler.
Needs
  • The Vivu connector with write access, to create a private project, open its upload page and search.
  • The road test videos as mp4 files on your computer, because Vivu indexes video files and every contact sheet is cut from the local file.
  • A shell with ffmpeg and ffprobe, to measure the batch and cut contact sheets and zoomed frames.
  • An upload path: the Vivu upload page opened in your own browser, or a browser tool that can attach local files.
  • Someone to review the CSVs, because Claude's labels are a first pass and the reviewer decides.
Your Vivu planIndex minutes equal the length of the batch you approve, sized to about an hour of footage, and each scenario uses one precise search, 15 search credits per batch for the default scenarios (estimate). Before anything is uploaded, Claude measures the batch with ffprobe, shows it next to what vivu_get_usage reports for your plan and waits for your approval. If a search comes back full, getting past the cap means uploading the whole batch again as two projects and paying its index minutes again, and Claude asks first. In our test run 18.3 minutes were indexed and gave 47 candidates to check across the scenarios.
Asks you firstConfirming your rights to the footage and your team's privacy rules before upload, approving the batch and its index minutes before upload, approving the cost of uploading the batch again as two projects, and approving one sample CSV row with its field mapping before the full CSVs are written. Any upload or post of the CSVs needs its own approval of the destination and a sample.

The skill does its video work through the Vivu connector. If the connector is not in your Claude yet, add it first; the skill checks that it is connected before it does anything else.

Before you run it

  • Search results for every scenario are candidates only: in our test run most cut in candidates and many pedestrian candidates were wrong, so Claude checks each one on a contact sheet and a reviewer makes the final call.
  • In our test run Vivu search failed at telling who moved in the picture: most cut in candidates were wrong, so treat the cut in list as a loose pre-filter, and turns across a crosswalk are not searched at all. Far pedestrians and cyclists need a zoomed frame to see, and night, rain and strong backlight footage is untested.
  • Part of the test footage carried a drawn band over the car's planned path and a speed readout that exported camera streams do not have, so results on your plain streams may differ from the test counts.
  • Claude labels each candidate in a single pass, and in our test run a second look changed labels in both directions, so your reviewer goes through every row, including the ones labeled looks right.
  • A search returns a capped number of candidates and misses some, so an empty or short list does not prove the batch has no such moment, and the CSVs carry no recall or precision figures.
  • People on the street did not agree to be filmed: the skill identifies no one, does no face or license plate recognition, and must not be used to collect footage of particular people or of children.
  • Use footage your team owns or has the rights to analyze, and follow your team's privacy rules before upload.
  • Claude in Chrome uploads at most 10 MB per call, so camera files go through your own browser or the Vivu web app, never split or recompressed.
  • The videos stay in your Vivu project until you delete them; the project is created private because the default is visible to your whole organization.
  • Claude writes only local files and does not post, send or upload the CSVs anywhere unless you approve a destination and a sample first.

Start it

Once the skill is installed, ask for the task in your own words. Naming the skill is the most reliable way to have Claude use it. For example:

Here is a folder of last week's front camera road test clips, exported as mp4. Build review lists of pedestrian crossings, cut ins and cyclists so our reviewers can pick eval set candidates.

In Claude Code you can also type /road-test-scenario-shortlists. Claude asks for anything the request leaves out, most important first.

What is inside

  1. When to use
  2. Working principles
  3. What you need before starting
  4. Inputs to collect
  5. Files and state
  6. Step 1: Check the Vivu connector and the setup
  7. Step 2: Measure the batch and price it
  8. Step 3: Upload and index
  9. Step 4: Search each scenario
  10. Step 5: Make a contact sheet for every candidate and label it
  11. Step 6: Write one CSV per scenario and hand it over
  12. Compliance
  13. Known failure modes

The full skill

This is road-test-scenario-shortlists/SKILL.md from the download, as Claude reads it: the frontmatter first, then the instructions.

---
name: road-test-scenario-shortlists
description: "Turn front camera road test videos into a candidate CSV per scenario with Vivu, each window labeled on a contact sheet for human review. Use when building eval sets from drive logs."
---

Road test scenario shortlists with Vivu

This skill turns a batch of front camera road test videos (camera streams your team has already exported from its drive logs as mp4 files) into one CSV per scenario on the team's list: a pedestrian crossing in front of the car, another vehicle cutting in close ahead, and a cyclist on the road. Claude measures the batch, prices it against the user's Vivu plan, uploads it to a private Vivu project, runs one precise search per scenario, cuts a contact sheet at half second steps for every window Vivu returns, labels each sheet looks right, looks wrong or can't tell, and writes every candidate with its label into the scenario's CSV. The CSV keeps an empty reviewer_verdict column for the person who decides what goes into the eval set.

The value is in what only the pictures hold. Road test file names carry a date, a vehicle and a route, and the log fields carry speed and position; whether someone stepped off the curb in front of the car, or a van swung into the lane just ahead, is only in the video, and exported camera streams have no audio or captions to search. Vivu's search is used only to shortlist candidates. In our test run most cut in candidates were cars that had been ahead all along, parked cars or the camera car's own lane changes, and many pedestrian candidates were people on the sidewalk, while the reason text described them as cut ins and crossings. So every candidate is labeled from its contact sheet, rejected candidates stay in the CSV with their label, and a short or empty list never means the batch had no such moments.

When to use

Use when someone asks to pull pedestrian crossings or cut ins out of this week's drives, to shortlist scenario candidates from road test videos for an eval set, or to find which clips in a batch have cyclists close to the car.

Not for these:

  1. A one off question about one clip ("is there a cut in in this file?"): search Vivu directly and look at the frames.
  2. Hundreds of hours at once. Pick a subset with the log fields first (route, time of day, weather tags); one run is sized for about an hour of footage.
  3. Frame accurate labels, bounding boxes or ground truth files: use the team's annotation tool. This skill makes review lists, not labels.
  4. Judging whether the car drove correctly or whether anyone broke a traffic rule. The skill lists what happens in the picture and nothing more.
  5. Turns across a crosswalk with people on it. The skill does not search for them: in our test run a turn search returned 8 candidates, 3 real and 5 false, and its reasons described turns where the car went straight. Reviewers can note turns while they go through the pedestrian list.
  6. Night, rain or strong backlight footage as the main material: untested. The skill can run, but treat its output as unmeasured and have the reviewers look at every row.

Working principles

  1. Report measured numbers, not estimates. When a number is an estimate, say so.
  2. Nothing is verified until it has been checked against the source video. Vivu's results are candidates, Claude's labels are a first pass, and the reviewer's verdict is the decision.
  3. Stop and tell the user when a required capability or tool is missing. Do not guess around it.
  4. Ask the user before anything that is expensive to redo (the batch and its index minutes, before upload) and before writing the full CSVs (one sample row first). The skill posts and sends nothing; if the user asks Claude to put the CSVs somewhere, ask again before each write.
  5. Labels come from frames, never from Vivu's reason text, which describes crossings and lane changes that are not in the picture.
  6. The CSVs keep every candidate. They carry no recall or precision figures, because a search caps how many candidates come back and misses some, so an empty or short list does not prove the footage has no such moment.

What you need before starting

Check each item at the start of the run and tell the user plainly what is missing before doing anything else.

Requirement Why How to check
Vivu connector with write access create a private project, open its upload page, search vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the user reconnects Vivu and allows write access
The road test videos as mp4 files on the user's computer Vivu indexes video files, and every sheet is cut from the local file ffprobe -v error -show_entries format=duration -of csv=p=0 FILE prints a length for each file; logs such as ROS bags or MCAP files must be exported to video by the team's own tools first
A shell on that computer with ffmpeg and ffprobe measure the batch, cut contact sheets and zoomed frames ffmpeg -version, ffprobe -version
An upload path move the files into Vivu vivu_open_upload_page plus the user's own browser, or a browser tool that can attach local files
Someone who reviews the CSVs Claude's labels are a first pass; the reviewer decides ask who fills reviewer_verdict; the default is the user

This skill needs Claude Code on the user's computer (the terminal or the Code tab of Claude Desktop), because it reads local files and runs ffmpeg. Claude on the web and cloud sessions cannot reach the files. It downloads nothing, so no residential IP is needed. Nothing recurs, so no scheduler is involved; run it once per batch.

Inputs to collect

Ask for anything missing, most important first.

  1. BATCH_DIR: the folder that holds the exported mp4 files for this batch. Required.
  2. The scenario list and what counts for each. Default: the three scenarios in Step 4 with the label rules in Step 5. If the reviewers define a scenario differently (for example, a pedestrian crossing any lane of a junction counts, not only the camera car's lane), their definition wins; write it into config.json before labeling.
  3. BATCH_NAME: a short name for the batch, such as a date and route. It names the working folder and the Vivu project. Default: the name of BATCH_DIR.
  4. Batch size. Default: up to 60 minutes of video, or what remains of the plan this month, whichever is smaller.
  5. Who reviews the CSVs. Default: the user.

Files and state

Keep everything in one working folder next to BATCH_DIR:

scenarios-BATCH_NAME/
  config.json      BATCH_DIR, scenarios with query, maximum_results and label rules, Vivu project id (plus the A and B project ids if Step 4 needs them)
  files.csv        source_file, duration_s, listed_name, video_id
  results/         search results without the result page link, one JSON per scenario
  sheets/          one folder per scenario with contact sheets and zoomed frames
  SCENARIO.csv     one per scenario: every candidate, Claude's label, empty reviewer_verdict
  state.json       steps done, files uploaded, scenarios searched with job ids, scenarios cut off, candidate keys labeled

The commands in Steps 2 to 5 run from the folder that holds both BATCH_DIR and scenarios-BATCH_NAME/, so the relative paths in them resolve.

A rerun reads state.json first and skips finished steps. A file already marked uploaded to a project is never uploaded to that project again, a scenario with a saved result is not searched again unless Step 4 moves it to two new projects, and a candidate key (video_id, start_ms and scenario) that already has a label is not labeled again. The next batch gets its own folder and its own Vivu project, so its searches do not return last week's candidates.

Step 1: Check the Vivu connector and the setup

Goal: confirm every requirement before spending anything.

  1. Call vivu_get_account. If the tool does not exist, tell the user to add the Vivu connector in Claude (https://mcp.vivu.ai/mcp) and stop. If the account does not show can_create_projects: true, or a later write call fails with "has not granted vivu.write", ask the user to reconnect Vivu and allow write access, then stop until they have.
  2. Run ffmpeg -version and ffprobe -version.
  3. List BATCH_DIR and run the ffprobe length command on one file.

Done when vivu_get_account shows can_create_projects: true, ffmpeg and ffprobe print their versions, and ffprobe prints a length for a file in BATCH_DIR.

Step 2: Measure the batch and price it

Goal: an indexed set the user's plan can pay for, approved before upload.

  1. Run the ffprobe length command on every mp4 in BATCH_DIR and write files.csv with source_file and duration_s. Add listed_name: the file name with each space, bracket and plus sign replaced by an underscore, because that is how Vivu will list it.
  2. Call vivu_get_usage and show one table:
This batch Plan allowance Remaining this month
Index minutes sum of duration_s / 60 from vivu_get_usage from vivu_get_usage
Search credits 5 per scenario, 15 for the default three (estimate), plus 5 for each rewording from vivu_get_usage from vivu_get_usage

Plan facts: Free is $0 a month with 20 index minutes a month and 50 search credits a month; Premium is $30 a month with 180 index minutes and 500 search credits. A precise search uses 5 credits in total. Whether a larger maximum_results costs more credits is untested.

Also tell the user how much checking the batch will bring, because every candidate gets its own contact sheets and a row the reviewer goes through. In our test run 18.3 minutes of city driving gave 47 candidates across the scenarios. An hour of similar driving would give roughly three times as many rows (estimate), so the reviewer's time belongs in the approval too.

  1. If the batch does not fit, offer levers in this order: pick a smaller subset with the log fields (the routes, times of day or conditions the eval set needs most), split the batch across two runs or two months, and only then a larger plan. Never drop files the user asked for without saying which ones.
  2. Confirm the Compliance items with the user.

The account used in our test run returns no allowance figures, so the comparison with a real plan and the user's approval were not exercised in our test run.

Done when the user approves the file list and its index minutes, and config.json records both.

Step 3: Upload and index

Goal: every file in the batch indexed in a private Vivu project for this batch.

  1. Call vivu_list_projects and reuse a project named "Road test BATCH_NAME" if one exists. Otherwise call vivu_create_project with that name and visibility "private". The default visibility is organization, which shows the project to everyone in the user's Vivu organization, and road test footage is usually confidential.
  2. Call vivu_open_upload_page with the project ID immediately before uploading. The link expires in 180 seconds and works once, so never post or store it. Give it to the user to open in their own browser and select the files, or open it in a browser tool that can attach local files. Claude in Chrome accepts at most 10 MB per upload call and exported camera files are usually far larger, so they normally go through the user's own browser or the Vivu web app. Never split or recompress a file to fit.
  3. Poll vivu_list_videos until every file shows ready. Match each Vivu file name to files.csv by listed_name and record its video_id there and in state.json.

In our test run the 11 videos (18.3 minutes) were all ready by the first check, 591 seconds after the upload link was requested. They went through the same upload page with an automated browser, so opening the link in the user's own browser was not exercised in our test run.

Done when every file in files.csv shows ready and has a video_id.

Step 4: Search each scenario

Goal: a saved list of candidate windows for every scenario.

Field Query Mode maximum_results
ped_crossing a pedestrian whose whole body is on the asphalt of the road in front of the camera car, part way between the two curbs, stepping across the lane the car is driving in precise 100
cut_in a vehicle crosses the lane marking from the neighboring lane and ends up in the camera car's lane directly ahead, where no vehicle was a moment before; not parked cars, not cars already ahead, not cross traffic at a junction precise 40
cyclist a person riding a bicycle, or a wheelchair user, on the road close to or in front of the camera car precise 100

The searches run side by side, one per scenario, because each scenario is its own list for the reviewers. Within a scenario the chain is: the precise search finds candidate windows, the contact sheet shows what happened in each window, Claude labels it, and the reviewer decides. One link runs across the lists: a cut in or pedestrian candidate that turns out to be a rider is labeled looks wrong in its own list with a note, because the cyclist search is the one that covers riders. Every search is precise because only precise returns a time window; fast returns whole files with an empty reason, and the batch is already a chosen subset. Exported camera streams have no audio and no on screen text about the scene, so there is no second modality to switch to.

How far to trust each search (numbers in the worked example after Step 5): cyclist results were mostly right; pedestrian results were right more often than not but included people on the sidewalk and people crossing other lanes; cut in results were mostly wrong, so treat the cut in list as a loose pre-filter. Vivu's picture search in road footage finds things that are there (a person, a bike, a car ahead) better than it judges who moved from where to where. That is why every result for every scenario goes through Step 5.

maximum_results is also the ceiling on how many candidates come back. Pedestrians and cyclists are common in city driving, so a batch of up to an hour gets 100 (the limit) for those two and 40 for cut ins, which are rarer. A search that returns exactly its maximum_results was cut off: the batch holds more candidates than one search can return. A search covers a whole project and cannot be limited to some of its files, so every file left in this project keeps sharing its cap. To get past the cap, upload the whole batch again as two new projects of about half the files each and search the scenario in both; a second search is capped too and can still miss moments. Before doing that, show the user what it costs, the index minutes of the whole batch again plus 5 credits per scenario per project (estimate), and wait for approval as in Step 2. Create both projects as in Step 3, private and named "Road test BATCH_NAME A" and "Road test BATCH_NAME B", with each file in only one of them. Then build results/FIELD.json for that scenario from the two new projects only, keep their ranks apart (for example A01 and B01), and set the first project's list for it aside, so no moment is listed twice. If a search in either new project also returns exactly its maximum_results, mark the scenario cut off as well. If the user declines or the plan cannot pay for it, keep the list as it is, mark the scenario cut off in state.json, and say so in the summary. A later batch that is likely to be this dense can go into two projects at upload, each file uploaded once, so nothing is indexed twice.

In our test run every search used maximum_results 40 on 18.3 minutes and none came close to that limit; the larger value for pedestrians and cyclists and the move to two new projects were not exercised in our test run.

Run each query with vivu_search_videos (project_id, query, mode, maximum_results). It returns a job ID. Call vivu_get_search_results until complete is true; each status call can wait up to 45 seconds, so a pending search is not a stalled one. Save each completed result as results/FIELD.json without its result_page_url field, and record the job_id in state.json. Show the Vivu result page link in the live reply only; it expires after four hours, so it never goes into a saved file or a CSV.

The pedestrian and cut in wordings were tuned on our test corpus. If the team changes a wording, the new wording is untested until its first candidates have been through Step 5.

Done when every scenario has a completed result saved in results/, and every scenario whose search returned exactly its maximum_results has either been searched again in two new projects holding the whole batch, with the user's approval and its results/FIELD.json built from those two projects only, or is marked cut off in state.json.

Step 5: Make a contact sheet for every candidate and label it

Goal: every candidate labeled looks right, looks wrong or can't tell from what its frames show.

For each result in results/FIELD.json, cut contact sheets of its window from the local file at two frames a second:

mkdir -p scenarios-BATCH_NAME/sheets/FIELD
ffmpeg -v error -y -ss START -to END -i FILE -vf "fps=2,scale=384:-2,tile=5x4" scenarios-BATCH_NAME/sheets/FIELD/rRANK_STEM_STARTMS_%02d.png

FIELD is the scenario, FILE the local source file (BATCH_DIR/ followed by its source_file), STEM its name without .mp4, RANK the result_number written with two digits (r01, r09, r19) so the files sort by rank, with the project letter in front (rA01, rB01) when Step 4 built the list from two projects, STARTMS the result's start_ms, and START and END the result's start_ms and end_ms divided by 1000. Each sheet holds 20 tiles covering 10 seconds of the window; the last sheet is padded with black tiles. Tile k on sheet n (both counted from 0, tiles left to right and top to bottom) is at START + 10 * n + k / 2 seconds. The tile position gives each frame's time, so the command needs no text overlay. A single frame is not enough: windows are wider than the events, and the event can sit in a few tiles anywhere in the window.

Look at every sheet of the window. Far subjects are too small to see on a tile, so when a tile shows something small in the distance, or the sheet looks empty but the reason names someone, extract a zoomed full size frame at that second:

ffmpeg -v error -y -ss SECONDS -i FILE -frames:v 1 -vf "crop=iw/2:ih/3:X:ih/3,scale=iw*2:-2" scenarios-BATCH_NAME/sheets/FIELD/rRANK_STEM_SECMS_zoom_SIDE.png

SECONDS is the time from the tile, SECMS the same time in milliseconds, X picks the part of the picture (0 for the left half, iw/4 for the center, iw/2 for the right half), and SIDE is left, center or right to match X. SECMS and SIDE keep the file names unique when you zoom into two parts of the same second. The crop keeps the middle band of the frame, where the road ahead and far subjects are in a front camera view.

Label rules (defaults; the reviewers' definitions from Inputs win):

  1. ped_crossing looks right when a person on foot is on the road surface inside the camera car's lane, or its continuation ahead, at some point in the window, near or far. It looks wrong for people on the sidewalk or waiting at the curb, an empty crosswalk, people crossing a side road, people crossing other lanes of a junction without entering the camera car's lane, and a crossing that starts after the window ends.
  2. cut_in looks right when a vehicle that is in a neighboring lane (or entering from the side) in the early tiles is in the camera car's lane directly ahead, within about two car lengths, in the later tiles. It looks wrong for a vehicle that is ahead from the first tile, a parked car, cross traffic at a junction, a vehicle that passes in the next lane and stays there, the camera car's own lane change (the lane markings slide sideways under it), and bicycles, which belong in the cyclist list.
  3. cyclist looks right when someone rides a bicycle on the carriageway or a cycle lane on the road, ahead of or beside the car. It looks wrong for parked bicycles and riders on the sidewalk. Check that someone is on the bicycle and that it moves between tiles: a bicycle standing at a rack beside an empty bike lane is parked, whatever the reason says.
  4. can't tell when the zoomed frame still does not settle it (the subject is hidden, too small or blurred). The reviewer looks at the source.

Write a note of what the frames show for every candidate, such as "people only on the sidewalks" or "scooter enters from the left, the car follows it". Never copy the reason into the note. Record each candidate key and its label in state.json.

Worked example from our test run

The test corpus was public daytime city driving from front cameras: 11 videos (18.3 minutes) with the audio removed. All but the last were 105.0 second segments cut at fixed offsets, chosen before looking at them: half from an autonomous driving company's published raw drive in London and half from a city drive in Krakow published under a Creative Commons license. The last was a 48.0 second public domain dashcam clip with a close lane change. Crossings, cut ins, cyclists and turns were marked by hand on contact sheets before any search, and events first seen in search results were added to the hand marks afterwards, so the missed counts below leave out anything that neither the hand marks nor any search caught.

The test footage also carried overlays that exported camera streams do not have. The London segments show the publisher's speed readout and a green band drawn over the camera car's planned path, and the dashcam clip shows a timestamp and a purple patch drawn over a car; the Krakow segments have none. The band marks the camera car's own lane, which is what the pedestrian and cut in labels turn on, so it may have helped the search or the labels on the London segments. Results on plain exported streams were not measured separately, so treat the counts below as a guide, not a forecast for your footage.

  1. ped_crossing, with the wording in the table after 2 rewordings on the same corpus: 18 candidates, 11 real and 7 false, and 6 crossings in the hand marks were missed. The false ones were empty zebras, people on the sidewalk beside the camera car's lane or crossing a side road, and people crossing other lanes of a junction without entering the camera car's lane.
  2. The earlier ped_crossing wordings missed fewer. They returned 20 and 17 candidates, with 5 and 3 false, and each missed 3 crossings in the hand marks; one person labeled them, and the second look described below changed that person's labels on the final wording, so these counts may be optimistic. When every candidate is checked by hand, a miss costs more than a false candidate, so a team that cares most about misses can try the first wording, "a pedestrian walks across the road in front of the camera car, on a crosswalk or at an intersection", and check its first candidates through this step before relying on it.
  3. cut_in, with the wording in the table after 2 rewordings on the same corpus: 10 candidates, 2 real and 8 false, and 1 of the 3 cut ins in the hand marks was missed. Only the dashcam clip's cut in was marked before any search, and that clip was added to the corpus because it had one; the motor scooter and the black cab were first found by earlier wordings and then added to the hand marks, so how many cut ins no wording returned is unknown. The real ones were a white car cutting across from the right just before a collision and the motor scooter entering from the left. The missed one was the black cab. The false ones were cars that had been ahead all along, the camera car's own lane changes, another cab ahead pulling over to the curb and a bicycle.
  4. The earlier cut_in wordings did no better: 11 returned with 9 false, then 16 returned with 13 false. Writing exclusions into the query did not help; the reasons repeated the exclusions to describe cars that had been ahead the whole time.
  5. cyclist, first wording: 19 candidates, 15 real and 4 false (bicycles parked at a rack or at the curb, riders on the sidewalk), and none of the cyclists in the hand marks was missed. The search also found real riders the hand marks did not list, so the hand marks are not a full count and the search's own misses are unknown.
  6. A turn search (the car turning across a crosswalk with people on it) returned 8 candidates, 3 real and 5 false, and 1 of the 4 turns in the hand marks was missed; half of those turns were added to the hand marks only after the search found them. In the false ones the car went straight or turned over an empty crossing while the reason described a turn across people, so the skill leaves turns out.
  7. Far subjects were the hard part of labeling. A far pedestrian and several far cyclists could not be seen on the contact sheets and were confirmed only on zoomed full size frames; the first labeling pass missed that pedestrian.
  8. A second, blind look at every window of the three scenarios (a separate judge, with a third look at full resolution frames where the two disagreed) agreed with every cut in label, changed pedestrian labels in both directions, and moved one cyclist window from looks right to looks wrong. People beside the camera car's lane had been labeled looks right, and the far pedestrian looks wrong. The cyclist window, 01:19 to 01:27 in a Krakow segment, held a bicycle parked at a rack beside an empty bike lane; the first pass had taken it for a rider, as the reason did. The counts above are after that second look. A user's batch gets only Claude's single pass, which makes mistakes of these kinds, so the reviewer goes through every row.
  9. The 11 videos were all ready by the first check, 591 seconds after the upload link was requested. Searches took from 42 to 326 seconds to complete. The most any search returned was 20, against a maximum_results of 40. The three CSVs held 47 rows in total.

Done when every candidate in results/ has at least one contact sheet in sheets/ and a label in state.json.

Step 6: Write one CSV per scenario and hand it over

Goal: a review list per scenario that a reviewer can work through without opening Vivu.

Show the user one sample row and the field mapping, and wait for an OK before writing the rest:

source_file,rank,start_mmss,end_mmss,search_reason,contact_sheets,claude_label,claude_note,reviewer_verdict
lon_s0240.mp4,9,00:55,01:03,"A scooter enters the camera car's lane from the left side, crossing into the lane directly ahead of the camera car where there was previously empty space.",sheets/cut_in/r09_lon_s0240_55000_01.png,looks right,"a motor scooter enters the lane from the left close ahead and the camera car follows it",

The row above is from our test run; the user's rows carry their own file names.

Column Source If unavailable
source_file files.csv, matched by listed_name keep the row with Vivu's file name, label it can't tell with the note "no local file matched", and tell the user
rank result_number in results/FIELD.json, with the project letter (A01, B01) when Step 4 built the list from two projects none
start_mmss, end_mmss start_ms and end_ms, written MM:SS (with a decimal when the time has half seconds) none
search_reason the result's reason, copied as is; it is Vivu's description, not evidence blank
contact_sheets the sheet files and zoomed frames from Step 5, separated by semicolons none; the row cannot be labeled without them
claude_label Step 5 (inferred from frames) can't tell
claude_note what the frames show, in Claude's words (inferred) blank
reviewer_verdict left empty for the reviewer empty

claude_label and claude_note are Claude's reading of the frames, not a measurement. If a label is wrong, a real moment sits among the looks wrong rows or a false one among the looks right rows. In our test run a second look changed labels in both directions (see the worked example), so every row keeps its sheets and the reviewer column, and the reviewer goes through every row, the looks right ones included, before anything enters the eval set.

Write FIELD.csv for every scenario with one row per result, sorted by rank, including the rows labeled looks wrong: a reviewer may disagree with a label, and the rejected rows show what the search confuses. The CSV uses the user's own file names and times, never a Vivu result page link, so it stays usable after the link expires.

Then tell the user, per scenario, how many candidates came back, how many Claude labeled looks right, looks wrong and can't tell, and whether the list was cut off at maximum_results (Step 4). Say plainly that an empty or short list does not mean the batch has no such moment. The CSVs and this summary carry counts of candidates and labels only. Do not add recall or precision figures, from our test run or from the batch: Claude's labels are a first pass, and the hand marks needed to measure recall do not exist for the user's batch.

The CSVs stay in the working folder. Passing them to the reviewers is the user's step. If the user asks Claude to upload or post them (a shared drive, a tracker, a chat), name the destination, show the rendered first rows and wait for a yes before each write.

In our test run the CSVs were written from the dry run's results and sheets; showing the sample row to a user and handing the CSVs to reviewers were not exercised in our test run.

Done when every scenario has a CSV whose row count equals the number of results in results/FIELD.json, the user approved the sample row and the mapping, and state.json marks the batch done.

Compliance

  1. Use road test footage the team owns or has the rights to analyze. For public driving videos used as extra material, check the license and the platform's terms first and keep them for internal analysis only. Have the user confirm this before Step 3.
  2. People on the street did not agree to be filmed. The skill lists moments by what happens on the road; it does not identify anyone, does no face, license plate or identity recognition, and never searches for people by appearance or age. Street footage can include children passing by; do not use this skill to collect footage of particular people or of children. Follow the team's privacy rules (some teams blur faces before any upload), and confirm before Step 3.
  3. The skill does not judge whether the camera car drove correctly or whether anyone broke a traffic rule, and it produces no annotation files. The CSVs are review lists for people.
  4. The videos stay in the user's Vivu project until the user deletes them. Create the project as private; the default is visible to the whole Vivu organization. Delete videos or the project only when the user asks, and confirm first.
  5. The skill writes only local files. Any upload or post of the CSVs goes through the approval in Step 6.

Known failure modes

Symptom Cause Fix
(observed) most cut in candidates are not cut ins: 10 returned, 2 real and 8 false, and 1 of the 3 cut ins in the hand marks was missed (most of those hand marks were added after earlier searches found them) the search finds a vehicle ahead and the reason adds the lane change label every candidate from its sheet: a cut in needs the vehicle in the next lane in early tiles and in the camera car's lane in later tiles
(observed) a cut in reason says "where there was no vehicle a moment before" about a car that was ahead all along the reason reuses the query's own words ignore the reason; label from the frames
(observed) a cut in candidate is the camera car's own lane change, a parked car or cross traffic picture search does not tell who moved compare the first and last tiles; if the lane markings slide sideways under the camera car, it was the camera car that moved
(observed) pedestrian candidates with nobody in the camera car's lane: 18 returned, 11 real and 7 false, and 6 crossings in the hand marks were missed people on the sidewalk, empty zebras and side road crossings match the query label each candidate from its sheet; never count a crossing from the reason
(observed) a reason has a person "running across the crosswalk" who stands on the sidewalk in every frame the reason describes what the query asked for label from the frames
(observed) people cross other lanes of a junction beside the camera car's lane, never in it the default rule counts only the camera car's lane; the reviewers may count any lane agree on the definition before labeling and write it into config.json
(observed) a far pedestrian or cyclist cannot be seen on the contact sheet the sheet scales each frame down extract a zoomed full size frame at that second (Step 5)
(observed) a search for turns across an occupied crosswalk returns straight drives: 8 returned, 3 real and 5 false the reason asserts a turn that the frames do not show the skill does not search for turns; reviewers note them in the pedestrian list
(observed) a cyclist candidate is a bicycle parked at a rack or a rider on the sidewalk, while the reason describes someone riding on the road a bicycle anywhere in the picture matches, and the reason follows the query label it looks wrong; zoom in on any bicycle near the curb to see whether anyone is on it (the first labeling pass in our test run once took a parked bicycle for a rider)
(observed) a window covers almost the whole clip: 44.5 seconds for a cut in that lasted a few seconds windows are wider than events look at every sheet of the window, not only the first
(observed) "No such filter: 'drawtext'" when stamping times on a sheet this ffmpeg build has no drawtext filter use the Step 5 command; the tile position gives the time
upload page asks to sign in or shows an error the one time upload link expires after 180 seconds request a new link right before opening it
a file is rejected by the browser upload tool the tool accepts at most 10 MB per call the user adds the file in their own browser or the Vivu web app
"has not granted vivu.write" Vivu connected read only the user reconnects Vivu with write access
a result's file name does not match BATCH_DIR Vivu replaced a space, bracket or plus sign in the name match on listed_name in files.csv
a search returns exactly its maximum_results the batch holds more candidates than one search returns getting past the cap means uploading the whole batch again as two projects, so quote the cost and wait for approval; if the user declines, mark the list cut off and say so in the summary
a scenario comes back empty none in the batch, or the search missed them an empty result does not prove the footage has no such moment; say so in the summary
labels on night, rain or backlit footage look unreliable this footage is untested treat every row as unmeasured and have the reviewers check each one

Connect Vivu, then add the skill.

The skill runs through the Vivu connector. Add it to Claude from the connector directory, then install the skill.