---
name: robot-episode-label-spot-check
description: "Check a robot teleop batch's task labels against the camera video with Vivu and get a review CSV: matched or needs a human look, with frames. Use before a batch goes into training."
---Robot episode label spot-check with Vivu
This skill takes one collection batch of robot manipulation episodes (one video file per episode, exported from your logs) and the task label each episode got at collection time. Claude turns each label into a description of what the camera shows when the task is done, indexes the batch in Vivu, runs one precise search per label, checks every returned episode on its own frames, and writes a CSV with one row per episode: the label, matched or needs a human look, the match window, a last frame and a contact sheet, and an empty column for the reviewer's verdict.
The value is in what only the footage shows. A label is text an operator picked from a list; whether the cup actually ended up inside the other cup, whether the run was cut short, or whether the arm blocked the camera at the end is only visible in the recording. Success flags and gripper states in the log do not say where the object ended up. An episode that Vivu does not return under its own label goes to a human, never to "failed", because an empty search result does not prove the task was not done.
When to use
Use when someone asks to spot-check the task labels on a teleop batch, find episodes whose video does not match their label, review a LeRobot or RLDS batch before training, or build a review list for today's collection. For one episode, open the video. This skill does not produce annotations or training labels, grade operators, or explain why an episode failed; for those, use your annotation tool and your own review.
Working principles
- Report real numbers, not estimates. When a number is an estimate, say so.
- Nothing is verified until it has been checked against the episode's frames. A Vivu match is a candidate until Claude has looked at the start, middle and end of the window.
- Stop and tell the user when a required capability or tool is missing. Do not guess around it.
- Ask the user before anything that costs index minutes or search credits (Step 3) and before anything leaves the working folder.
- A label counts as matched only when the search for that same label returned the episode and the frames show the end state. A match under a different label is an UNCERTAIN hint for the reviewer, never a relabel.
- The CSV is a review list for the data team. It carries no operator scores, no failure reasons and no accuracy figures to share outside the team.
What you need before starting
Check each item at the start of the run and tell the user plainly what is missing before doing anything else.
| Requirement | Why | How to check |
|---|---|---|
| Vivu connector with write access | create a private project, open its upload page, search | vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the connection is read only: ask the user to reconnect Vivu with write access |
| A shell on the machine that holds the episodes, with ffmpeg and ffprobe | minutes for pricing, frames for every check, last frames and contact sheets for the CSV | ffmpeg -version and ffprobe -version |
| One video file per episode (mp4 or another common video format) | Vivu indexes video files; ROS bag, MCAP or RLDS logs have to be exported first | ls the batch folder and ffprobe one file |
| The label list for the batch (LeRobot meta/episodes.jsonl, a CSV or a JSON export from the collection tool) | every row is compared with its label | read it and count episodes per label |
| A browser the user can open the upload page in, or a browser tool that can attach local files | moving files into Vivu | ask the user; for Claude in Chrome, check that it is connected |
The skill runs in Claude Code on the user's computer (terminal or the Code tab of Claude Desktop), because it needs a shell and the local episode files. It needs no residential IP, since nothing is downloaded from YouTube, and it sets up no scheduled task: run it once per batch.
Inputs to collect
Ask for anything missing, most important first.
- The batch folder and the camera to check (required). Default camera: the third person view; a wrist camera rarely shows where the object ended up.
- The label list and which field holds the label (required).
- Scope. Default: every episode in the batch. When the batch is larger than the plan allows, Step 3 offers a per label sample.
- The working folder. Default: label-check/ next to the batch folder.
- The Vivu project name. Default: "Label check BATCH_ID", created private.
Files and state
Keep everything in one working folder:
label-check/
config.json batch id, batch folder, camera, label file, Vivu project id, a LABEL_KEY and one description per label
episodes.csv episode_file, path, label, duration_s, stored_name, video_id
state.json uploaded video ids, finished label searches with job ids, triaged episodes
h264/ H.264 copies of AV1 episodes (Step 2), same base names; nothing is written into the batch folder
results/LABEL_KEY.json raw result of each label search, saved as returned
frames/ start, middle and end frames of every returned episode
review/ last frame and contact sheet per episode
label_spot_check.csv the review list
LABEL_KEY is a short lowercase key for each distinct label, stored in config.json (for example cloth or ranch); results/LABEL_KEY.json and the frame names in Step 6 use it.
state.json lets a rerun resume: skip uploads whose video id is already there, skip label searches that already have a saved result, and skip episodes already triaged. Each batch gets its own working folder, so a row is never reported twice.
Step 1: Check the Vivu connector and the tools
Goal: confirm that everything in the table above is in place.
- Call vivu_get_account. If the tool does not exist, tell the user to add the Vivu connector in Claude (https://mcp.vivu.ai/mcp) and stop. If can_create_projects is not true, or a later write fails with "has not granted vivu.write", ask the user to reconnect Vivu with write access and stop.
- Run ffmpeg -version and ffprobe -version. If either is missing, ask the user to install ffmpeg.
- Confirm the batch folder and the label list can be read.
Done when vivu_get_account shows can_create_projects: true, both ffmpeg commands print a version, and the batch folder and label list are readable.
Step 2: Line up episodes and labels
Goal: episodes.csv with one row per episode file, its label and its length.
- If the batch is still in ROS bag, MCAP or RLDS form, export the chosen camera to one mp4 per episode with the team's own export script. Log export was not exercised in our test run: the public dataset we used already ships a video file per episode.
- LeRobot datasets encode video as AV1. We did not test uploading AV1 to Vivu; in our test run we re-encoded each file to H.264 first. This is a format conversion, not a way to fit an upload limit:
ffmpeg -v error -i EPISODE_FILE -c:v libx264 -pix_fmt yuv420p -crf 20 -an OUTPUT_FILE
EPISODE_FILE is the exported episode. OUTPUT_FILE is label-check/h264/ plus the episode's base name. When a copy exists, Step 4 uploads the copy and Steps 6 and 7 extract frames from it.
- Measure each file:
ffprobe -v error -show_entries format=duration -of csv=p=0 EPISODE_FILE
- Join files to labels by episode index, not by position in a list. Show the user the count per label and any file without a label or label without a file.
Done when episodes.csv has one row per episode with label and duration_s, and the user has seen the counts per label.
Step 3: Price the batch and ask
Goal: the user approves the cost before anything is uploaded.
- Index minutes = the sum of duration_s divided by 60. Search credits = one precise search per distinct label; a precise search uses 5 credits. Label it an estimate, since a rewording costs another search.
- Call vivu_get_usage for the plan and what remains this month. For reference, Free is 20 indexing minutes a month and 50 search credits a month; Premium is $30 a month for 180 indexing minutes a month and 500 search credits a month.
- Show one table:
| This batch | Remaining this month | Fits | |
|---|---|---|---|
| Index minutes | sum of durations | from vivu_get_usage | yes or no |
| Search credits (estimate) | labels × 5 | from vivu_get_usage | yes or no |
- If it does not fit, offer these levers in order, each with its saving: check a sample per label (for example 10 episodes per label) instead of the whole batch; leave out episodes already reviewed in an earlier batch; check one camera only; move to a larger plan last. For example, 50 episodes of about 30 seconds is about 25 index minutes, more than a Free month.
Approval was not exercised in our test run (it ran unattended on an admin account that reports no remaining allowance).
Done when the user has seen the table and said yes to a scope.
Step 4: Create a private project and upload
Goal: every episode in the approved scope shows ready in Vivu.
- Ask the user to confirm the Compliance points before uploading.
- Call vivu_list_projects and reuse a project for this batch if one exists. Otherwise call vivu_create_project with the project name and visibility "private". The default visibility is the whole organization, which is wrong for internal lab footage.
- Call vivu_open_upload_page with the project ID right before the upload. The upload_url is a one time sign in link that expires in 180 seconds, so request it only when the user or the browser tool is ready, never post it in a message, and request a new one if it lapses.
- Hand the link to the user to open in their own browser and select the episode files, or open it with a browser tool that can attach local files. Claude in Chrome takes at most 10 MB per upload call, so send files in batches under that. If a file cannot go through, ask the user to add it in the Vivu web app. Never split or recompress an episode to fit.
- Poll vivu_list_videos until every video shows ready. Vivu replaces spaces and brackets in file names with underscores, so match rows back by the episode index in the name and check that duration_ms agrees with duration_s. Record the name Vivu shows (stored_name) and video_id in episodes.csv and the video ids in state.json.
In our test run the upload page link was opened by an automated browser on our machine, not by a user.
Done when vivu_list_videos shows every approved episode as ready and each one is matched to a row in episodes.csv.
Step 5: Turn each label into a picture description
Goal: one search description per distinct label, approved by the user.
A label is an instruction ("put the ranch bottle into the pot"). A search needs what the camera shows when it is done. Rewrite each label with this template:
the robot arm ACTION, so OBJECT ends up WHERE
- Keep the label's objects, colors and direction. The end state clause is what keeps an unfinished episode from matching.
- Name objects the way they look on camera (a red bowl that looks pink is still "the red bowl" if the label says so; add the color the camera shows only if the user agrees).
- Show the user the list, one sample CSV row and the field mapping below, and wait for a yes. Fixing a description now costs nothing; after the search it costs another 5 credit search.
| CSV field | Source | If unavailable |
|---|---|---|
| matched | yes, when its own label search returned it and the frames show the end state | needs a human look |
| match_window | start_ms and end_ms of that result, as MM:SS | blank |
| last_frame, window_sheet | ffmpeg on the local file (Step 7) | blank, and say why |
| other_label_hint | another label's search returned it and its frames were judged real | blank |
| review_verdict | the reviewer | blank |
Sample row:
episode_file,label,matched,match_window,last_frame,window_sheet,other_label_hint,review_verdict
episode_000008.mp4,put the ranch bottle into the pot,needs a human look,,review/episode_000008_last.png,review/episode_000008_window.png,UNCERTAIN: sweep the green cloth to the left side of the table,
User confirmation was not exercised in our test run.
Done when config.json holds one description per label and the user has approved them and the sample row.
Step 6: Search each label and check every match
Goal: for every label, the episodes whose frames show that task done.
The standing queries are the label descriptions, one per label:
| Field | Query (template) | Mode |
|---|---|---|
| matched, for label L | the robot arm ACTION, so OBJECT ends up WHERE, written from L | precise |
The chain is short on purpose. The label is what was said; the search looks for what was done on camera. One precise search per label finds the episodes that show it; the frames then confirm or reject each one. A fast pass adds nothing: fast returns whole files with no reason, and every episode here is a short file that gets checked anyway.
- For each label, call vivu_search_videos with project_id, query, mode "precise" and maximum_results equal to the number of episodes in the project (at most 100). The value is the recall ceiling: a smaller number silently drops matches. If the project holds more than 100 episodes, split the batch into projects of 100 or fewer.
- Call vivu_get_search_results with the job ID until complete is true. Each call can wait up to 45 seconds, so a pending search is not a stalled one. Save the returned JSON unchanged as results/LABEL_KEY.json and record the job ID in state.json. The Vivu result page link from a completed search can go in the live reply; it expires after four hours, so it never goes in the CSV.
- For every returned episode, extract three frames inside the window (start, middle, end; use end minus 0.3 seconds when the window ends at the end of the file) and look at each one:
ffmpeg -v error -ss SECONDS -i EPISODE_FILE -frames:v 1 -q:v 3 label-check/frames/NAME_POSITION.png
SECONDS is the frame time, NAME the episode file's base name plus LABEL_KEY, POSITION start, mid or end. Judge real only if the end frame shows the label's end state. The reason text is a paraphrase and has invented objects, so judge from the frames. 4. Returned and judged real, same label: matched is yes. Returned and judged false: count it as a false positive, and the row stays needs a human look. Returned under another label and judged real: add that label to other_label_hint, marked UNCERTAIN. Not returned under its own label: needs a human look. 5. Tell the user the counts per label: returned, real, false.
How far to trust the search: finding "the task was done" is action and visual state search, which is new and lightly tested for Vivu, and lookalike objects are a known weakness. That is why no row is marked matched without the frame check, and why a row Vivu did not return goes to a human instead of being called a failure: an empty result does not prove the footage has no such moment.
Worked example from our test run
In our test run on 12 public robot episodes (3.73 minutes; a UR5 arm doing tabletop pick and place and cloth sweeping, one third person camera that is not on the arm and shifts during some episodes, a low frame rate, no audio), all 12 were ready within 76 seconds of the upload starting. We changed the batch labels on purpose before searching: we swapped labels between tasks, kept one label that the public dataset itself gets wrong (episode_000499 is labeled "put the marker into the bowl" but shows the cup task), and cut episode_000038 short so the bottle never reaches the pot. We wrote the truth from contact sheets before any search. The descriptions we ran, each with maximum_results set to the batch size:
| Field | Query | Mode |
|---|---|---|
| matched, cloth label | the robot arm sweeps the green cloth across the table so the cloth ends up on the left side of the table | precise |
| matched, ranch label | the robot arm puts the white ranch dressing bottle into the metal pot, so the bottle ends up inside the pot | precise |
| matched, cup label | the robot arm picks up the blue cup and puts it into the brown cup, so the blue cup ends up inside the brown cup | precise |
| matched, tiger label | the robot arm takes the stuffed tiger out of the red bowl and puts it in the grey bowl, so the tiger ends up in the grey bowl | precise |
| matched, marker label | the robot arm puts a marker pen into a bowl, so the marker ends up inside the bowl | precise |
| control, not part of the skill | the reverse of the tiger task and of the cup task | precise |
The label searches whose task was in the batch returned 11 matches: 11 real, 0 false, 0 missed against that truth. Each episode with a changed or wrong label, and the cut short one, came out as needs a human look, and each correctly labeled episode came out as matched. Median windows ran from 14.8 to 21.6 seconds, most of each episode, so a window says little about when the task finished; the last frame does.
The label with no footage in the batch, "put the marker into the bowl", returned 2 matches, both false: the reason text called the ranch bottle a marker pen. A description that gave the marker's shape and size, after 1 rewording on the same corpus, again returned 2, both false, one of them a tiger episode with an invented marker. Neither changed a row, because a match only counts under the episode's own label, but it is why hints need frames too. We also searched the reverse of the tiger task and of the cup task: the reversed tiger task returned 0, and the reversed cup task returned 1 false (a tiger episode described as cups).
In this dataset each task has its own objects on the table, so most wrong labels could be caught from which objects are present. Batches where several tasks share one scene and differ only in the action were not tested; expect weaker results there and review more rows by hand. The run used 8 precise searches, an estimated 40 credits.
Done when every label has a saved result, every returned episode has a real or false judgment in state.json, and the user has seen the counts.
Step 7: Build the review CSV
Goal: label_spot_check.csv, one row per episode, each with a last frame and a contact sheet.
- Last frame of every episode:
ffmpeg -v error -y -sseof -0.3 -i EPISODE_FILE -frames:v 1 label-check/review/NAME_last.png
NAME is the episode file's base name.
- Contact sheet at one frame per second: the match window for matched rows, the whole episode for the rest. START and LENGTH are in seconds; for episodes longer than 25 seconds, replace fps=1 with fps=25/LENGTH so the sheet still fits 25 tiles:
ffmpeg -v error -y -ss START -t LENGTH -i EPISODE_FILE -vf "fps=1,scale=320:-1,tile=5x5" -frames:v 1 label-check/review/NAME_window.png
- For needs a human look rows, vivu_get_video_summary (project_id, video_id, include_segments true) gives a short description of what the episode shows; it can go in the reply as a second signal. In our test run it described the cut short episode as the arm hovering over the pot (it called it a metal bowl) and then grasping the bottle, without saying the bottle went in, and episode_000499 as cup stacking. It is not a verdict.
- Write label_spot_check.csv with the columns from Step 5, file names as the only link to the video (no Vivu links). Show the user the matched and needs a human look counts and the first rows.
- The CSV stays in the working folder. If the user wants it in a shared drive or a tracker, name the destination, show the rendered file and wait for approval; that was not exercised in our test run.
Done when label_spot_check.csv has one row per episode in episodes.csv, every last_frame and window_sheet path exists, and the user has the file.
Compliance
- Rights to the footage: use it on data your team owns or may process. For public datasets, follow the dataset license and keep the videos for internal checks. Confirm before Step 4.
- People on camera: lab recordings often catch operators or passers-by. Ask whether they know the recordings are processed; leave out episodes that show people who did not agree. Confirm before Step 4.
- Minors: if a home collection might show children, get explicit guardian consent or leave those episodes out. Confirm before Step 4.
- Where the videos go: uploaded episodes stay in the user's Vivu project until the user deletes them. The skill creates the project as private and never deletes projects or videos unless the user asks, and then confirms first.
- What the skill does not do: no face or identity recognition, no operator scoring, no failure diagnosis, no annotation export.
Known failure modes
| Symptom | Cause | Fix |
|---|---|---|
| (observed) a label with no footage in the batch still gets matches, with reason text such as "the white marker pen with a green cap" | Vivu matched a lookalike object (the ranch bottle) and described it as the labeled object | count a match only under the episode's own label and only after the frames show the object; treat other labels' matches as hints |
| (observed) a reworded description returns an episode with an object that is not there, "a black, slim pen-shaped marker" | the reason text invents objects | judge from the start, middle and end frames, never from the reason |
| (observed) a reversed task description returns an unrelated episode, "the blue cup inside the brown tiger-striped cup" | the search can return a confident match for an action nobody did | keep the end state clause in every description and check the end frame |
| (observed) match windows cover most of each episode | precise windows contain the moment but are wide | read the end state from the last frame and the contact sheet, not from the window bounds |
| (observed) a public dataset label names a task the video does not show | the collection label was wrong at the source | this is what needs a human look is for; do not trust labels over frames |
| an episode that did the task is not returned | recall is below perfect for actions and visual states, and maximum_results caps it | leave it as needs a human look; raise maximum_results to the episode count |
| every episode of one label comes back as needs a human look | the description does not match what the camera shows (wrong color name, wrong camera) | compare one last frame with the description, fix the wording with the user, run that label again |
| tasks that share one scene and differ only in the action are confused | untested; objects alone cannot separate them | review more rows by hand and add the end state to each description |
| "has not granted vivu.write" | Vivu connected read only | reconnect Vivu with write access |
| upload page asks to sign in or shows an error | the one time upload link expires after 180 seconds | request a new link right before opening it |
| a file is rejected by the browser upload tool | file above the tool's size limit | the user adds that file in the Vivu web app |
| rows do not line up with results | Vivu normalized the file name | match on the episode index and duration_ms |
| the camera is blocked by the arm in the last frame | the arm parks over the scene | look at the contact sheet, or check a second camera for that row |