SkillsSkill for Claude

Robot episode label spot-check

Give Claude one collection batch of robot manipulation episodes and the task label each one got at collection time. Claude rewrites each label as what the camera shows when the task is done, indexes the batch in Vivu, runs one precise search per label, checks every returned episode on its own frames, and writes a CSV with one row per episode: matched or needs a human look, the match window, a last frame and a contact sheet, and an empty column for your verdict. Whether the object really ended up in the container, or the run was cut short, is only visible in the video; an episode Vivu does not return goes to a human, never to failed.

Maintained by Vivu. Updated 2026-09-28.

Download

robot-episode-label-spot-check.zip

9 KB. Unzips to robot-episode-label-spot-check/SKILL.md. Upload the zip as it is in the Claude app, or unzip it into your skills folder for Claude Code.

SHA-256 ec6da7fbe0a094168b3b44eda144a5ff8188ab104f770b242a078c6cd652c081

At a glance

What the Robot episode label spot-check skill does, where it runs, what it needs, and when it asks
Looks forThe point where an episode shows its labeled task finished, such as the cup sitting inside the other cup or the cloth on the left side of the table. Labels and log flags cannot show where the object ended up; the camera can. Every match is checked against the episode's own frames before it goes in the CSV.
Runs onClaude Code on your computer (terminal or the Code tab of Claude Desktop), because it needs a shell with ffmpeg and ffprobe and the episode files on local disk. No residential IP is needed and nothing is scheduled: you run it once per batch.
Needs
  • The Vivu connector with write access, to create a private project, open its upload page and search.
  • A shell with ffmpeg and ffprobe on the machine that holds the episodes, for durations, frame checks, last frames and contact sheets.
  • One video file per episode; ROS bag, MCAP or RLDS logs have to be exported to video first.
  • The batch's label list (LeRobot episodes.jsonl, a CSV or a JSON export), so each row can be compared with its label.
  • A browser you can open the upload page in, or a browser tool that can attach local files.
Your Vivu planIndex minutes equal the total length of the episode videos you check, and each label uses one precise search; a precise search uses 5 credits. The skill prices the batch against your plan with vivu_get_usage and waits for your yes before uploading, offering a per label sample when the batch is too large.
Asks you firstBefore uploading (the cost against your plan and the compliance points), before searching (the picture description for each label and a sample CSV row), and before the CSV goes anywhere outside the working folder.

The skill does its video work through the Vivu connector. If the connector is not in your Claude yet, add it first; the skill checks that it is connected before it does anything else.

Before you run it

  • Recognizing a finished action is lightly tested in Vivu, so every match is checked on frames and episodes Vivu does not return go to a human instead of being marked failed.
  • A label whose object is absent from the batch can still get lookalike matches with invented descriptions; matches under another label are only UNCERTAIN hints.
  • Batches where several tasks share one scene and differ only in the action were not tested; plan to review more rows by hand.
  • Lab recordings can show operators or passers-by: confirm they know, and leave out children unless a guardian agreed.
  • Uploaded episodes stay in your Vivu project until you delete them; the skill creates the project as private.
  • The upload link expires in 180 seconds, and Claude in Chrome uploads at most 10 MB per call, so larger files go through the Vivu web app.
  • The CSV is a review list, not annotations: it carries no operator scores, failure reasons or accuracy figures.
  • The skill writes only to your working folder unless you approve another destination.

Start it

Once the skill is installed, ask for the task in your own words. Naming the skill is the most reliable way to have Claude use it. For example:

Spot-check the labels on today's teleop batch in ~/data/batch_0924: the labels are in meta/episodes.jsonl, use the front camera, and give me the review CSV.

In Claude Code you can also type /robot-episode-label-spot-check. Claude asks for anything the request leaves out, most important first.

What is inside

  1. When to use
  2. Working principles
  3. What you need before starting
  4. Inputs to collect
  5. Files and state
  6. Step 1: Check the Vivu connector and the tools
  7. Step 2: Line up episodes and labels
  8. Step 3: Price the batch and ask
  9. Step 4: Create a private project and upload
  10. Step 5: Turn each label into a picture description
  11. Step 6: Search each label and check every match
  12. Step 7: Build the review CSV
  13. Compliance
  14. Known failure modes

The full skill

This is robot-episode-label-spot-check/SKILL.md from the download, as Claude reads it: the frontmatter first, then the instructions.

---
name: robot-episode-label-spot-check
description: "Check a robot teleop batch's task labels against the camera video with Vivu and get a review CSV: matched or needs a human look, with frames. Use before a batch goes into training."
---

Robot episode label spot-check with Vivu

This skill takes one collection batch of robot manipulation episodes (one video file per episode, exported from your logs) and the task label each episode got at collection time. Claude turns each label into a description of what the camera shows when the task is done, indexes the batch in Vivu, runs one precise search per label, checks every returned episode on its own frames, and writes a CSV with one row per episode: the label, matched or needs a human look, the match window, a last frame and a contact sheet, and an empty column for the reviewer's verdict.

The value is in what only the footage shows. A label is text an operator picked from a list; whether the cup actually ended up inside the other cup, whether the run was cut short, or whether the arm blocked the camera at the end is only visible in the recording. Success flags and gripper states in the log do not say where the object ended up. An episode that Vivu does not return under its own label goes to a human, never to "failed", because an empty search result does not prove the task was not done.

When to use

Use when someone asks to spot-check the task labels on a teleop batch, find episodes whose video does not match their label, review a LeRobot or RLDS batch before training, or build a review list for today's collection. For one episode, open the video. This skill does not produce annotations or training labels, grade operators, or explain why an episode failed; for those, use your annotation tool and your own review.

Working principles

  1. Report real numbers, not estimates. When a number is an estimate, say so.
  2. Nothing is verified until it has been checked against the episode's frames. A Vivu match is a candidate until Claude has looked at the start, middle and end of the window.
  3. Stop and tell the user when a required capability or tool is missing. Do not guess around it.
  4. Ask the user before anything that costs index minutes or search credits (Step 3) and before anything leaves the working folder.
  5. A label counts as matched only when the search for that same label returned the episode and the frames show the end state. A match under a different label is an UNCERTAIN hint for the reviewer, never a relabel.
  6. The CSV is a review list for the data team. It carries no operator scores, no failure reasons and no accuracy figures to share outside the team.

What you need before starting

Check each item at the start of the run and tell the user plainly what is missing before doing anything else.

Requirement Why How to check
Vivu connector with write access create a private project, open its upload page, search vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the connection is read only: ask the user to reconnect Vivu with write access
A shell on the machine that holds the episodes, with ffmpeg and ffprobe minutes for pricing, frames for every check, last frames and contact sheets for the CSV ffmpeg -version and ffprobe -version
One video file per episode (mp4 or another common video format) Vivu indexes video files; ROS bag, MCAP or RLDS logs have to be exported first ls the batch folder and ffprobe one file
The label list for the batch (LeRobot meta/episodes.jsonl, a CSV or a JSON export from the collection tool) every row is compared with its label read it and count episodes per label
A browser the user can open the upload page in, or a browser tool that can attach local files moving files into Vivu ask the user; for Claude in Chrome, check that it is connected

The skill runs in Claude Code on the user's computer (terminal or the Code tab of Claude Desktop), because it needs a shell and the local episode files. It needs no residential IP, since nothing is downloaded from YouTube, and it sets up no scheduled task: run it once per batch.

Inputs to collect

Ask for anything missing, most important first.

  1. The batch folder and the camera to check (required). Default camera: the third person view; a wrist camera rarely shows where the object ended up.
  2. The label list and which field holds the label (required).
  3. Scope. Default: every episode in the batch. When the batch is larger than the plan allows, Step 3 offers a per label sample.
  4. The working folder. Default: label-check/ next to the batch folder.
  5. The Vivu project name. Default: "Label check BATCH_ID", created private.

Files and state

Keep everything in one working folder:

label-check/
  config.json            batch id, batch folder, camera, label file, Vivu project id, a LABEL_KEY and one description per label
  episodes.csv           episode_file, path, label, duration_s, stored_name, video_id
  state.json             uploaded video ids, finished label searches with job ids, triaged episodes
  h264/                  H.264 copies of AV1 episodes (Step 2), same base names; nothing is written into the batch folder
  results/LABEL_KEY.json raw result of each label search, saved as returned
  frames/                start, middle and end frames of every returned episode
  review/                last frame and contact sheet per episode
  label_spot_check.csv   the review list

LABEL_KEY is a short lowercase key for each distinct label, stored in config.json (for example cloth or ranch); results/LABEL_KEY.json and the frame names in Step 6 use it.

state.json lets a rerun resume: skip uploads whose video id is already there, skip label searches that already have a saved result, and skip episodes already triaged. Each batch gets its own working folder, so a row is never reported twice.

Step 1: Check the Vivu connector and the tools

Goal: confirm that everything in the table above is in place.

  1. Call vivu_get_account. If the tool does not exist, tell the user to add the Vivu connector in Claude (https://mcp.vivu.ai/mcp) and stop. If can_create_projects is not true, or a later write fails with "has not granted vivu.write", ask the user to reconnect Vivu with write access and stop.
  2. Run ffmpeg -version and ffprobe -version. If either is missing, ask the user to install ffmpeg.
  3. Confirm the batch folder and the label list can be read.

Done when vivu_get_account shows can_create_projects: true, both ffmpeg commands print a version, and the batch folder and label list are readable.

Step 2: Line up episodes and labels

Goal: episodes.csv with one row per episode file, its label and its length.

  1. If the batch is still in ROS bag, MCAP or RLDS form, export the chosen camera to one mp4 per episode with the team's own export script. Log export was not exercised in our test run: the public dataset we used already ships a video file per episode.
  2. LeRobot datasets encode video as AV1. We did not test uploading AV1 to Vivu; in our test run we re-encoded each file to H.264 first. This is a format conversion, not a way to fit an upload limit:
ffmpeg -v error -i EPISODE_FILE -c:v libx264 -pix_fmt yuv420p -crf 20 -an OUTPUT_FILE

EPISODE_FILE is the exported episode. OUTPUT_FILE is label-check/h264/ plus the episode's base name. When a copy exists, Step 4 uploads the copy and Steps 6 and 7 extract frames from it.

  1. Measure each file:
ffprobe -v error -show_entries format=duration -of csv=p=0 EPISODE_FILE
  1. Join files to labels by episode index, not by position in a list. Show the user the count per label and any file without a label or label without a file.

Done when episodes.csv has one row per episode with label and duration_s, and the user has seen the counts per label.

Step 3: Price the batch and ask

Goal: the user approves the cost before anything is uploaded.

  1. Index minutes = the sum of duration_s divided by 60. Search credits = one precise search per distinct label; a precise search uses 5 credits. Label it an estimate, since a rewording costs another search.
  2. Call vivu_get_usage for the plan and what remains this month. For reference, Free is 20 indexing minutes a month and 50 search credits a month; Premium is $30 a month for 180 indexing minutes a month and 500 search credits a month.
  3. Show one table:
This batch Remaining this month Fits
Index minutes sum of durations from vivu_get_usage yes or no
Search credits (estimate) labels × 5 from vivu_get_usage yes or no
  1. If it does not fit, offer these levers in order, each with its saving: check a sample per label (for example 10 episodes per label) instead of the whole batch; leave out episodes already reviewed in an earlier batch; check one camera only; move to a larger plan last. For example, 50 episodes of about 30 seconds is about 25 index minutes, more than a Free month.

Approval was not exercised in our test run (it ran unattended on an admin account that reports no remaining allowance).

Done when the user has seen the table and said yes to a scope.

Step 4: Create a private project and upload

Goal: every episode in the approved scope shows ready in Vivu.

  1. Ask the user to confirm the Compliance points before uploading.
  2. Call vivu_list_projects and reuse a project for this batch if one exists. Otherwise call vivu_create_project with the project name and visibility "private". The default visibility is the whole organization, which is wrong for internal lab footage.
  3. Call vivu_open_upload_page with the project ID right before the upload. The upload_url is a one time sign in link that expires in 180 seconds, so request it only when the user or the browser tool is ready, never post it in a message, and request a new one if it lapses.
  4. Hand the link to the user to open in their own browser and select the episode files, or open it with a browser tool that can attach local files. Claude in Chrome takes at most 10 MB per upload call, so send files in batches under that. If a file cannot go through, ask the user to add it in the Vivu web app. Never split or recompress an episode to fit.
  5. Poll vivu_list_videos until every video shows ready. Vivu replaces spaces and brackets in file names with underscores, so match rows back by the episode index in the name and check that duration_ms agrees with duration_s. Record the name Vivu shows (stored_name) and video_id in episodes.csv and the video ids in state.json.

In our test run the upload page link was opened by an automated browser on our machine, not by a user.

Done when vivu_list_videos shows every approved episode as ready and each one is matched to a row in episodes.csv.

Step 5: Turn each label into a picture description

Goal: one search description per distinct label, approved by the user.

A label is an instruction ("put the ranch bottle into the pot"). A search needs what the camera shows when it is done. Rewrite each label with this template:

the robot arm ACTION, so OBJECT ends up WHERE
  1. Keep the label's objects, colors and direction. The end state clause is what keeps an unfinished episode from matching.
  2. Name objects the way they look on camera (a red bowl that looks pink is still "the red bowl" if the label says so; add the color the camera shows only if the user agrees).
  3. Show the user the list, one sample CSV row and the field mapping below, and wait for a yes. Fixing a description now costs nothing; after the search it costs another 5 credit search.
CSV field Source If unavailable
matched yes, when its own label search returned it and the frames show the end state needs a human look
match_window start_ms and end_ms of that result, as MM:SS blank
last_frame, window_sheet ffmpeg on the local file (Step 7) blank, and say why
other_label_hint another label's search returned it and its frames were judged real blank
review_verdict the reviewer blank

Sample row:

episode_file,label,matched,match_window,last_frame,window_sheet,other_label_hint,review_verdict
episode_000008.mp4,put the ranch bottle into the pot,needs a human look,,review/episode_000008_last.png,review/episode_000008_window.png,UNCERTAIN: sweep the green cloth to the left side of the table,

User confirmation was not exercised in our test run.

Done when config.json holds one description per label and the user has approved them and the sample row.

Step 6: Search each label and check every match

Goal: for every label, the episodes whose frames show that task done.

The standing queries are the label descriptions, one per label:

Field Query (template) Mode
matched, for label L the robot arm ACTION, so OBJECT ends up WHERE, written from L precise

The chain is short on purpose. The label is what was said; the search looks for what was done on camera. One precise search per label finds the episodes that show it; the frames then confirm or reject each one. A fast pass adds nothing: fast returns whole files with no reason, and every episode here is a short file that gets checked anyway.

  1. For each label, call vivu_search_videos with project_id, query, mode "precise" and maximum_results equal to the number of episodes in the project (at most 100). The value is the recall ceiling: a smaller number silently drops matches. If the project holds more than 100 episodes, split the batch into projects of 100 or fewer.
  2. Call vivu_get_search_results with the job ID until complete is true. Each call can wait up to 45 seconds, so a pending search is not a stalled one. Save the returned JSON unchanged as results/LABEL_KEY.json and record the job ID in state.json. The Vivu result page link from a completed search can go in the live reply; it expires after four hours, so it never goes in the CSV.
  3. For every returned episode, extract three frames inside the window (start, middle, end; use end minus 0.3 seconds when the window ends at the end of the file) and look at each one:
ffmpeg -v error -ss SECONDS -i EPISODE_FILE -frames:v 1 -q:v 3 label-check/frames/NAME_POSITION.png

SECONDS is the frame time, NAME the episode file's base name plus LABEL_KEY, POSITION start, mid or end. Judge real only if the end frame shows the label's end state. The reason text is a paraphrase and has invented objects, so judge from the frames. 4. Returned and judged real, same label: matched is yes. Returned and judged false: count it as a false positive, and the row stays needs a human look. Returned under another label and judged real: add that label to other_label_hint, marked UNCERTAIN. Not returned under its own label: needs a human look. 5. Tell the user the counts per label: returned, real, false.

How far to trust the search: finding "the task was done" is action and visual state search, which is new and lightly tested for Vivu, and lookalike objects are a known weakness. That is why no row is marked matched without the frame check, and why a row Vivu did not return goes to a human instead of being called a failure: an empty result does not prove the footage has no such moment.

Worked example from our test run

In our test run on 12 public robot episodes (3.73 minutes; a UR5 arm doing tabletop pick and place and cloth sweeping, one third person camera that is not on the arm and shifts during some episodes, a low frame rate, no audio), all 12 were ready within 76 seconds of the upload starting. We changed the batch labels on purpose before searching: we swapped labels between tasks, kept one label that the public dataset itself gets wrong (episode_000499 is labeled "put the marker into the bowl" but shows the cup task), and cut episode_000038 short so the bottle never reaches the pot. We wrote the truth from contact sheets before any search. The descriptions we ran, each with maximum_results set to the batch size:

Field Query Mode
matched, cloth label the robot arm sweeps the green cloth across the table so the cloth ends up on the left side of the table precise
matched, ranch label the robot arm puts the white ranch dressing bottle into the metal pot, so the bottle ends up inside the pot precise
matched, cup label the robot arm picks up the blue cup and puts it into the brown cup, so the blue cup ends up inside the brown cup precise
matched, tiger label the robot arm takes the stuffed tiger out of the red bowl and puts it in the grey bowl, so the tiger ends up in the grey bowl precise
matched, marker label the robot arm puts a marker pen into a bowl, so the marker ends up inside the bowl precise
control, not part of the skill the reverse of the tiger task and of the cup task precise

The label searches whose task was in the batch returned 11 matches: 11 real, 0 false, 0 missed against that truth. Each episode with a changed or wrong label, and the cut short one, came out as needs a human look, and each correctly labeled episode came out as matched. Median windows ran from 14.8 to 21.6 seconds, most of each episode, so a window says little about when the task finished; the last frame does.

The label with no footage in the batch, "put the marker into the bowl", returned 2 matches, both false: the reason text called the ranch bottle a marker pen. A description that gave the marker's shape and size, after 1 rewording on the same corpus, again returned 2, both false, one of them a tiger episode with an invented marker. Neither changed a row, because a match only counts under the episode's own label, but it is why hints need frames too. We also searched the reverse of the tiger task and of the cup task: the reversed tiger task returned 0, and the reversed cup task returned 1 false (a tiger episode described as cups).

In this dataset each task has its own objects on the table, so most wrong labels could be caught from which objects are present. Batches where several tasks share one scene and differ only in the action were not tested; expect weaker results there and review more rows by hand. The run used 8 precise searches, an estimated 40 credits.

Done when every label has a saved result, every returned episode has a real or false judgment in state.json, and the user has seen the counts.

Step 7: Build the review CSV

Goal: label_spot_check.csv, one row per episode, each with a last frame and a contact sheet.

  1. Last frame of every episode:
ffmpeg -v error -y -sseof -0.3 -i EPISODE_FILE -frames:v 1 label-check/review/NAME_last.png

NAME is the episode file's base name.

  1. Contact sheet at one frame per second: the match window for matched rows, the whole episode for the rest. START and LENGTH are in seconds; for episodes longer than 25 seconds, replace fps=1 with fps=25/LENGTH so the sheet still fits 25 tiles:
ffmpeg -v error -y -ss START -t LENGTH -i EPISODE_FILE -vf "fps=1,scale=320:-1,tile=5x5" -frames:v 1 label-check/review/NAME_window.png
  1. For needs a human look rows, vivu_get_video_summary (project_id, video_id, include_segments true) gives a short description of what the episode shows; it can go in the reply as a second signal. In our test run it described the cut short episode as the arm hovering over the pot (it called it a metal bowl) and then grasping the bottle, without saying the bottle went in, and episode_000499 as cup stacking. It is not a verdict.
  2. Write label_spot_check.csv with the columns from Step 5, file names as the only link to the video (no Vivu links). Show the user the matched and needs a human look counts and the first rows.
  3. The CSV stays in the working folder. If the user wants it in a shared drive or a tracker, name the destination, show the rendered file and wait for approval; that was not exercised in our test run.

Done when label_spot_check.csv has one row per episode in episodes.csv, every last_frame and window_sheet path exists, and the user has the file.

Compliance

  1. Rights to the footage: use it on data your team owns or may process. For public datasets, follow the dataset license and keep the videos for internal checks. Confirm before Step 4.
  2. People on camera: lab recordings often catch operators or passers-by. Ask whether they know the recordings are processed; leave out episodes that show people who did not agree. Confirm before Step 4.
  3. Minors: if a home collection might show children, get explicit guardian consent or leave those episodes out. Confirm before Step 4.
  4. Where the videos go: uploaded episodes stay in the user's Vivu project until the user deletes them. The skill creates the project as private and never deletes projects or videos unless the user asks, and then confirms first.
  5. What the skill does not do: no face or identity recognition, no operator scoring, no failure diagnosis, no annotation export.

Known failure modes

Symptom Cause Fix
(observed) a label with no footage in the batch still gets matches, with reason text such as "the white marker pen with a green cap" Vivu matched a lookalike object (the ranch bottle) and described it as the labeled object count a match only under the episode's own label and only after the frames show the object; treat other labels' matches as hints
(observed) a reworded description returns an episode with an object that is not there, "a black, slim pen-shaped marker" the reason text invents objects judge from the start, middle and end frames, never from the reason
(observed) a reversed task description returns an unrelated episode, "the blue cup inside the brown tiger-striped cup" the search can return a confident match for an action nobody did keep the end state clause in every description and check the end frame
(observed) match windows cover most of each episode precise windows contain the moment but are wide read the end state from the last frame and the contact sheet, not from the window bounds
(observed) a public dataset label names a task the video does not show the collection label was wrong at the source this is what needs a human look is for; do not trust labels over frames
an episode that did the task is not returned recall is below perfect for actions and visual states, and maximum_results caps it leave it as needs a human look; raise maximum_results to the episode count
every episode of one label comes back as needs a human look the description does not match what the camera shows (wrong color name, wrong camera) compare one last frame with the description, fix the wording with the user, run that label again
tasks that share one scene and differ only in the action are confused untested; objects alone cannot separate them review more rows by hand and add the end state to each description
"has not granted vivu.write" Vivu connected read only reconnect Vivu with write access
upload page asks to sign in or shows an error the one time upload link expires after 180 seconds request a new link right before opening it
a file is rejected by the browser upload tool file above the tool's size limit the user adds that file in the Vivu web app
rows do not line up with results Vivu normalized the file name match on the episode index and duration_ms
the camera is blocked by the arm in the last frame the arm parks over the scene look at the contact sheet, or check a second camera for that row

Connect Vivu, then add the skill.

The skill runs through the Vivu connector. Add it to Claude from the connector directory, then install the skill.