---
name: drone-flight-environment-shortlists
description: "Sort a batch of drone flight videos into one candidate CSV per environment (water, trees or wires, building gaps, landing pads) with Vivu. Use when curating flight training data."
---Drone flight environment shortlists with Vivu
This skill turns a batch of onboard camera videos from recent drone flights (forward or downward camera streams your team has already exported to mp4) into one CSV per environment class on the team's list: flying low over water, flying close to trees or power lines, flying through a narrow gap between buildings, and a landing pad or ground marker in view below. Claude measures the batch, prices it against the user's Vivu plan, uploads it to a private Vivu project, runs one precise search per class, cuts a contact sheet at half second steps for every window Vivu returns, labels each window looks right, looks wrong or can't tell, and writes every candidate into its class's CSV with an empty reviewer_verdict column. The data engineer decides what goes into the dataset.
The value is in what only the picture holds. Flight IDs and file names carry a date and an airframe, and the flight log carries position, altitude and attitude; none of them says whether the camera was looking at a lake, a tree line, a power line or a painted H. Two flights over the same GPS point can face water or forest depending on heading, and exported camera streams have no audio or captions to search. Vivu's search is used only to shortlist: in our test run the same wording, run twice on the same footage, returned different candidate lists, and thin wires showed up on a contact sheet only as faint lines. So every candidate is labeled from its frames, rejected candidates stay in the CSV, and a short list never means the batch had no such stretch.
When to use
Use when someone asks to find the low over water stretches in this week's flights, to pull tree or power line close passes for an obstacle dataset, to shortlist landing pad approaches for precision landing data, or to sort a batch of flight videos by environment before labeling.
Not for these:
- A one off question about one flight ("is there water in this file?"): search Vivu directly and look at the frames.
- Relative motion events such as another aircraft crossing, a bird strike or a gust drift. The skill lists static environments only; motion between objects is not something it searches for.
- Safety judgments, crash or failure analysis, or anything real time or on the aircraft. The CSVs are review lists made after the flight.
- Frame accurate labels, masks or bounding boxes: use the team's annotation tool. This skill makes review lists, not labels.
- Hundreds of hours at once. Pick a subset from the flight logs first (sites, dates, altitude bands); one run is sized for about an hour of footage.
- Night, fog, rain or thermal footage as the main material: untested. The skill can run, but treat its output as unmeasured and have the reviewer look at every row.
Working principles
- Report measured numbers, not estimates. When a number is an estimate, say so.
- Nothing is verified until it has been checked against the source video. Vivu's results are candidates, Claude's labels are a first pass, and the reviewer's verdict is the decision.
- Stop and tell the user when a required capability or tool is missing. Do not guess around it.
- Ask the user before anything that is expensive to redo (the batch and its index minutes, before upload) and before writing the full CSVs (one sample row first). The skill posts and sends nothing; if the user asks Claude to put the CSVs somewhere, ask again before each write.
- Labels come from frames, never from Vivu's reason text, which describes the class it was asked for even when the frame shows something else.
- The CSVs keep every candidate and carry no recall or precision figures. A search misses some stretches and caps how many come back, so an empty or short list does not prove the batch has none.
What you need before starting
Check each item at the start of the run and tell the user plainly what is missing before doing anything else.
| Requirement | Why | How to check |
|---|---|---|
| Vivu connector with write access | create a private project, open its upload page, search | vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the user reconnects Vivu and allows write access |
| The flight videos as mp4 files on the user's computer | Vivu indexes video files, and every sheet is cut from the local file | ffprobe -v error -show_entries format=duration -of csv=p=0 FILE prints a length for each file; ROS bags, MCAP or ULog camera topics must be exported to video by the team's own tools first |
| A shell on that computer with ffmpeg and ffprobe | measure the batch, cut contact sheets and zoomed frames | ffmpeg -version, ffprobe -version |
| An upload path | move the files into Vivu | vivu_open_upload_page plus the user's own browser, or a browser tool that can attach local files |
| Someone who reviews the CSVs | Claude's labels are a first pass; the reviewer decides | ask who fills reviewer_verdict; the default is the user |
This skill needs Claude Code on the user's computer (the terminal or the Code tab of Claude Desktop), because it reads local files and runs ffmpeg. Claude on the web and cloud sessions cannot reach the files. It downloads nothing, so no residential IP is needed. Nothing recurs, so no scheduler is involved; run it once per batch.
Inputs to collect
Ask for anything missing, most important first.
- BATCH_DIR: the folder that holds the exported mp4 files for this batch. Required.
- The environment classes and what counts for each. Default: the four classes in Step 4 with the label rules in Step 5. If the team defines a class differently (for example, a river seen from 50 meters counts as water, or a single wall counts as a building gap), the team's definition wins; write it into config.json before labeling.
- BATCH_NAME: a short name such as a site and a date. It names the working folder and the Vivu project. Default: the name of BATCH_DIR.
- Batch size. Default: up to 60 minutes of video, or what remains of the plan this month, whichever is smaller.
- Who reviews the CSVs. Default: the user.
Files and state
Keep everything in one working folder next to BATCH_DIR:
envlists-BATCH_NAME/
config.json BATCH_DIR, classes with query, maximum_results and label rules, Vivu project id
files.csv source_file, duration_s, listed_name, video_id
results/ search results without the result page link, one JSON per class
sheets/ one folder per class with contact sheets and zoomed frames
CLASS.csv one per class: every candidate, Claude's label, empty reviewer_verdict
state.json steps done, files uploaded, classes searched with job ids, classes cut off, candidate keys labeled
The commands in Steps 2 to 5 run from the folder that holds both BATCH_DIR and envlists-BATCH_NAME/, so the relative paths in them resolve.
A rerun reads state.json first and skips finished steps. A file already marked uploaded is never uploaded again, a class with a saved result is not searched again, and a candidate key (video_id, start_ms and class) that already has a label is not labeled again. The next batch gets its own folder and its own Vivu project, so its searches do not return last week's candidates.
Step 1: Check the Vivu connector and the setup
Goal: confirm every requirement before spending anything.
- Call vivu_get_account. If the tool does not exist, tell the user to add the Vivu connector in Claude (
https://mcp.vivu.ai/mcp) and stop. If the account does not showcan_create_projects: true, or a later write call fails with "has not granted vivu.write", ask the user to reconnect Vivu and allow write access, then stop until they have. - Run
ffmpeg -versionandffprobe -version. - List BATCH_DIR and run the ffprobe length command on one file.
Done when vivu_get_account shows can_create_projects: true, ffmpeg and ffprobe print their versions, and ffprobe prints a length for a file in BATCH_DIR.
Step 2: Measure the batch and price it
Goal: an indexed set the user's plan can pay for, approved before upload.
- Run the ffprobe length command on every mp4 in BATCH_DIR and write files.csv with source_file and duration_s. Add listed_name: the file name with each space, bracket and plus sign replaced by an underscore, because that is how Vivu will list it.
- Call vivu_get_usage and show one table:
| This batch | Plan allowance | Remaining this month | |
|---|---|---|---|
| Index minutes | sum of duration_s / 60 | from vivu_get_usage | from vivu_get_usage |
| Search credits | 5 per class, 20 for the default four (estimate), plus 5 for each rewording | from vivu_get_usage | from vivu_get_usage |
Plan facts: Free is $0 a month with 20 index minutes a month and 50 search credits a month; Premium is $30 a month with 180 index minutes and 500 search credits. A precise search uses 5 credits in total. A typical batch of ten flights of four minutes each is about 40 index minutes (estimate), which is more than Free covers in a month and fits in Premium.
- If the batch does not fit, offer levers in this order: pick a smaller subset with the flight logs (the sites, altitude bands or dates the dataset needs most), split the batch across two runs or two months, and only then a larger plan. Never drop files the user asked for without saying which ones.
- Confirm the Compliance items with the user.
The account used in our test run returns no allowance figures, so the comparison with a real plan and the user's approval were not exercised in our test run.
Done when the user approves the file list and its index minutes, and config.json records both.
Step 3: Upload and index
Goal: every file in the batch indexed in a private Vivu project for this batch.
- Call vivu_list_projects and reuse a project named "Flight env BATCH_NAME" if one exists. Otherwise call vivu_create_project with that name and visibility "private". The default visibility is organization, which shows the project to everyone in the user's Vivu organization, and flight footage over private land is often sensitive.
- Call vivu_open_upload_page with the project ID immediately before uploading. The link expires in 180 seconds and works once, so never post or store it. Give it to the user to open in their own browser and select the files, or open it in a browser tool that can attach local files. Claude in Chrome accepts at most 10 MB per upload call and flight videos are usually larger, so they normally go through the user's own browser or the Vivu web app. Never split or recompress a file to fit.
- Poll vivu_list_videos until every file shows ready. Match each Vivu file name to files.csv by listed_name and record its video_id there and in state.json.
In our test run the 12 videos (19.17 minutes) were all ready 244 seconds after the upload started. They went through the same upload page with an automated browser, so opening the link in the user's own browser was not exercised in our test run.
Done when every file in files.csv shows ready and has a video_id.
Step 4: Search each environment class
Goal: a saved list of candidate windows for every class.
| Field | Query | Mode | maximum_results |
|---|---|---|---|
| over_water | the drone flies low over water, so a lake, river or sea surface fills most of the picture below the horizon | precise | 40 |
| trees_or_wires | the drone flies close to trees or power lines, so tree canopy, branches, wires or a transmission tower fill a large part of the picture | precise | 40 |
| between_buildings | the drone flies through a narrow gap between buildings, with building walls close on both sides of the picture | precise | 40 |
| landing_marker | a landing pad or helipad marking, such as a painted H, a circle or a printed target board, is visible on the ground below the drone | precise | 40 |
The searches run side by side, one per class, because each class is its own list for the reviewer. Within a class the chain is: the precise search finds candidate windows, the contact sheet shows what the camera saw across each window, Claude labels it, and the reviewer decides. Every search is precise because only precise returns a time window; fast returns whole files with an empty reason, and the batch is already a chosen subset. Exported onboard camera streams have no audio and no on screen text about the scene, so there is no second modality to switch to.
maximum_results is also the ceiling on how many candidates come back. 40 leaves room for an hour of footage with several stretches per flight; a search that returns exactly 40 was cut off, and the batch holds more candidates than one search can return. Then tell the user, mark the class cut off in state.json, and offer to split the batch into two smaller batches with their own projects (quote the index minutes again and wait for approval as in Step 2). In our test run no search came close to 40, so a cut off search was not exercised in our test run.
Run each query with vivu_search_videos (project_id, query, mode, maximum_results). It returns a job ID. Call vivu_get_search_results until complete is true; each status call can wait up to 45 seconds, so a pending search is not a stalled one. Save each completed result as results/FIELD.json without its result_page_url field, and record the job_id in state.json. Show the Vivu result page link in the live reply only; it expires after four hours, so it never goes into a saved file or a CSV.
Run each class once. Running the same query again on the same project can return a different list (see the worked example), so a second run is not a check on the first. If the team changes a wording, the new wording is untested until its first candidates have been through Step 5.
Done when every class has a completed result saved in results/, and every class whose search returned exactly its maximum_results is marked cut off in state.json.
Step 5: Make contact sheets for every candidate and label it
Goal: every candidate labeled looks right, looks wrong or can't tell from what its frames show.
For each result in results/FIELD.json, cut contact sheets of its window from the local file at two frames a second:
mkdir -p envlists-BATCH_NAME/sheets/FIELD
ffmpeg -v error -y -ss START -to END -i FILE -vf "fps=2,scale=384:-2,tile=5x4" envlists-BATCH_NAME/sheets/FIELD/rRANK_STEM_STARTMS_%02d.png
FIELD is the class, FILE the local source file (BATCH_DIR/ followed by its source_file), STEM its name without .mp4, RANK the result_number written with two digits (r01, r09) so the files sort by rank, STARTMS the result's start_ms, and START and END the result's start_ms and end_ms divided by 1000. Each sheet holds 20 tiles covering 10 seconds of the window; the last sheet is padded with black tiles. Tile k on sheet n (both counted from 0, left to right and top to bottom) is at START + 10 * n + k / 2 seconds. The tile position gives each frame's time, so the command needs no text overlay. Look at every sheet of a window: windows can be shorter or longer than the stretch they point at.
Wires, poles and far markers are only a few pixels wide on a tile. When a tile shows a thin line, a pole or a small shape on the ground, or the sheet looks empty but the reason names wires or a marker, extract a zoomed full size frame at that tile's second:
ffmpeg -v error -y -ss SECONDS -i FILE -frames:v 1 -vf "crop=iw/2:ih/3:X:Y,scale=iw*2:-2" envlists-BATCH_NAME/sheets/FIELD/rRANK_STEM_SECMS_zoom_PART.png
SECONDS is the time of the tile where you saw the line, SECMS the same time in milliseconds, X is 0 for the left half or iw/2 for the right half, Y is 0, ih/3 or ih*2/3 for the top, middle or bottom band, and PART names the crop (for example left_bottom) so two crops of the same second get different files. Take the crop at the tile's own second and the band where the line sat: a wire moves through the frame fast, and a crop one or two seconds off can miss it entirely.
Label rules (defaults; the team's definitions from Inputs win):
- over_water looks right when water (a lake, river, sea or frozen lake surface) fills most of the picture below the horizon and the drone is low enough that ripples, reflections or ice texture are clear. It looks wrong for a river seen from high above, wet roads, glass or car roofs reflecting the sky, and tiles where the water is only at the far edge.
- trees_or_wires looks right when a tree crown, branches, a hedge of trees, a wire or a transmission tower is close to the camera and takes up a large part of the picture (roughly a quarter or more), or when wires clearly cross the frame. It looks wrong for trees in the middle distance across a field, a forest seen from high above, crop rows, low bushes on grass, and bridges or other steel structures.
- between_buildings looks right when building walls are close on both sides of the picture at the same time. It looks wrong for flying past a single wall, over rooftops, or down a wide street with low houses.
- landing_marker looks right when a landing pad, helipad H, painted circle or printed target board is visible on the ground below the drone. It looks wrong for crosswalks, runway or road lines and other paint that is not a landing marker.
- can't tell when the zoomed frame still does not settle it (the object is hidden, too small or blurred). The reviewer looks at the source.
Write a note of what the frames show for every candidate, such as "calm river fills the lower half, bare trees reflected" or "four wires cross the river below, clear only on the zoom at 00:32". Never copy the reason into the note. Record each candidate key and its label in state.json.
Worked example from our test run
The test corpus was public onboard footage published under a Creative Commons license: 12 videos (19.17 minutes) cut at fixed offsets from forward and downward cameras on quadcopters, FPV quads and one fixed wing, with the audio removed and the files renamed so the names said nothing about content. Low over water, trees and wires, a gap between buildings and landing markers were marked by hand on contact sheets before any search, together with look alikes (crop rows, a truss bridge, a crosswalk seen from above, a river seen from high up) and stretches too ambiguous to count either way. Some clips carried overlays that exported camera streams do not have (a detection overlay on the indoor landing pad clip, the wing and a date stamp on the fixed wing clip, propellers in view on one FPV clip).
- over_water: 5 candidates, 5 real and 0 false, and 2 of the 7 stretches marked by hand were missed, both over a frozen lake. A river seen from high above and a car roof reflecting the sky were not returned.
- trees_or_wires, with the wording in the table: 6 candidates, 6 real and 0 false, and 2 stretches marked by hand were missed (the drone just above a tree crown, and a pass beside bushes and small trees). This is the wording we settled on after 2 rewordings on the same corpus; the final wording is the first wording run again, so the numbers below show how much one run can differ from another.
- The first run of that same wording returned 9 candidates, 8 real and 1 false, and missed 2 stretches. The false one was a flight over dry grass with trees in the middle distance, while the reason said trees and branches filled the view. Three of the nine came from a suburban clip whose second run returned none of them.
- A tighter wording ("within a few meters ... trees standing in the distance across a field do not count") did worse: 7 candidates, 6 real and 1 false, and 3 stretches missed, including both passes low over a forest canopy. The false one was a flight over a grassy bank with small bushes.
- The one wire stretch (power lines crossing a river, seen from above) was returned in every run, but on the contact sheet the wires were faint lines. A zoomed crop at the tile's second showed the wires and a pole clearly; crops taken a little before and after that second showed none.
- between_buildings: 1 candidate, 1 real, 0 missed. The corpus had only one real gap between buildings, so this says little; passes beside a single wall were not returned. The window ended while the walls were still on both sides.
- landing_marker: 3 candidates, 3 real and 0 false, and none of the 3 stretches marked by hand missed: a printed H under a downward camera, printed target boards on a lawn and a painted circle with an H on a runway apron. A crosswalk seen from above and runway lines were not returned. Two of the windows covered the whole clip. All three markers were in view for many seconds; a marker that is far away or in view for a second or two is untested.
- The 12 videos were all ready 244 seconds after the upload started. The CSVs held 15 rows in total.
Done when every candidate in results/ has at least one contact sheet in sheets/ and a label in state.json.
Step 6: Write one CSV per class and hand it over
Goal: a review list per class that a data engineer can work through without opening Vivu.
Show the user one sample row and the field mapping, and wait for an OK before writing the rest:
source_file,rank,start_mmss,end_mmss,search_reason,contact_sheets,claude_label,claude_note,reviewer_verdict
FLT_08.mp4,4,00:25,00:38,"The drone flies low and close over electrical power lines stretched across the river, with the wires clearly traversing the screen below.",sheets/trees_or_wires/r04_FLT_08_25000_01.png;sheets/trees_or_wires/r04_FLT_08_32000_zoom_left_bottom.png,looks right,"four thin wires and a wooden pole cross the river well below the drone; clear only on the zoomed frame at 00:32",
The row above is adapted from our test run (file names shortened); the user's rows carry their own file names.
| Column | Source | If unavailable |
|---|---|---|
| source_file | files.csv, matched by listed_name | keep the row with Vivu's file name, label it can't tell with the note "no local file matched", and tell the user |
| rank | result_number in results/FIELD.json | none |
| start_mmss, end_mmss | start_ms and end_ms, written MM:SS | none |
| search_reason | the result's reason, copied as is; it is Vivu's description, not evidence | blank |
| contact_sheets | the sheet files and zoomed frames from Step 5, separated by semicolons | none; the row cannot be labeled without them |
| claude_label | Step 5 (inferred from frames) | can't tell |
| claude_note | what the frames show, in Claude's words (inferred) | blank |
| reviewer_verdict | left empty for the reviewer | empty |
claude_label and claude_note are Claude's reading of the frames, not a measurement. Borderline cases (bankside trees seen from over a river, a bare tree beside a building) depend on the team's definition, so the reviewer goes through every row, the looks right ones included, before anything enters the dataset.
Write FIELD.csv for every class with one row per result, sorted by rank, including the rows labeled looks wrong: a reviewer may disagree with a label, and the rejected rows show what the search confuses. The CSV uses the user's own file names and times, never a Vivu result page link, so it stays usable after the link expires.
Then tell the user, per class, how many candidates came back, how many Claude labeled looks right, looks wrong and can't tell, and whether the list was cut off at maximum_results. Say plainly that an empty or short list does not mean the batch has no such stretch, and that thin wires and far markers are the easiest to miss. Do not add recall or precision figures: the hand marks needed to measure them do not exist for the user's batch.
The CSVs stay in the working folder. Passing them on is the user's step. If the user asks Claude to upload or post them (a shared drive, a tracker, a chat), name the destination, show the rendered first rows and wait for a yes before each write.
In our test run the CSVs were written from the dry run's results and sheets; showing the sample row to a user and handing the CSVs over were not exercised in our test run.
Done when every class has a CSV whose row count equals the number of results in results/FIELD.json, the user approved the sample row and the mapping, and state.json marks the batch done.
Compliance
- Use flight footage the team owns or has the rights to analyze. For public flight videos used as extra material, check the license and the platform's terms first and keep them for internal analysis only. Have the user confirm this before Step 3.
- People and private property on the ground did not agree to be filmed. The skill lists stretches by environment only; it does no face, license plate or identity recognition and never searches for people, homes or vehicles. Do not use it to collect footage of particular people, particular homes or of children. Follow the team's data use rules (some teams blur faces or crop private yards before any upload), and confirm before Step 3.
- The skill makes no safety judgment, does not say why a flight failed, and does nothing on the aircraft or in real time. The CSVs are review lists for people, not annotation files.
- The videos stay in the user's Vivu project until the user deletes them. Create the project as private; the default is visible to the whole Vivu organization. Delete videos or the project only when the user asks, and confirm first.
- The skill writes only local files. Any upload or post of the CSVs goes through the approval in Step 6.
Known failure modes
| Symptom | Cause | Fix |
|---|---|---|
| (observed) the same wording on the same project returned 9 candidates on one run and 6 on the next | search results vary between runs | run each class once and label what came back; never treat a second run as a check on the first |
| (observed) a tree candidate shows a field with trees in the middle distance: 9 returned, 8 real and 1 false | the reason describes trees filling the view that the frames do not show | label from the sheet; trees across a field are looks wrong |
| (observed) a tighter tree wording lost the passes over a forest canopy: 7 returned, 6 real and 1 false, 3 missed | extra exclusions pushed out real close passes | keep the wording in the table; if the team rewrites it, treat it as untested until its candidates have been labeled |
| (observed) power lines look like faint lines or nothing on the contact sheet | a tile is a small image and a wire is only a few pixels wide | zoom at the tile's own second and band (Step 5); in our test run crops taken at other seconds missed the wires |
| (observed) stretches low over a frozen lake were not returned: 5 returned, 5 real, 2 missed | ice reads less like water than open water does | an empty result does not prove the batch has none; say so in the summary |
| (observed) a window covers the whole clip for a landing marker, or ends while the walls of a building gap are still on both sides | windows do not match the length of the stretch | look at every sheet and set start_mmss and end_mmss from what the tiles show if the team needs tight edges |
| (observed) a bush on a grassy bank came back as a close tree pass | low vegetation matches the tree query | label it looks wrong under the default rule, or change the rule in config.json if the team counts bushes |
| upload page asks to sign in or shows an error | the one time upload link expires after 180 seconds | request a new link right before opening it |
| a file is rejected by the browser upload tool | the tool accepts at most 10 MB per call | the user adds the file in their own browser or the Vivu web app |
| "has not granted vivu.write" | Vivu connected read only | the user reconnects Vivu with write access |
| a result's file name does not match BATCH_DIR | Vivu replaced a space, bracket or plus sign in the name | match on listed_name in files.csv |
| a search returns exactly its maximum_results | the batch holds more candidates than one search returns | mark the class cut off, say so in the summary, and offer to split the batch |
| a class comes back empty | none in the batch, or the search missed them | an empty result does not prove the footage has no such stretch; say so in the summary |
| labels on night, fog or thermal footage look unreliable | this footage is untested | treat every row as unmeasured and have the reviewer check each one |