Footage for training and evaluation

How to find similar failures across robot episodes

Which route works depends on whether the failure was recorded as a failure. If something in the stack flagged it, filter on the flag and the job is mostly done. If nothing flagged it, there are three routes that actually exist: describe what the failure looks like on camera and search the video for it, build a detector for it, or watch episodes. One translation is required before any of them. "Similar to this episode" has to become "similar to this description", because a description is what a search can act on.

Exhaust the metadata first

Episode metadata is the cheapest place to look: success and failure flags, task identifiers, operator notes, error codes, the timestamp where the controller gave up. If the failure mode you care about was already written down at record time, this is the whole job and nothing below it is worth reading.

The limit is that metadata contains only what someone decided to record. A flag saying the episode failed does not say the gripper closed early, and it does not say where in the episode that happened. Ask a general-purpose assistant how to find a specific thing across a pile of footage and you get a reasonable two-part answer: be disciplined about metadata up front, and use software that makes the footage searchable. The first half is correct and it is also the catch, because metadata discipline only covers failure modes you knew to name before you had them, and the ones worth hunting are usually the ones you did not.

Build a detector for the failure

If a failure mode is well defined, recurring, and visible in signals you already log, a classifier or a hand-written heuristic over your own data is the strongest route available, and it gets sharper every time you run it.

It has two costs. One is engineering time, and it recurs: the detector gets rebuilt when the gripper changes, when the scene changes, or when a camera moves. The other is an ordering problem. You need a set of labeled examples to build the thing whose job is to find examples, which is the same wall you were trying to get over.

Describe what it looks like and search the video

The newer route treats recorded video as something you can question directly. Episodes go into a search layer, get read once, and then you ask in plain language for what the failure looks like: the object slipping out of the gripper, the arm continuing after contact, the object still on the table while the arm retracts. What comes back is candidate ranges with a line of reasoning attached to each, and a person opens them one at a time before any of them count as data. Vivu is one implementation of this route: episode video is indexed once on the way into a project, and after that a failure you can put into words becomes a query whose answer is ranges to open. Camera streams living inside ROS bags or MCAP files have to be exported to ordinary video files first, which is its own small pipeline job. If the asking happens inside an assistant rather than a search box, running the search from an agent is the same route with a different front end.

Two limits decide whether this is any use to you. It sees what the camera saw, so a failure that exists as a force spike, a controller timeout, or a plan that was already wrong before the arm moved is invisible to it. And borderline candidates are genuinely borderline. A segment that touches the failure indirectly, or shows a near-miss instead of the failure, arrives in the same list as the clean matches, and sorting those apart is human work. That is the argument for a review step between the search and the set, and reviewing what a search handed back is where data quality is actually decided.

Watching episodes is still the baseline

Sampling by hand is what every other route gets measured against. It recognizes everything, including failure modes that do not have a name yet, and its cost rises directly with the size of the archive. On a small archive it wins outright and the rest of this is premature. The other routes exist because archives grow and sampling does not.

When you don't need any of this

If the failure is already caught by a metric in your evaluation loop, searching video for it is redundant work. If what distinguishes the failure is not on camera, video search cannot reach it at all and no amount of query rewriting will fix that. And if you are chasing one particular incident rather than a class of them, the shape of the job is different: that is closer to hunting a rare scenario in driving footage than to assembling a set.

So the dividing line is whether the failure has a look. If it does, describing it is the cheapest way to assemble candidates out of episodes nobody flagged, and the review step is where the real cost lands. If it does not, the work belongs in your signals and your detectors, and video search will only produce confident-looking noise. Both kinds of failure usually live in the same archive, and the two routes do not substitute for each other.

FAQ

Can I give it an example clip of a failure and get similar episodes back?

Not in the way that phrasing suggests. The query these tools take is a description, so "similar to this episode" has to be rewritten into something like "shows the object slipping out of the gripper while the arm lifts". That is usually a better query anyway, because it forces you to name what the similarity actually is. If the only thing two episodes share is a pattern in joint or force data, no description will capture it and this is the wrong layer.

Do I have to export camera streams out of ROS bags before searching them?

Yes, if the recordings live in ROS bags or MCAP files. Video search works on ordinary video files, so the camera topics have to be written out as video first. For most setups this is a one-time script. The decision worth making carefully is which cameras you export, because the views you skip are invisible to every search you run later, and discovering that after you have indexed a campaign is an expensive way to learn it.

Can one search cover every robot and every project at once?

No. Each search runs inside a single project, so how you split episodes into projects decides what one question can reach. Splitting by robot, by task, or by collection campaign each makes a different set of questions easy and a different set awkward. It is worth choosing that split based on the questions you expect to ask, not on how the files happened to land on disk.

Does the search tell me why the episode failed?

No, and treating it as though it does is the fastest route to a bad dataset. It returns segments that match a description you wrote. Why the robot failed is a question for your logs, your signals, and a person watching the segment, and the search result is only the thing that put that segment in front of them.