SkillsSkill for Claude

Recipe card from cooking videos

Hand Claude the cooking videos on your drive and a rough list of steps, and it returns a recipe card draft for your blog or video description: each step with the amount caption read from the picture, what you said when there is narration, a step photo from the footage, a finished-dish photo, and a link to your public video at that second. Many recipe videos print amounts on screen and never say them, and some have only music, so transcripts miss the numbers and the photos can only come from the footage. Every amount is copied from a full frame, never from Vivu's reason text, and anything unreadable is marked for you to fill in.

Maintained by Vivu. Updated 2026-10-03.

Download

recipe-card-step-frames.zip

9 KB. Unzips to recipe-card-step-frames/SKILL.md. Upload the zip as it is in the Claude app, or unzip it into your skills folder for Claude Code.

SHA-256 02442a8fb20d990ef3fd3bf3528b0e73d5b74f2438a48282fa6060875071ae4e

At a glance

What the Recipe card from cooking videos skill does, where it runs, what it needs, and when it asks
Looks forAmount captions printed on screen (3 TBSP, 210°C, 8 to 10 minutes), the moment each step happens, and the finished-dish shot. Each window is checked against full frames before it reaches the card, and a local contact sheet of every video catches captions the search missed.
Runs onClaude Code on your computer (the terminal or the Code tab of Claude Desktop), because it reads your local video files and runs ffmpeg. A residential IP is needed only if you download your own published video with yt-dlp. Nothing recurs, so there is no scheduler. Claude on the web and cloud sessions cannot run it.
Needs
  • The Vivu connector with write access, to create a private project, open its upload page and search.
  • ffmpeg and ffprobe on the computer that holds the videos, to measure them and extract frames, contact sheets and step photos.
  • Your video files on that computer, because the photos and caption reads come from them.
  • An upload path: the upload link opened in your own browser, or a browser tool that can attach local files.
  • Optional: your exported captions file or a local transcript, so quotes of what you said are exact.
  • Optional: yt-dlp and node on a residential IP, only to download your own published video.
Your Vivu planIndexing uses one index minute per minute of video; a recipe takes about six precise searches, around 30 search credits (estimate). Claude measures the videos and shows the cost against your plan from vivu_get_usage before uploading.
Asks you firstConfirms you own or may use the videos, shows the index minutes and credits against your plan and waits for approval before upload, and shows one sample row of the card before writing the rest. It never publishes the card; you paste it yourself.

The skill does its video work through the Vivu connector. If the connector is not in your Claude yet, add it first; the skill checks that it is connected before it does anything else.

Before you run it

  • Use only videos you made or have the rights to; leave faces, especially children's, out of the step photos unless everyone agreed.
  • Search recall for on-screen captions is weak, especially on vertical or low resolution video; the skill sweeps every video with a local contact sheet, and an empty result does not mean there are no amounts.
  • Finding a step from the picture alone is weak; those windows are candidates that Claude checks frame by frame before using them.
  • Vivu's reason text paraphrases; amounts come only from full frames and quotes only from your captions file or a local transcript.
  • Claude in Chrome uploads at most 10 MB per call, so most recipe videos go through your own browser or the Vivu web app.
  • The videos stay in your Vivu project until you delete them; the project is created as private.
  • The skill writes local files only; publishing the card is your step.

Start it

Once the skill is installed, ask for the task in your own words. Naming the skill is the most reliable way to have Claude use it. For example:

Make a recipe card from my salmon meal prep video in ~/Videos/salmon: read the amounts off the screen, give me a photo for each step and a finished shot, and link each step to https://www.youtube.com/watch?v=VIDEO_ID at the right second.

In Claude Code you can also type /recipe-card-step-frames. Claude asks for anything the request leaves out, most important first.

What is inside

  1. When to use
  2. Working principles
  3. What you need before starting
  4. Inputs to collect
  5. Files and state
  6. Step 1: Check the Vivu connector and the setup
  7. Step 2: Collect the videos, the step list and the links
  8. Step 3: Measure and price
  9. Step 4: Upload and index
  10. Step 5: Search for amount captions
  11. Step 6: Read every caption from full frames
  12. Step 7: Sweep each video for captions the search missed
  13. Step 8: Find the steps that have no amount caption
  14. Step 9: Pick the step photos and the finished shot
  15. Step 10: Assemble the recipe card
  16. Compliance
  17. Known failure modes

The full skill

This is recipe-card-step-frames/SKILL.md from the download, as Claude reads it: the frontmatter first, then the instructions.

---
name: recipe-card-step-frames
description: "Turn your own cooking videos into a recipe card with Vivu: amounts read from the on-screen captions, one photo per step, a finished-dish shot and timestamped links. Use before writing the blog post."
---

Recipe card with step frames from your cooking videos

This skill takes the cooking videos a creator has on their own drive, plus a rough list of the recipe's steps, and returns a recipe card draft to paste into a blog post or a video description: one row per step with the source file, the time, the amount caption read from the picture (for example "3 TBSP | SOY SAUCE"), what the creator said in that step when there is narration, a photo of the step taken from the footage, a finished-dish photo, and a link to the creator's public video at that second. Claude measures the videos, prices the job against the user's Vivu plan, uploads them to a private Vivu project, runs one search for every amount caption and a few searches for the steps that have none, checks every window on full frames, sweeps each video with a cheap local contact sheet for captions the search missed, and writes recipe_card.csv and recipe_card.md.

The value is in what only the picture holds. Many recipe videos print the amounts on screen and never say them; plenty have no narration at all, only music. Auto captions and transcripts miss every one of those numbers, and the step photos can only come from the footage. The risk is a wrong amount in a published recipe, so every amount in the card is copied from a full frame, never from Vivu's reason text, and anything Claude could not read is marked UNCERTAIN for the creator to fill in.

When to use

Use when a creator says "make a recipe card from my salmon video", "pull the ingredient amounts and step photos from these three cooking videos for my blog", "write the recipe for my video description with timestamps", or "I need step by step photos from my recipe footage".

Not for these:

  1. One question about one video ("how long did I bake it for?"): search Vivu directly and look at the frame.
  2. Subtitles, captions files, edits, short clips or social cuts. This skill reads the footage and writes a card; it does not cut or deliver video.
  3. Someone else's videos. Use only footage the creator owns or has the rights to.
  4. Nutrition facts, scaling a recipe, or converting units. The card copies what the video shows; Claude can help with those separately, labeled as its own calculation.

Working principles

  1. Report measured numbers, not estimates. When a number is an estimate, say so.
  2. Nothing is verified until it has been checked against the source. A Vivu window is a candidate until a full frame shows the caption; the amount in the card is copied from that frame. Spoken words come from the creator's own captions file or a local transcript, never from the reason text, which paraphrases.
  3. Stop and tell the user when a required capability or tool is missing. Do not guess around it.
  4. Ask the user before anything that is expensive to redo (the index minutes, before upload) and before anything that acts on their behalf. This skill writes local files only; publishing the card is the creator's own step.
  5. A step with no amount caption gets "NO AMOUNT ON SCREEN", and an amount Claude cannot read gets UNCERTAIN. Never fill an amount from memory, from a similar recipe or from the reason text.

What you need before starting

Check each item at the start of the run and tell the user plainly what is missing before doing anything else.

Requirement Why How to check
Vivu connector with write access create a private project, open its upload page, search vivu_get_account shows can_create_projects: true (tool names may carry a server prefix). A write call failing with "has not granted vivu.write" means the user reconnects Vivu and allows write access
A shell with ffmpeg and ffprobe, on the computer that holds the videos measure the videos, extract frames and contact sheets, save step photos ffmpeg -version, ffprobe -version
The video files on that computer the photos and the caption reads come from the local files ls VIDEO_FOLDER
An upload path move the files into Vivu vivu_open_upload_page plus the user's own browser, or a browser tool that can attach local files
Optional: the captions file the creator exported (.srt or .vtt) or a local speech to text tool such as whisper.cpp quotes of what was said come from it open the file; without it the said column is NOT VERIFIED
Optional: yt-dlp and node, on a machine with a residential IP only when the creator no longer has a local copy and downloads their own published video yt-dlp --version, node --version

This skill needs Claude Code on the user's computer (the terminal or the Code tab of Claude Desktop), because it reads local video files and runs ffmpeg. Claude on the web and cloud sessions cannot do that. Nothing recurs, so no scheduler is involved; run it once per recipe. The only connector it needs is Vivu.

Inputs to collect

Ask for anything missing, most important first.

  1. VIDEO_FOLDER and the files that belong to this recipe. Required.
  2. The step list: the creator's draft, one line per step. Default: Claude drafts it from vivu_get_video_summary and the creator corrects it in Step 2.
  3. The public link for each video (YouTube or other). Default: none; the card then gives the file name and the time instead of a link.
  4. WORK_NAME: a short name for the working folder and the Vivu project. Default: the recipe name, for example maple-salmon.
  5. Where the card goes. Default: local files only; the creator pastes the card into their blog or description themselves.

Files and state

Keep everything in one working folder:

recipe-WORK_NAME/
  config.json         VIDEO_FOLDER, Vivu project id, public links, query wording
  videos.csv          file, duration_s, video_id, public_url
  steps.csv           the step list: step, step_text
  results/            one JSON per search, without the result page link
  sheets/             contact sheets of windows and whole videos
  read/               full frames the amounts were read from
  steps/              the step photos and finished.jpg
  candidates.csv      every window Vivu returned or the sweep found, with the verdict
  recipe_card.csv     one row per step
  recipe_card.md      the card to paste
  state.json          steps done, files uploaded, searches run with job ids, windows checked

Run every command from the folder that contains recipe-WORK_NAME/. A rerun reads state.json first: a file marked uploaded is not uploaded again, a search with a saved result is not run again, and a window with a verdict in candidates.csv is not checked again.

Step 1: Check the Vivu connector and the setup

Goal: confirm every requirement before spending anything.

  1. Call vivu_get_account. If the tool does not exist, tell the user to add the Vivu connector in Claude (https://mcp.vivu.ai/mcp) and stop. If the account does not show can_create_projects: true, or a later write call fails with "has not granted vivu.write", ask the user to reconnect Vivu and allow write access, then stop until they have.
  2. Run ffmpeg -version and ffprobe -version, and list VIDEO_FOLDER.

Done when vivu_get_account shows can_create_projects: true, both commands print a version, and the video files are listed.

Goal: the files, the step list and the public links written down.

  1. Confirm the Compliance items with the user.
  2. Write videos.csv with each file and its public link. Rename files that contain spaces or brackets to plain names (for example salmon.mp4), because Vivu replaces spaces and punctuation with underscores and plain names make results easy to match.
  3. Write steps.csv from the creator's draft. When there is no draft, write one after Step 4 from vivu_get_video_summary and ask the creator to correct it.
  4. Only if the creator has no local copy: download their own published video at 720p from a machine with a residential IP (YouTube blocks cloud IPs). 720p keeps caption text readable. VIDEO_ID is the 11 character ID in the creator's link:
yt-dlp --js-runtimes node --restrict-filenames \
  -f "bv*[height<=720]+ba/b[height<=720]" --merge-output-format mp4 \
  --match-filter "duration<=3600" \
  --download-archive recipe-WORK_NAME/archive.txt \
  -P "home:VIDEO_FOLDER" -o "%(id)s.%(ext)s" --retries 5 \
  "https://www.youtube.com/watch?v=VIDEO_ID"

The youtube-competitor-watch skill (https://vivu.ai/skills/youtube-competitor-watch) covers channel verification and download troubleshooting in more depth. If yt-dlp fails with "unable to download video data", retry once with --extractor-args "youtube:player_client=web_embedded".

In our test run the videos were three public Creative Commons recipe videos downloaded this way, standing in for a creator's own files, so collecting a creator's draft step list was not exercised in our test run.

Done when videos.csv lists every file and steps.csv has one row per step (or the user has agreed to draft it after Step 4).

Step 3: Measure and price

Goal: the job priced against the user's plan and approved before upload.

  1. Measure each file and write duration_s:
ffprobe -v error -show_entries format=duration -of csv=p=0 VIDEO_FOLDER/FILE.mp4

FILE is the video's file name without .mp4, as written in videos.csv (for example salmon).

  1. Call vivu_get_usage and show one table:
This recipe Plan allowance Remaining this month
Index minutes sum of duration_s / 60 from vivu_get_usage from vivu_get_usage
Search credits about six precise searches, 30 credits (estimate) from vivu_get_usage from vivu_get_usage

Plan facts: Free is $0 a month with 20 index minutes a month and 50 search credits a month; Premium is $30 a month with 180 index minutes and 500 search credits. A precise search uses 5 credits in total. A fast search uses 1 credit; this skill does not use fast. A 12 minute recipe video uses 12 index minutes, so Free covers one such video a month (estimate).

  1. If the job does not fit, offer levers in this order: index only the final cut, not the raw takes; index only the videos for this recipe; split the work across two months; and only then move to a larger plan. Never trim or recompress a video to save minutes.

The account in our test run returns no allowance figures, so the plan table and the user's approval were not exercised in our test run.

Done when every file has a duration_s and the user has approved the minutes and credits.

Step 4: Upload and index

Goal: every video indexed in a private Vivu project.

  1. Call vivu_list_projects and reuse a project named "Recipe WORK_NAME" if one exists. Otherwise call vivu_create_project with that name and visibility "private". The default visibility is organization, which shows the project to everyone in the user's Vivu organization.
  2. Call vivu_open_upload_page with the project ID immediately before uploading. The link expires in 180 seconds and works once, so never post or store it. Give it to the user to open in their own browser and select the files, or open it in a browser tool that can attach local files. Claude in Chrome accepts at most 10 MB per upload call, and recipe videos are usually larger, so most users will open the link themselves or add the files in the Vivu web app. Never split or recompress a file to fit.
  3. Poll vivu_list_videos until every file shows ready, then record each video_id in videos.csv. Match listed names to local names; Vivu replaces spaces and punctuation with underscores.

In our test run the files went through the upload page with an automated browser, so opening the link in the user's own browser was not exercised in our test run.

Done when every row of videos.csv has a video_id and every file shows ready.

Step 5: Search for amount captions

Goal: candidate windows for every amount printed on screen.

One precise search covers all the videos in the project. The query describes the kind of caption, not its words: a query that quotes the caption itself has led Vivu to attach that caption to windows that do not show it.

Field Query Mode maximum_results
amount_caption text overlaid on the video that states an ingredient amount, a cooking time or an oven temperature as a number with a unit (tablespoons, grams, cloves, minutes, degrees, centimeters); a caption without a number does not count precise 30

maximum_results 30 fits three to five recipe videos with up to about ten amount captions each; it is also the recall ceiling, so raise it for more videos. fast is not used: it returns whole files with an empty reason, and the files are already known.

Run vivu_search_videos (project_id, query, mode "precise", maximum_results 30). It returns a job ID; call vivu_get_search_results until complete is true. Each status call can wait up to 45 seconds, so a pending search is not a stalled one. Save the result to results/amount_caption.json without the result page link, and show the result page link only in the live reply: it expires after four hours, so it never goes into the card or any file.

Done when results/amount_caption.json is saved and its job id is in state.json.

Step 6: Read every caption from full frames

Goal: each window labeled kept or rejected, with the amount copied from a full frame.

  1. Make a contact sheet across the window, one tile every half second. A window often holds several captions in a row:
ffmpeg -v error -y -ss START -to END -i VIDEO_FOLDER/FILE.mp4 -vf "fps=2,scale=320:-1,tile=8x6" -frames:v 1 recipe-WORK_NAME/sheets/FILE_START.png

START and END are the window in seconds (start_ms and end_ms divided by 1000). Tile k is at START plus k/2 seconds, counting from 0 at the top left; one sheet covers 24 seconds, so make another sheet from START plus 24 for a longer window. 2. For each caption on the sheet, extract a full frame where it is fully shown and read it:

ffmpeg -v error -y -ss SECONDS -i VIDEO_FOLDER/FILE.mp4 -frames:v 1 -q:v 3 recipe-WORK_NAME/read/FILE_SECONDS.png
  1. Copy the caption exactly as written, units and symbols included (210°C/410°F, 8'-10'). Kept: a caption with a number and a unit. Rejected: a window with no amount caption on any frame; count it as a false positive. Text that is not legible is UNCERTAIN.
  2. Write every window into candidates.csv, one line per caption found in it.
Worked example from our test run

The test corpus was three public recipe videos (15.95 minutes): a salmon meal prep with amount labels in the lower left and music only during the steps, a vertical phone video of french fries with instruction text at the top and music only, and a banana bread video whose voiceover is burned in as full captions, so its amounts are both said and written. Before searching we went through every video on a contact sheet and wrote down every amount caption, the captions without numbers ("GRAB A BOWL", "SPRINKLE SALT & PEPPER") and the finished-dish shots. All files were ready about 11 minutes after the upload started.

The amount caption search returned 8 windows: 8 real, 0 false, and 6 captions missed. No caption without a number was returned, and the reason text matched the frames. The median window was 14 seconds wide, but the widest held the whole marinade sequence, several captions in a row, while the reason named only the first; the half second sheet is what turned it into separate rows. Every caption in the vertical phone video was missed although each is legible on a full frame, and the remaining misses were short caption lines in the banana bread video. Step 7 found all of them.

Done when every window has a verdict in candidates.csv and each kept caption has a full frame in read/.

Step 7: Sweep each video for captions the search missed

Goal: no amount caption left out because the search did not return it.

Recall is the weak side of the search, especially on vertical or low resolution video. Sweep every video with a contact sheet of the whole file, one tile every 2 seconds. This is local and costs no credits:

ffmpeg -v error -y -i VIDEO_FOLDER/FILE.mp4 -vf "fps=1/2,scale=400:-1,tile=6x8" recipe-WORK_NAME/sheets/FILE_sweep_%02d.png

Use scale=270:-1 and tile=8x4 for a vertical video. Sheet n (counting from 1) starts at (n - 1) times 96 seconds for the 6x8 layout, or 64 seconds for 8x4; tile k is 2 times k seconds after that. Captions shorter than 2 seconds can fall between tiles, so check the 2 seconds around each caption the search returned and around each step boundary with the Step 6 half second command. For each caption on the sweep that is not already in candidates.csv, extract a full frame with the Step 6 command, read it, and add it with the verdict "added from sweep".

In our test run the sweep sheets showed all 6 captions the search missed, and the vertical video's captions were legible on full frames at its low resolution.

Done when every video has sweep sheets and every amount caption on them is in candidates.csv.

Step 8: Find the steps that have no amount caption

Goal: a time and a check for every step in steps.csv.

Map each kept caption to its step in steps.csv by time and by what the frame shows. For each step still without a time, run one precise search with maximum_results 5 (one step happens once or twice in a video, and 5 keeps a little room). Write the query in the creator's words for that step. When the video has narration, describe what is said; when it has music only, describe what is done, which turns the same step from said to shown:

Field Query example Mode maximum_results
step_said the baker explains the oven temperature and how long to bake the loaf precise 5
step_shown the cut potato strips are lowered into hot oil in a pan to fry precise 5

Check a step_said window with vivu_get_video_summary (project_id, video_id, include_segments true, start_ms and end_ms around the window), then copy the words from the creator's captions file or a local transcript into the said column; mark it NOT VERIFIED when neither exists. Check a step_shown window on a half second contact sheet (Step 6 command): kept only when the frames show that step being done. Finding an action from the picture alone is a weak Vivu capability, so a step_shown window is never kept without that check.

In our test run the step_said search returned 1 window, 1 real, and the summary segment agreed with the burned-in captions. The step_shown search on the music-only video returned 2 windows: 1 real and 1 false, the false one being the opening montage of fries already frying. A control search for a step none of the videos contain (adding vanilla extract) returned 0 windows.

Done when every step in steps.csv has a time with a kept window, or "not found in the video".

Step 9: Pick the step photos and the finished shot

Goal: one photo per step and one finished-dish photo.

  1. Run one precise search for the finished dish:
Field Query Mode maximum_results
finished_dish a close-up of the finished dish plated or sliced, ready to serve precise 10

maximum_results 10 leaves room for two or three hero shots per video. 2. For each step and for the finished dish, pick the clearest tile on the window's half second sheet: food in focus, no motion blur, no hands covering it. Skip frames with people in them; windows often start or end on the host, and the creator decides whether faces go on the card. Save a full frame:

ffmpeg -v error -y -ss SECONDS -i VIDEO_FOLDER/FILE.mp4 -frames:v 1 -q:v 2 recipe-WORK_NAME/steps/NN.jpg

NN is the step number from steps.csv with two digits (01, 02, ...); for the finished dish save it as recipe-WORK_NAME/steps/finished.jpg instead.

In our test run the finished-dish search returned 5 windows, all 5 real, at least one in every video; some of them ended on the host at the counter, so the photo came from earlier tiles.

Done when steps/ has a photo for every step that has a time, plus finished.jpg.

Step 10: Assemble the recipe card

Goal: a card the creator can paste after checking it.

  1. Show the creator one sample row and the field mapping, and wait for a yes before writing the rest:
step,step_text,source_file,time,amount_caption,amount_source,said,said_checked,photo,video_link
2,Add the salmon and marinate,salmon.mp4,01:48,"170G/6OZ | SALMON FILLETS, SKIN ON OR OFF; MARINATE | 30 MINS - OVERNIGHT",search,,no narration,steps/02.jpg,https://www.youtube.com/watch?v=VIDEO_ID&t=108
Field Source If unavailable
step, step_text steps.csv none
source_file videos.csv none
time the kept frame, as MM:SS "not found in the video"
amount_caption copied from the full frame NO AMOUNT ON SCREEN; UNCERTAIN when illegible
amount_source search, or sweep (Step 7) none
said the creator's captions file or a local transcript NOT VERIFIED; "no narration"
photo steps/ blank
video_link the public link with &t= and the time in seconds file name and time
  1. Write recipe_card.csv and recipe_card.md. The markdown has the finished photo, an ingredient list built from the kept captions, the numbered steps with photos and time links, and a "Check before publishing" list of every NO AMOUNT ON SCREEN, UNCERTAIN and NOT VERIFIED item.
  2. Tell the creator the counts: steps with an amount, steps without, captions found only by the sweep, windows rejected.

The creator's approval of the sample row was not exercised in our test run; the card was written for the salmon video and nothing was published.

Done when recipe_card.csv has one row per step in steps.csv, recipe_card.md exists, and the creator has approved the sample row.

Compliance

  1. Use only videos the creator made or has the rights to use. Confirm this before Step 2. When the creator downloads their own published video, it is for this card only.
  2. People on camera: step photos show food and hands. Leave out frames with faces unless the creator wants them, and leave out anyone who did not agree to appear, including children; a child's face never goes on the card.
  3. The videos stay in the user's Vivu project until the user deletes them. Create the project as private. Delete a project or video only when the user asks, and confirm first.
  4. The skill writes local files only. Publishing the card to a blog, a description or social media is the creator's own action after checking it.
  5. No face or logo recognition. Brand names on packaging are copied only when they are part of an amount caption the creator wrote.

Known failure modes

Symptom Cause Fix
(observed) no window for any caption in a vertical phone video, though each caption is legible on a full frame search recall is weak on vertical, low resolution video Step 7 sweep; an empty result does not mean the video has no amount captions
(observed) one wide window holds a whole run of captions and the reason names only the first captions that follow each other are merged into one window read every tile of the half second sheet and write one row per caption
(observed) the opening montage is returned as a step: "Potato strips are lowered on a slotted spoon into a pan of hot oil to fry." the intro shows the finished food in the same pan check step windows on the sheet; drop windows from the intro
(observed) a finished-dish window ends on the host at the counter hero shots and the host's closing share one window pick the photo from earlier tiles
(observed) caption download from YouTube fails with "Too Many Requests" rate limit on caption downloads use the captions file the creator exported from their editor, or a local transcript
(observed) yt-dlp fails with "unable to download video data" the default player client was refused retry with the web_embedded player client, from the user's own machine
the reason text gives an amount the frame does not show the reason paraphrases and has attached the asked-for caption to other windows on other footage copy amounts only from a full frame
a quoted step is worded differently from what was said the reason is a paraphrase, not a transcript copy quotes from the captions file or a local transcript
upload page asks to sign in or shows an error the one time upload link expires after 180 seconds request a new link right before opening it
a file is rejected by the browser upload tool Claude in Chrome accepts at most 10 MB per upload call the user opens the upload link in their own browser or adds the file in the Vivu web app
a listed file name does not match videos.csv Vivu replaced spaces or punctuation with underscores match by the plain name from Step 2, or by length
"has not granted vivu.write" Vivu connected read only the user reconnects Vivu with write access

Connect Vivu, then add the skill.

The skill runs through the Vivu connector. Add it to Claude from the connector directory, then install the skill.