Claude cannot watch a video. Paste a YouTube link and it guesses from the title or reads a transcript — and on a screen recording the transcript is the smallest part. Every budget field, toggle, dropdown and formula on screen is invisible to it.

I fixed that, pointed it at a 29-minute Meta ads walkthrough for law firms, and pulled 132 labelled lines of playbook out of it for about five cents.

Here is the whole thing. The repo, the setup, the two-model recipe, and the step that did the actual work.

The repo

bradautomates/claude-video — "Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude."

MIT licence, Python, 17,531 stars and 1,799 forks — checked live on the GitHub API on 22 September 2026, not read off a screenshot.

It downloads the video, extracts frames as JPEGs, pulls a timestamped transcript, and prints the frame paths so Claude reads each one as an image. By the time it answers, it has seen the screen.

1. Install it. Claude Code:

/plugin marketplace add bradautomates/claude-video
/plugin install watch@claude-video

Codex, Cursor, Copilot, Gemini CLI and 50+ other hosts:

npx skills add bradautomates/claude-video -g

First run installs yt-dlp and ffmpeg if they are missing. Captions are free. A Whisper key is only needed when a video has none.

2. Try it plain.

/watch https://youtu.be/8Ff9nY7U53Y what daily budget does he set, and where on screen?

That already beats a transcript. It is not the method. The method is the three steps below, and the third one is the one that matters.

Step 1 — index it with a cheap model

Point a cheap, fast model at the same video and ask it for timestamps and a category only. Never a value.

That constraint is the whole idea, so do not soften it. Here is the prompt, verbatim:

You are building a COVERAGE INDEX for this video, not a summary.

Scan the whole video end to end. Output ONLY a list of timestamps where something
VISUALLY READABLE appears on screen that a reader would need to see the pixels to
get right: a numeric field value, a dollar amount, a toggle/checkbox state, a
dropdown selection, a named setting, a menu label, a formula, a filename, an
entity name, or a metric column.

STRICT OUTPUT CONTRACT — violating this makes the output useless:
- One line per timestamp, format exactly: MM:SS | <category>
- <category> is ONE of: BUDGET, SETTING, TOGGLE, FORMULA, METRIC, FORM, CREATIVE,
  TARGETING, TOOL, NAMING, OTHER
- NEVER state the value, the number, the setting name, or what it says. Category only.
- WRONG: "07:18 | BUDGET | $100 a day"
- WRONG: "08:31 | TOGGLE | A/B test is off"
- RIGHT: "07:18 | BUDGET"
- RIGHT: "08:31 | TOGGLE"

Be exhaustive and evenly distributed across the FULL runtime including the final
minutes. Prefer more timestamps over fewer. No preamble, no headings, no commentary.

3. Run it. I used Gemini 3.8 Flash, which takes a YouTube URL directly. Save the prompt into this and run it:

# index.py — timestamps only, never values
import json, os, sys, urllib.request

PROMPT = """<paste the prompt above>"""

url, model = sys.argv[1], "gemini-3.8-flash"
body = {"contents": [{"parts": [{"file_data": {"file_uri": url}},
                                {"text": PROMPT}]}],
        "generationConfig": {"temperature": 0.0, "maxOutputTokens": 8000}}
req = urllib.request.Request(
    f"https://generativelanguage.googleapis.com/v1beta/models/{model}:generateContent",
    data=json.dumps(body).encode(),
    headers={"Content-Type": "application/json",
             "x-goog-api-key": os.environ["GOOGLE_API_KEY"]})
r = json.load(urllib.request.urlopen(req, timeout=900))
print("\n".join(p["text"] for p in r["candidates"][0]["content"]["parts"]))
python3 index.py "https://youtu.be/8Ff9nY7U53Y" > index.txt

It read 29 minutes in 17.7 seconds and returned 74 lines that look like this:

00:01 | TOOL
00:53 | OTHER
03:48 | BUDGET
08:31 | TOGGLE
24:26 | FORMULA

Why it must never emit a value

This is the bit people get wrong, so here it is straight.

A second model asked "what does this video say" is a second opinion, and a second opinion you cannot check is worse than none. In my first run I let it answer freely. It inverted a metric definition and mislabelled its own provenance twice.

Constraining it to timestamps does not make it more accurate. It removes its opportunity to be inaccurate. That is a design win, not a capability win. It holds whether or not you are paying attention that day.

What you get instead is an attention map, and the map was measurably good. Six of its timestamps landed on content my dense frame grid straddled, and those six yielded about eight facts I would otherwise have lost. Eleven grid frames it ignored yielded nothing — they were talking-head slides. It spent its budget on the Ads Manager and starved the filler. That is real discrimination, not noise.

I checked the contract held with awk, looking for any third field or any digit, $ or % in the category column. Zero leaks. Not one number from the second model entered the final document.

Step 2 — get the pixels, in windows

The repo's scene detector is built for cuts. A screen recording has almost none, so it under-samples badly.

My first pass used the default scene-aware mode. It returned 36 frames with a 383-second blind gap — six and a half minutes of live Ads Manager with not one frame. Across five stretches, roughly 19 of the 29 minutes were blind.

Do not reach for the detail dial to fix this. I did, and measured it: --detail token-burner returned 36 frames from 36 candidates, byte-identical timestamps to the default. Once scene detection finds eight or more shots, the detail mode, the frame cap and the fps setting are all inert. It reports a budget of 100 and hands you 36 anyway.

4. Force the uniform sampler with windows instead. Ten chaptered passes across a 29-minute video. $SKILL_DIR is wherever the skill installed — the folder holding its SKILL.md:

for i in $(seq 0 9); do
  S=$((i*174)); E=$((S+174))
  python3 "$SKILL_DIR/scripts/watch.py" video.mp4 \
    --start $S --end $E --max-frames 12 --out-dir dense/w$i
done

5. Then pull a frame at every timestamp the index gave you, at the video's native width so the on-screen text is actually legible:

python3 "$SKILL_DIR/scripts/watch.py" video.mp4 \
  --detail transcript \
  --timestamps 00:01,00:53,03:48,08:31,24:26 \
  --resolution 1114 --out-dir cues

Windows gave me 112 frames and cut the worst gap to 136s. The index cue frames filled the hole the windows left. Together: 56 seconds, worst case. Down from 383.

Step 3 — read the numbers off the pixels yourself

This is the step everyone skips, and it is the one that earns the whole exercise.

Near the end of the video the presenter builds a custom Meta metric. He names it Hook Rate, sets the format to Percentage, and enters the formula:

Video plays ÷ 3-second video plays

That is backwards. Three-second video plays are a subset of video plays, so the numerator is always the larger number. The metric can only ever return 100% or more. His own stated threshold — anything under 80% is bad — is unreachable. The video teaches an inverted metric to everyone who follows along. The correct version is 3-second video plays ÷ Video plays.

No transcript contains that. He never says the formula out loud. It exists only as text in a dialog box, and the only way to catch it is to look at the dialog box.

One rule makes step 3 work: never settle a question from a single frame. If a setting is being typed or selected, pull ±5 seconds either side before you call it.

I learned that the hard way. My first run adjudicated a language setting from one frame at 11:05, ruled the second model had invented two values, and wrote the wrong answer into the document as a confirmed finding. The frame at 11:08 shows all three values going in, exactly as the second model had said.

I had scored that run "Claude 3, Gemini 0". The real score was Claude 2, Gemini 1, no-contest 3.

So the finding underneath all of this is not "denser sampling finds more facts", though it does — the document went from 109 to 132 labelled lines, with 31 new facts and 2 corrections.

It is that sparse sampling does not just lose facts. It manufactures false confidence in the arbitration. A single frame gives you a tidy, plausible, wrong answer, stamped as verified. That is a worse failure than the one you started with, and it is invisible unless you go back and look again.

What this still does not do

  • The blind gaps never hit zero. Best case here was 56 seconds. On a fast-moving UI, things happen in 56 seconds.

  • The index removes the model's chance to be wrong. It does not remove the ambiguity. It relocates the misread risk from the model to you. On that language question, the index pointed at 11:03 — almost certainly the exact frame that would have produced the same wrong answer in reverse. I only got it right because I pulled three more frames myself.

  • Context is the real budget, not money. The run extracted 225 frames and I read 37. Frames are images and images are expensive. Point the density at the minutes that matter.

  • The cost figures are calculated, not billed. About $0.05 for the second run, about $0.20 for the first, worked out from published per-token rates against the usage metadata each call returned. Not from a statement.

  • YouTube blocks the downloader from datacentre IPs. On a VPS you will need to route the download through a residential connection. On a laptop it just works.

Run it on something you were going to watch at 2x anyway, and reply to this email with what it caught. Mine was a formula in a dialog box that nobody said out loud.

If you would rather have this kind of thing built and running for you, that is what we do at OptiMAX.

SOURCES (verified live 22 Sep 2026)

  • Repo, stars, forks, licence — GitHub API, api.github.com/repos/bradautomates/claude-video

  • The video — https://youtu.be/8Ff9nY7U53Y (PPC4, 28:52, published 9 Dec 2025)

  • Frame counts, blind gaps, line counts, index-contract check and cost — measured on my own run, 22 Sep 2026