Claude Code skill · macOS · MIT

video-lens

Measures video instead of eyeballing it, so Claude can answer when, for how long and how. Everything runs on your Mac.

Claude cannot watch video. video-lens measures it frame by frame, hands Claude numbers and text first, and shows images only where a look is needed.

A 300 ms animation at 60 fps spans 18 frames. Sampling 2 frames per second catches 0 or 1 of them.

With video-lens, against the same model alone (9 tasks scored against answer keys, 27 runs per condition):

  • Claude Opus 5.5: cost per run 23% lower, time per run 37% lower, mean score 0.956 against 0.949.
  • Claude Sonnet 5.5: cost per run 35% lower, time per run 58% lower, mean score 0.940 against 0.970.
  • Grok 4.7: cost per run 27% lower, time per run 35% higher, mean score 0.942 against 0.819.

What it does

UI motion as CSS

Start, duration, easing as cubic-bezier, stagger and travel of each animated element, to the frame, written out as CSS.

Talks and demos with timestamps

Scene cuts, keyframes, Korean and English on-screen text (macOS Vision), a speech transcript and audio/video sync.

Nothing uploaded

ffmpeg, OpenCV, macOS Vision, Apple on-device speech recognition and whisper.cpp. Video and audio stay on your Mac.

Why

Compared with /watch and video-use

/watch (claude-video 0.3.2) is built for a quick look at what a video is about. video-use edits videos by conversation: cuts, colour and subtitles. video-lens is for when something happens, for how long and how.

Aspect video-lens /watch (claude-video 0.3.2) video-use (browser-use)
Frames looked atEvery framePicked at scene changes (evenly when there are none), at most 100, 512 px wide10 frames per requested range, 320 px wide, when it asks for them
Time precisionOne frame (16.7 ms at 60 fps)Whole secondsWord timestamps for speech; frames at the times it asks for
A 300 ms animationStart, duration, easing and CSS measured0 or 1 framesOnly the frames it samples; motion is not measured
SpeechTranscribed on your Mac, Korean includedCaptions first; otherwise WhisperX on your Mac (a 1.5 GB install) or uploaded to Groq or OpenAIAudio uploaded to ElevenLabs Scribe (paid API key)
Scene cuts and on-screen textCut times and the text of each slideUsed to pick frames; cut times and slide text are not reportedNot detected

Shaded: what video-lens does that the other two do not.

In the benchmark, Opus with /watch 0.3.2 matched the video-lens average but had a lower worst run, and with video-use it scored a little higher. Both mostly measured the clip with their own ffmpeg and OpenCV code on top of the skill, and cost about 1.8 and 2.4 times as much. Details: Against other video skills.

Compared with the model alone

Without the skill, Claude writes new ffmpeg and Python code for each video and measures with it. That often works, but the code differs every run.

Task "Modal enter and exit": Grok 4.7 alone had a median score of 0.875 over 3 runs and 0.000 in its worst run.

video-lens runs one tested pipeline instead, so every run starts from the same measurements.

  • Claude Opus 5.5: cost 23% lower, time 37% lower, mean score 0.956 against 0.949, worst run 0.800 against 0.167.
  • Claude Sonnet 5.5: cost 35% lower, time 58% lower, mean score 0.940 against 0.970, worst run 0.625 against 0.600.
  • Grok 4.7: cost 27% lower, time 35% higher, mean score 0.942 against 0.819, worst run 0.667 against 0.000.

Example: a 3-second toast recording

The toast was rendered from real CSS in headless Chrome, so the right answer is known.

Measure video-lens Answer key
Duration 298 ms (range 284 to 313) 300 ms
Easing cubic-bezier(0.22, 1, 0.36, 1) cubic-bezier(0.22, 1, 0.36, 1)

Benchmarks

9 tasks scored against answer keys, 3 runs each, so 27 runs per condition, measured from September 29, 2026 to October 3, 2026. Each model ran alone (writing its own ffmpeg and Python) and with video-lens, with the same prompt, no MCP servers and one run at a time. The tool and effort level each model ran with are under How we measured.

Loading the benchmark data…

Loading the benchmark data…

Loading the benchmark data…

Loading the benchmark data…

Both conditions solve most tasks, so mean scores sit close together. The worst run shows how far a single run can go wrong.

Grok 4.7's time with video-lens is rough: those runs shared the Mac with other heavy work for part of the pass, and one of them took 4.3 hours. Its median time was 446 s against 778 s alone.

Skill runs named the skill in the prompt. In a separate test where the prompt did not name it, Opus 5.5 chose it on its own in 8 of 9 tasks.

How the tasks were built and scored: How we measured and BENCHMARK.md.

When to use it

Use it for

  • Rebuilding or reviewing UI motion. Start, duration, easing and stagger as numbers and CSS, each with the range it could fall in. Questions such as "is this transition really 400 ms ease-out?" get a measured answer.
  • Long recordings. On the 10-minute Korean lecture, Claude Opus 5.5 with video-lens: cost 30% lower, time 74% lower (median of 3 runs).
  • Repeated motion. Carousels and loops are measured once as a group, with each repeat's start time.
  • Private talks and meetings. Speech is transcribed on your Mac, and slide text comes with timestamps.

You do not need it for

  • A quick gist of a short clip. The model alone handles it, and loading the skill adds a little cost.
  • Windows or Linux. video-lens runs on macOS only.
  • Speaker labels, language detection or music tempo. These are not supported.

Use case: rebuilding viral showreels

In the launch week of Claude Opus 5.5, a one-line prompt for a 15-second motion graphics showreel spread across X. We took 2 of the most-shared reels and gave each to Claude Opus 5.5 in Claude Code with video-lens and a single prompt: work out how the video moves and rebuild it as one HTML page. Below is what came back, rendered frame by frame and left as it was, scored against its original.

What video-lens changed here: not how close the rebuilds came. Without it, the same model's rebuilds looked at least as close in 2 of 2 reels (looks alike 0.828 on average, against 0.798 with it). With it, total cost was 24% lower, and 1 of 2 were faster. One run per reel: examples, not a measurement.

@stephanlivera · made with Claude Opus 5.5

Original: @stephanlivera on X, 16k likes and 1.7M views

  • Looks alike: 0.786 (two different reels from the trend: 0.344)
  • Cuts: 9 of the original's 9 hard cuts are in the rebuild, within 0.1 s
  • Words: 36 of the 39 words on screen in the original are in the rebuild
  • Run: 45 min, $9.49 at API prices, no human edits
  • Same model, no video-lens: looks alike 0.821 · cuts 9/9 · words 35/39 · 33 min, $11.43

Still off: the motion blur and RGB split are imitated with blur filters, the 3D dot shapes have no lines between the dots, and the width wobble on CLAUDE is weaker than in the original. The page asks for SF Pro and JetBrains Mono, so a browser without them shows other fonts; the video is the reference.

@ajith_io · made with Claude Opus 5.5

Original: @ajith_io on X, 3.8k likes and 546k views

  • Looks alike: 0.810 (two different reels from the trend: 0.344)
  • Cuts: 5 of the original's 5 hard cuts are in the rebuild, within 0.1 s
  • Words: 22 of the 29 words on screen in the original are in the rebuild
  • Run: 29 min, $6.21 at API prices, no human edits
  • Same model, no video-lens: looks alike 0.835 · cuts 5/5 · words 23/29 · 37 min, $9.30

Still off: the dot sphere on cream is flatter than the original's spiky one, the red square outline above the blue floor at 7.35 s is missing, and film grain and motion blur were left out. The fonts are stand-ins that ship with macOS. Most of the missed words are pieces of text bent into a circle, which the OCR reads in fragments.

How the rebuilds were scored

  • Looks alike is SSIM, a standard image similarity score, between the original and the rebuild at the same moment, 10 times a second (1 = identical). Two different reels from the same trend give the floor shown next to it.
  • Cuts are found by ffmpeg's scene detector in both clips. A cut counts when the rebuild cuts within 0.1 s of the original.
  • Words are read by the video-lens OCR in both clips, counting only dictionary words so misreadings drop out.
  • Same model, no video-lens is the same prompt without the skill: no skills loaded, video-lens hidden on disk, and a working folder where the other rebuilds cannot be seen. The model wrote its own frame extraction and rendering. Its final page is rendered and scored the same way.
  • Each video shows the original on the left, muted and at half size, for comparison, and the rebuild on the right, which is code Claude wrote; the link under it goes to the original post. Likes and views are from Revid's gallery of launch-week videos on September 28, 2026. Scripts: bench/cases.
The prompt, the same for every reel
/video-lens clip.mp4 in this folder is a 15-second motion-graphics showreel posted on X. Reverse-engineer it and rebuild it as one self-contained file, recreation.html (a 1920x1080 page, no audio), that plays the same video: the same shots in the same order at the same times, with the same layout, colors, text and motion (delays, durations, easing). So that every frame can be rendered exactly, animate only with CSS animations or the Web Animations API, all of them starting at page load (page time 0 = video time 0); no requestAnimationFrame, timers, canvas, video, or external files, fonts or libraries (use installed system fonts close to the original). When it is built, render it with the skill's declared.mjs (--viewport 1920x1080 --render 60 15.06) and copy the render to recreation.mp4 in this folder, compare it with the original using the skill, and fix the biggest differences, at most two fix rounds. Look at as many sheets and frames as the rebuild needs; nobody will answer questions, so do not ask. Finish with a table of the shots (start, end, what is on screen) for the original and for your recreation.

Install

In Claude Code:

/plugin marketplace add Junhan2/video-lens
/plugin install video-lens@video-lens

In a terminal:

brew install ffmpeg
python3 -m pip install opencv-python numpy
xcode-select --install

Required

  • macOS. Tested on macOS 26 with Apple Silicon.
  • ffmpeg.
  • Python 3.10 or newer with opencv-python and numpy. Tested with Python 3.13, OpenCV 4.12 and numpy 2.2. If pip refuses with externally-managed-environment (Homebrew's Python), add --user --break-system-packages. A missing package or an older Python prints the exact fix.
  • Xcode Command Line Tools, to build the on-screen text and speech helpers.

Optional

  • macOS 26 for on-device speech recognition (Apple SpeechTranscriber).
  • whisper-cpp and a ggml model (for example ggml-large-v3-turbo-q5_0.bin) in ~/.local/share/whisper/, to transcribe with whisper.
  • Node 24 and Google Chrome, to read a web page's declared CSS animations and compare them with the measurement.
  • yt-dlp, to analyse a video from a URL.

Usage

Ask as usual. Claude picks the skill when a question needs measuring. To call it directly, start your message with /video-lens.

Example prompts

  • Analyse the animation in this screen recording so I can rebuild it in CSS
  • List when each slide appears in this lecture and what it says
  • What was said around 12:00?
  • Summarise this YouTube talk scene by scene with screenshots

When several elements move at once, name the area, for example "only the list on the left". Claude then narrows the measurement with --roi.

Scene digest, only when you ask

Ask for a scene-by-scene summary with screenshots and Claude writes digest.md and a self-contained digest.html: one capture per scene, one or two lines on what happens, what was said, and a link to that moment. A YouTube video with chapters is split by its chapters. On a 66-minute Korean talk with 21 chapters this took about 4 minutes on the Mac (download, analysis and digest). Ordinary analyses never make one.

How it works

One command, vl.py analyze, measures the video and prints a text report of at most 6,000 characters. Claude reads that first, then a few labelled images, and single frames only when needed.

  1. Probe

    ffprobe reads the streams, the frame rate and each frame's timestamp, so variable frame rate recordings keep their real timing.

  2. Audio

    Sound activity and onsets. A quiet clip of up to 2 minutes goes to motion mode, a longer one to content mode, a short clip with sound gets both.

  3. Speech and content

    Captions or an on-device transcript (Apple SpeechTranscriber or whisper.cpp), scene cuts, keyframes and on-screen text with macOS Vision.

  4. Motion

    OpenCV finds the moving elements and tracks them frame by frame. Start, duration, easing (a named curve or cubic-bezier), stagger and travel are fitted and written as CSS.

  5. Sync

    The offset between sound and picture, from paired onsets.

  6. Report

    A text report with ranges and confidence, a timeline, and labelled contact sheets with timing maps.

Reading order for Claude: report, then timeline or text search, then the overview image, then contact sheets, then single frames.

How we measured

The 9 tasks

IDTask
v1Card stagger (3 cards, 60 fps, named curve)
v2Modal enter and exit (VFR recording)
v3Off-screen toast with a click sound
h1Overlapping 6-row list, custom curve, 30 fps, Retina
h2Colour change, overshoot badge, bottom drawer (VFR)
orbitReal 3D card carousel recording, 16.5 s (approximate answer key)
v420-second Korean lecture, 3 slides
h310-minute Korean lecture, 24 slides
gistTwo-sentence summary of a clip

Full method, per-task numbers and caveats: BENCHMARK.md. Scripts to reproduce the runs: bench/.

Limits