UI motion as CSS
Start, duration, easing as cubic-bezier, stagger and travel of each animated element, to the frame, written out as CSS.
Claude Code skill · macOS · MIT
Measures video instead of eyeballing it, so Claude can answer when, for how long and how. Everything runs on your Mac.
Claude cannot watch video. video-lens measures it frame by frame, hands Claude numbers and text first, and shows images only where a look is needed.
With video-lens, against the same model alone (9 tasks scored against answer keys, 27 runs per condition):
Start, duration, easing as cubic-bezier, stagger and travel of each animated element, to the frame, written out as CSS.
Scene cuts, keyframes, Korean and English on-screen text (macOS Vision), a speech transcript and audio/video sync.
ffmpeg, OpenCV, macOS Vision, Apple on-device speech recognition and whisper.cpp. Video and audio stay on your Mac.
/watch (claude-video 0.3.2) is built for a quick look at what a video is about. video-use edits videos by conversation: cuts, colour and subtitles. video-lens is for when something happens, for how long and how.
| Aspect | video-lens | /watch (claude-video 0.3.2) | video-use (browser-use) |
|---|---|---|---|
| Frames looked at | Every frame | Picked at scene changes (evenly when there are none), at most 100, 512 px wide | 10 frames per requested range, 320 px wide, when it asks for them |
| Time precision | One frame (16.7 ms at 60 fps) | Whole seconds | Word timestamps for speech; frames at the times it asks for |
| A 300 ms animation | Start, duration, easing and CSS measured | 0 or 1 frames | Only the frames it samples; motion is not measured |
| Speech | Transcribed on your Mac, Korean included | Captions first; otherwise WhisperX on your Mac (a 1.5 GB install) or uploaded to Groq or OpenAI | Audio uploaded to ElevenLabs Scribe (paid API key) |
| Scene cuts and on-screen text | Cut times and the text of each slide | Used to pick frames; cut times and slide text are not reported | Not detected |
Shaded: what video-lens does that the other two do not.
In the benchmark, Opus with /watch 0.3.2 matched the video-lens average but had a lower worst run, and with video-use it scored a little higher. Both mostly measured the clip with their own ffmpeg and OpenCV code on top of the skill, and cost about 1.8 and 2.4 times as much. Details: Against other video skills.
Without the skill, Claude writes new ffmpeg and Python code for each video and measures with it. That often works, but the code differs every run.
Task "Modal enter and exit": Grok 4.7 alone had a median score of 0.875 over 3 runs and 0.000 in its worst run.
video-lens runs one tested pipeline instead, so every run starts from the same measurements.
The toast was rendered from real CSS in headless Chrome, so the right answer is known.
| Measure | video-lens | Answer key |
|---|---|---|
| Duration | 298 ms (range 284 to 313) | 300 ms |
| Easing | cubic-bezier(0.22, 1, 0.36, 1) |
cubic-bezier(0.22, 1, 0.36, 1) |
9 tasks scored against answer keys, 3 runs each, so 27 runs per condition, measured from September 29, 2026 to October 3, 2026. Each model ran alone (writing its own ffmpeg and Python) and with video-lens, with the same prompt, no MCP servers and one run at a time. The tool and effort level each model ran with are under How we measured.
Loading the benchmark data…
Loading the benchmark data…
Loading the benchmark data…
Loading the benchmark data…
Both conditions solve most tasks, so mean scores sit close together. The worst run shows how far a single run can go wrong.
Grok 4.7's time with video-lens is rough: those runs shared the Mac with other heavy work for part of the pass, and one of them took 4.3 hours. Its median time was 446 s against 778 s alone.
Skill runs named the skill in the prompt. In a separate test where the prompt did not name it, Opus 5.5 chose it on its own in 8 of 9 tasks.
Claude Opus 5.5 with each video skill on the same 9 tasks and prompts, 27 runs per condition, one run at a time. Each run named its skill in the prompt, and video-lens was hidden from the runs of the other skills.
Loading the benchmark data…
How the tasks were built and scored: How we measured and BENCHMARK.md.
In the launch week of Claude Opus 5.5, a one-line prompt for a 15-second motion graphics showreel spread across X. We took 2 of the most-shared reels and gave each to Claude Opus 5.5 in Claude Code with video-lens and a single prompt: work out how the video moves and rebuild it as one HTML page. Below is what came back, rendered frame by frame and left as it was, scored against its original.
What video-lens changed here: not how close the rebuilds came. Without it, the same model's rebuilds looked at least as close in 2 of 2 reels (looks alike 0.828 on average, against 0.798 with it). With it, total cost was 24% lower, and 1 of 2 were faster. One run per reel: examples, not a measurement.
Original: @stephanlivera on X, 16k likes and 1.7M views
Still off: the motion blur and RGB split are imitated with blur filters, the 3D dot shapes have no lines between the dots, and the width wobble on CLAUDE is weaker than in the original. The page asks for SF Pro and JetBrains Mono, so a browser without them shows other fonts; the video is the reference.
Open the rebuilt page (one HTML file that plays once when it opens)
Original: @ajith_io on X, 3.8k likes and 546k views
Still off: the dot sphere on cream is flatter than the original's spiky one, the red square outline above the blue floor at 7.35 s is missing, and film grain and motion blur were left out. The fonts are stand-ins that ship with macOS. Most of the missed words are pieces of text bent into a circle, which the OCR reads in fragments.
Open the rebuilt page (one HTML file that plays once when it opens)
/video-lens clip.mp4 in this folder is a 15-second motion-graphics showreel posted on X. Reverse-engineer it and rebuild it as one self-contained file, recreation.html (a 1920x1080 page, no audio), that plays the same video: the same shots in the same order at the same times, with the same layout, colors, text and motion (delays, durations, easing). So that every frame can be rendered exactly, animate only with CSS animations or the Web Animations API, all of them starting at page load (page time 0 = video time 0); no requestAnimationFrame, timers, canvas, video, or external files, fonts or libraries (use installed system fonts close to the original). When it is built, render it with the skill's declared.mjs (--viewport 1920x1080 --render 60 15.06) and copy the render to recreation.mp4 in this folder, compare it with the original using the skill, and fix the biggest differences, at most two fix rounds. Look at as many sheets and frames as the rebuild needs; nobody will answer questions, so do not ask. Finish with a table of the shots (start, end, what is on screen) for the original and for your recreation.
In Claude Code:
/plugin marketplace add Junhan2/video-lens /plugin install video-lens@video-lens
In a terminal:
brew install ffmpeg python3 -m pip install opencv-python numpy xcode-select --install
externally-managed-environment (Homebrew's Python), add --user --break-system-packages. A missing package or an older Python prints the exact fix.ggml-large-v3-turbo-q5_0.bin) in ~/.local/share/whisper/, to transcribe with whisper.Ask as usual. Claude picks the skill when a question needs measuring. To call it directly, start your message with /video-lens.
When several elements move at once, name the area, for example "only the list on the left". Claude then narrows the measurement with --roi.
Ask for a scene-by-scene summary with screenshots and Claude writes digest.md and a self-contained digest.html: one capture per scene, one or two lines on what happens, what was said, and a link to that moment. A YouTube video with chapters is split by its chapters. On a 66-minute Korean talk with 21 chapters this took about 4 minutes on the Mac (download, analysis and digest). Ordinary analyses never make one.
One command, vl.py analyze, measures the video and prints a text report of at most 6,000 characters. Claude reads that first, then a few labelled images, and single frames only when needed.
ffprobe reads the streams, the frame rate and each frame's timestamp, so variable frame rate recordings keep their real timing.
Sound activity and onsets. A quiet clip of up to 2 minutes goes to motion mode, a longer one to content mode, a short clip with sound gets both.
Captions or an on-device transcript (Apple SpeechTranscriber or whisper.cpp), scene cuts, keyframes and on-screen text with macOS Vision.
OpenCV finds the moving elements and tracks them frame by frame. Start, duration, easing (a named curve or cubic-bezier), stagger and travel are fitted and written as CSS.
The offset between sound and picture, from paired onsets.
A text report with ranges and confidence, a timeline, and labelled contact sheets with timing maps.
Reading order for Claude: report, then timeline or text search, then the overview image, then contact sheets, then single frames.
/video-lens in Claude Code). Each prompt asks for a JSON answer in a fixed format that a script scores against the key.total_cost_usd that Claude Code reports. Time is the run duration that Claude Code reports.total_cost_usd that Claude Code reports. Time is the run duration that Claude Code reports.total_cost_usd that Grok Build CLI reports. Time is the wall-clock duration of the run.| ID | Task |
|---|---|
| v1 | Card stagger (3 cards, 60 fps, named curve) |
| v2 | Modal enter and exit (VFR recording) |
| v3 | Off-screen toast with a click sound |
| h1 | Overlapping 6-row list, custom curve, 30 fps, Retina |
| h2 | Colour change, overshoot badge, bottom drawer (VFR) |
| orbit | Real 3D card carousel recording, 16.5 s (approximate answer key) |
| v4 | 20-second Korean lecture, 3 slides |
| h3 | 10-minute Korean lecture, 24 slides |
| gist | Two-sentence summary of a clip |
Full method, per-task numbers and caveats: BENCHMARK.md. Scripts to reproduce the runs: bench/.
--roi fixes it. In the benchmark Opus 5.5 did this on its own; Sonnet 5.5 did it less often.