cd /news/artificial-intelligence/the-underrated-ai-tool-that-lets-any… · home topics artificial-intelligence article
[ARTICLE · art-74129] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

The Underrated AI Tool That Lets Any LLM Watch Videos

A developer's claude-real-video tool lets any LLM watch videos by extracting scene-aware keyframes with timestamps and a timestamped transcript, bypassing middleman models. The tool, which runs locally, survived a 2,181-video stress test from a single user, revealing bugs that shaped the project's direction.

read2 min views1 publishedJul 26, 2026

I spent this month testing the ways people actually get an LLM to "watch" a video. One of them got a 2,181-video stress test from a single user. Here is what held up.

1. Upload it to a hosted multimodal model (Gemini).

This is the shortest path and I want to be fair: for "tell me about this one clip," Gemini is genuinely good. You upload, you ask, you get an answer. The trade-offs are structural, not quality: your footage leaves your machine, results vary between runs, and you cannot hand the same evidence to a different model. When the answer looks wrong, there is nothing to check — you re-run and hope.

2. Run a describe-it-for-you pipeline (byjlw/video-analyzer, 1.5k stars).

A solid tool and an honest design: OpenCV pulls key frames, Whisper transcribes, a local vision model (Llama 3.2 11B) describes each frame, and you read the reconstructed description. The catch is that your LLM never sees the video — it sees another model's opinion of the video. Every frame the vision model misreads becomes a fact your LLM confidently repeats. You also pay for a second model in the loop.

3. Hand your LLM the evidence directly (claude-real-video).

This is the underrated one, and yes, it is mine — judge the argument, not the author. crv

turns a video into what an LLM can actually read: scene-aware keyframes with true source timestamps, a timestamped transcript with optional speaker labels, and a manifest that tells the agent how to read the folder. No middleman model. Claude (or GPT, or a local model) looks at the actual frames and cites frame_012 @ 00:03:41

. Everything runs locally.

A user ran crv over his entire photo library — 2,181 videos in four days — and mailed me the failure list, worst first. Two bugs from that report shaped my view of this whole category:

frames.json

; the same user re-verified all 60 final frames of a 22-minute lecture against the source, 60/60.Neither fix came from my test suite. They came from one user with a big real library. That is the honest state of this category: the hard bugs live where nobody's demo goes.

pip install "claude-real-video[fast]"
npx skills add HUANGCHIHHUNGLeo/claude-real-video

MIT, runs on your machine, source: https://github.com/HUANGCHIHHUNGLeo/claude-real-video

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @claude-real-video 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-underrated-ai-to…] indexed:0 read:2min 2026-07-26 ·