What Claude can and cannot see in your footage
An agent has no eyes. What it works from instead, what that makes easy, and where you still have to look yourself.
People ask whether Claude "watches" the video. It does not. Knowing exactly what it works from tells you which asks will land and which will not.
No eyes, three senses
Claude works from three things in BetterEdits.
The files. A video is three JSON files written for reading: the tracks, every clip with its file and its in and out points, every text and caption bar. Claude reads the timeline the way you read a script. It knows what is on the timeline, in what order, for how long.
The transcript. When a file has been analysed, a small document sits next to it with every spoken word and its timing, who said it, the silences, the on-screen text, the scenes and cuts. Claude reads it whole. This is where "cut the pauses" comes from: the pauses are in the transcript as numbers.
Pictures it draws itself. The app renders a frame at any time, or a contact sheet, one frame per clip with the captions and text drawn on. Claude asks for these and reads them like any image. This is how it checks a cut, notices a caption sitting under the Reels UI, or sees that a title covers a face.
That is the whole sensorium. No motion, no sound, no sense of pacing beyond the numbers.
What that makes easy
Anything the transcript describes, Claude does well and fast.
- Cut every pause over a second. Remove the false starts. Leave a little air.
- Find the moment where I say the product name and start the clip there.
- Turn a twenty-minute recording into five clips, one per topic.
- Caption it, never across a cut, one speaker per bar.
- Move the sentence with the hook to the front.
Anything structural, it also does well: lay out a hook and three points as slots, put a title over the first two seconds, set the canvas, export at a size.
What it does poorly
Anything that needs looking at motion, or having taste.
- Pick the take where I look best.
- Cut this b-roll montage to the beat so it feels good.
- Choose a font that suits the mood.
- Notice that the camera drifted out of focus at 1:40.
It can approximate some of these from numbers. Music beats are in the analysis, so a montage cut on beats is possible, and it will be on the beat. Whether it feels good is a question it cannot answer. You answer it in a second, on the contact sheet, and then tell Claude what to change by pointing at a clip by its words.
How to work with that
Give it the jobs that live in the transcript and keep the judgement calls. The efficient loop looks like this:
- Ask for the cut in words: what goes, what stays, how much air.
- Ask for a contact sheet. Look at it for five seconds.
- Point at anything wrong by what is said there, not by a time.
- Do the last touches by hand in the editor, which is why it is a real editor.
Every change lands as one undo step, so a wrong guess costs nothing.
Why we did not fake it
We could have run a model over every frame and told you Claude "sees" the video. It would be slow, expensive, and mostly wrong about the things that matter, which are the things you can see instantly. The honest version is faster: analyse once, cut by words, check with pictures, and keep a person in the loop for the part a person is good at.