So, um, you edit the words, you know, like, and the video just follows.

Cutrova is a local, text-based video editor. Cut, caption, and publish by editing the transcript.

How it works

A finished video in three steps

1

Import and transcribe

Record screen, camera, and mic, or drag in any file. Cutrova transcribes it locally with word-level timing.

2

Edit the words

Delete text to cut video. Strike a sentence, remove every filler word in one click, tighten pauses.

3

Dress it up and publish

Add captions, titles, and b-roll. Reframe for vertical and export 16:9, 9:16, or 1:1.

The transcript is the edit

Delete words and the video skips them, gaplessly. Strike, restore, reorder scenes, and undo it all. Kill every "um" and awkward silence in one pass, and audition a cut before you commit it.

Vertical, without re-editing

One click reframes the story for 9:16. Captions stay styled and karaoke-highlighted, linted against broadcast limits, and export as SRT or VTT.

Pipeline

Renders without a window

A Scene Manifest goes in, a finished video comes out. Cutrova's render server turns manifests into 16:9 or 9:16 masters on your own hardware, no editor open.

cutrova render manifest.json --profile vertical

Also built in

The rest of the studio

Local first, by design

Transcription and rendering run on your machine. Your raw footage and your voice never have to leave it.

Captions that comply

Styled, karaoke-highlighted captions in a click, linted against FCC, BBC, and Netflix limits.

Studio Sound

Clean up noise and even out loudness. No fancy mic or treated room required.

Multitrack and podcast

Drop in everyone's recordings. Cutrova syncs them, builds one script, and follows the active speaker.

Underlord, your co-editor

"Remove the fillers, tighten pauses, and cut the boring intro."

It edits, you approve the diff.

"Most of editing spoken video is just editing words."

The idea behind Cutrova

Start with your last recording

Early access is open. Send a mail and we will get you set up.