How it works
A finished video in three steps
Import and transcribe
Record screen, camera, and mic, or drag in any file. Cutrova transcribes it locally with word-level timing.
Edit the words
Delete text to cut video. Strike a sentence, remove every filler word in one click, tighten pauses.
Dress it up and publish
Add captions, titles, and b-roll. Reframe for vertical and export 16:9, 9:16, or 1:1.
So, um, this whole part was, you know, recorded in one, like, single take.
The transcript is the edit
Delete words and the video skips them, gaplessly. Strike, restore, reorder scenes, and undo it all. Kill every "um" and awkward silence in one pass, and audition a cut before you commit it.
Vertical, without re-editing
One click reframes the story for 9:16. Captions stay styled and karaoke-highlighted, linted against broadcast limits, and export as SRT or VTT.
Deletethewords,keepthestory.
Pipeline
Renders without a window
A Scene Manifest goes in, a finished video comes out. Cutrova's render server turns manifests into 16:9 or 9:16 masters on your own hardware, no editor open.
$ cutrova render manifest.json --profile vertical
Also built in
The rest of the studio
Local first, by design
Transcription and rendering run on your machine. Your raw footage and your voice never have to leave it.
Captions that comply
Styled, karaoke-highlighted captions in a click, linted against FCC, BBC, and Netflix limits.
Studio Sound
Clean up noise and even out loudness. No fancy mic or treated room required.
Multitrack and podcast
Drop in everyone's recordings. Cutrova syncs them, builds one script, and follows the active speaker.
Underlord, your co-editor
"Remove the fillers, tighten pauses, and cut the boring intro."
It edits, you approve the diff.
"Most of editing spoken video is just editing words."
The idea behind Cutrova
Start with your last recording
Early access is open. Send a mail and we will get you set up.