← Help Center

Cleaning dialogue

Remove filler words without a robot voice

Every “um” you cut makes the next one more obvious. The point of this tool is not to remove all of them; it is to remove the ones that were never part of how you talk, and to leave the ones that were.

What it finds

Three kinds of thing, and they are not treated the same:

  • Pure hesitation. “Um”, “uh”, “er”, and their stretched forms: “ummm”, “uhhhh”. These are almost never meant. They are ticked by default.
  • Hesitation phrases. “I mean”, “you know”, “sort of” when they carry nothing. Ticked when the surrounding speech says they are padding.
  • Context-dependent words. “Like”, “actually”, “basically”, “literally”. These are real words at least as often as they are filler, so they are listed but never ticked for you. Removing “like” from “it looked like a bug” produces nonsense.

Step by step

  1. Open Remove filler words, then set the timeline area and voice track under Choose the source.
  2. Analyze. The first run on a machine downloads the speech model, which takes several minutes and looks like nothing is happening. Later runs start immediately.
  3. Read the list. Each entry shows the word, where it sits, and a confidence score. Anything the tool is not sure about is unticked.
  4. Untick anything you want to keep. A pause you took for effect is not a mistake.
  5. Apply. Um Out cuts picture and dialogue together, so the video track stays in sync with the audio.
Listen to one paragraph before running it on an hour. Two minutes of checking tells you more about whether the threshold suits your speech than any setting description can.

Why some are found and not ticked

Confidence is about the removal, not the transcription. A stretched “ummmm” is one of the least clearly transcribed things in any recording, and for a long time that counted against it: the tool scored it as an uncertain word and then declined to remove it, which meant the most obvious filler in the file was the one it skipped. Pure hesitations are now judged on what they are rather than on how cleanly the model heard them.

Everything else still weighs the transcription, because for a real word a bad transcription is a genuine reason not to cut.

Keeping it sounding human

  • Do not clear the list and accept all. The unticked ones are unticked for a reason.
  • Leave a breath. Um Out trims the word and a little either side, but never more than half the real gap around it, so words are not welded together.
  • Run silences first, then fillers, then review. Both passes cut the same track, and it is easier to judge a stumble once the dead air around it is gone.
If you have already added J/L cuts or fades by hand, run fillers before that, not after. A later pass cuts the underlying clip and the handles you built no longer describe the audio beneath them.

When it finds nothing

Check the source line is pointing at the track with your voice on it. A scan of a music track finds no filler words because there are none, and the tool has no way to know that is not what you meant. If the track is right and the list is still empty, you may simply not use many; try the same pass on a section you know contains one.

What it needs

The local speech model. If the panel says a transcription model is required, that is what is missing. Filler removal is included on every paid plan and the transcription runs on your own machine, so nothing about your recording is uploaded.