Learning Center Dialogue

Remove filler words without losing the speaker’s voice

Remove clear hesitation sounds first, leave context-dependent words unticked, and review each edit in the sentence around it.

7 min read · reviewed September 9, 2026

Filler words are the most dangerous thing to automate in a spoken edit, because the tooling is good enough to remove all of them and removing all of them is wrong. "Um" and "uh" are usually noise. "Like", "you know", "I mean" and "actually" are usually doing something, and a video with every one of them stripped out sounds like a person reading a transcript of themselves.

The goal is not zero filler. It is removing the ones a viewer would notice while leaving the ones that make the speech sound like speech.

Two categories, two policies

Pure hesitation sounds (um, uh, er, ah) carry no meaning. They are the sound of a brain buffering, and removing them is almost always an improvement. These are the safe category, and they are where the time savings are: a forty-minute recording can easily contain several hundred.

Real words used as filler are different. "Like" can be hedging, quoting, or approximating. "You know" can be checking that the listener is with you. "Actually" often signals a correction. Removing these by category changes meaning and flattens personality, which is why Um Out flags them separately rather than folding them in with the ums.

Read the sentence, not the word

Every candidate should be judged in the sentence around it. The same word in the same speaker can be disposable in one line and load-bearing in the next. "It was, like, four hours" is an approximation and the word is doing work. "It was like, um, really bad" is not.

This is slower than approving a list, and it is the part that matters. In practice it is quick: most candidates are obvious in the second it takes to read the line, and only a handful need actual thought.

Do filler cleanup after the silence pass

Order matters more than people expect. Silence removal changes the rhythm of the whole recording, and rhythm is what makes a stumble feel distracting or invisible. A hesitation that stood out in the raw file, surrounded by four seconds of dead air on either side, often disappears entirely once the dead air is gone.

Run silences first, watch the result, then run the filler pass on what remains. You will approve fewer removals and the edit will sound better for it.

Watch what happens to the join

Removing a word from the middle of a sentence leaves a join, and joins have the same problems as any other cut: clipped consonants, abrupt room tone, two words running together with no natural space. A filler removal that reads fine on the timeline can sound like a glitch on playback.

This is more noticeable with filler than with silence, because the surrounding speech is continuous: there is no pause to hide the seam in. Listen to each one in context rather than trusting the list.

Leave some in on purpose

A completely filler-free delivery reads as scripted, which is fine for a product video and wrong for a conversation, a tutorial, or anything where the appeal is that a real person is talking. If the goal is that someone sounds like themselves, a handful of surviving hesitations is not a failure of the pass. It is the pass working.

Common mistakes

Deleting every instance of "like"

It is a quotative, an approximator and a hedge as often as it is filler. Category deletion changes what sentences mean.

Running filler cleanup first

Do it after the silence pass. The tighter rhythm makes many stumbles stop mattering, so you remove less and the result sounds more natural.

Approving from the list without listening

The join is the risk, and the join is only audible on playback.

Aiming for zero

Perfectly compressed speech is not the same as natural speech, and viewers can tell the difference even when they cannot name it.

Questions

Which words does it detect?

Clear hesitation sounds are flagged as the safe category. Context-dependent words are flagged separately so you can decide them individually rather than by rule, because the right answer genuinely differs sentence to sentence.

Will removing fillers make the audio sound choppy?

It can if the padding either side of each removal is too tight, for the same reason silence removal can. Review the joins on playback; if a seam is audible, widen the margin or leave that one in.

Does it work on accented or non-native speech?

Hesitation sounds vary between speakers and languages, and detection is less reliable the further the speech is from what the model was trained on. Review is not optional here. It is the mechanism that makes the pass safe when detection is imperfect.

Can I keep a list of words to always ignore?

Your review decisions are per candidate rather than global, which is deliberate: a blanket always-ignore rule reintroduces exactly the category thinking that causes the damage.

The tools that do this