Writing
Use AI to catch filler words: a practical coaching prompt
updated 2026-09-17
To use AI as a filler-word coach, give it a short list of patterns, clear exclusions, and a speech sample. Ask for counts and examples. Then check the questionable flags against the recording. A transcript alone can miss fillers or change the punctuation.
I talk a lot. Calls, recordings, thinking out loud with agents. I wanted a way to notice the habits I miss while speaking. This is the checklist I use.

Which patterns should the agent flag?
These are my targets. They are not a universal list of words that good speakers must avoid.
| Pattern | Count it when | Leave it alone when |
|---|---|---|
| Um or uh | It appears as a filled pause | The speaker is quoting or discussing the word |
| You know | It repeats as a habit without adding meaning | It means something, as in “you know the client” |
| Opening so | It repeatedly starts a new thought without a clear role | It connects a result to what came before |
| Restarts and repeated words | I restart a phrase I want to practice | The repetition adds emphasis or the transcript is uncertain |
| Trailing right? | It becomes a repeated rhetorical tag | It is a genuine request for confirmation |
My watch list also includes opening “okay,” mid-sentence “right,” and “etcetera” used to abandon a list. I keep those separate until repeated examples make them worth practicing.
All speakers have moments of disfluency. Filler words, revisions, and pauses are not automatically problems. Stuttering is also not a bad habit that this checklist can diagnose or fix. ASHA explains these distinctions in its fluency overview.
How do you avoid transcript false positives?
Keep the original audio when you can. A cleaned transcript is not a complete record of delivery. For example, AssemblyAI documents an option to retain filler words. Check what your own transcription tool preserves.
Plain text cannot establish the length of a pause. Do not ask the agent to infer timing from an ellipsis. Use audio or timestamps for that.
The opening “so” punctuation test
A transcript might show:
We need a shorter workflow for this user. So I do not want to begin with inbox triage.
Join the clauses:
We need a shorter workflow for this user, so I do not want to begin with inbox triage.
That makes a causal reading plausible. It does not prove how the speaker used the word. My rule is to mark the case as uncertain, listen if audio is available, and avoid counting doubtful examples as confirmed fillers.
Context beats capitalization. A period inserted by software should not make the coaching decision for you.
What prompt should you use?
Paste this with a sample. Adjust the targets to the speaker's own goals.
Review this speech sample for my practice targets:
um/uh, habitual “you know,” habitual opening “so,”
phrase restarts, and habitual trailing “right?”
Watch list: opening “okay,” mid-sentence “right,”
and trailing “etcetera.” Keep these separate.
Do not rewrite my personality or diagnose a speech disorder.
Exclude quoted examples and discussion of these terms.
Check meaning before classifying a word as a filler.
Mark ambiguous cases as uncertain; do not force a count.
Do not infer pause length or timestamps from plain text.
Return a table: pattern, confirmed count, uncertain count,
one short example, and timestamp if supplied.
Then suggest one practice goal. Keep the feedback brief.
Review the first few outputs yourself. If the agent repeatedly counts a meaningful use, add that example to the exclusions. You are refining the instructions, not fine-tuning a model.
How should you track progress?
Practice one pattern at a time so you can still focus on the idea you are explaining.
Compare similar samples: the same approximate length, setting, and kind of task. A rehearsed presentation and a difficult brainstorming call are different tests.
| Record | Why it helps |
|---|---|
| Sample length and setting | Gives the count context |
| Confirmed and uncertain counts | Keeps guesses out of the score |
| One audio example | Makes the feedback easy to check |
| One practice goal | Keeps the next attempt manageable |
If you compare recordings of different lengths, a count per minute can help. Use the actual audio duration. Still listen for clarity and comfort; a lower number is not the only useful outcome.
Before uploading a call, make sure you have permission to use it and understand the tool's data settings. A short solo practice recording is an easy place to start.
Frequently asked questions
Does “you know” always count as a filler?
No. It can carry meaning or help a conversation. Count only the uses that match the practice goal, and leave doubtful cases marked as uncertain.
Can AI count filler words accurately from a transcript?
It can help find candidates. Accuracy depends on what the transcript preserves and whether the agent reads the context correctly. Check a sample against the audio before trusting totals.
Do I need to train a custom model?
No. Start with instructions, examples, and a repeatable output format. In this workflow, “training the agent” means teaching it your checklist through context.
Should I aim for zero fillers?
That is not my goal. I want to communicate clearly and notice habits that distract me. Natural pauses and informal language still belong in a conversation.
Can I use this with someone else?
Yes, if they want that feedback. Agree on the targets and recording permissions first. Treat the output as practice feedback, not a judgment of the person or a clinical assessment.
Turn your reviewed checklist into a reusable agent skill with the Life = Content loop.