“Verbatim” sounds like it should mean one thing. In practice it names a spectrum, and the transcript you actually want for a sermon sits in a specific place on it.
The fastest way to understand the difference is to see it, so this page starts with the same forty seconds of preaching rendered four ways. Then it covers which style to use for archives, captions, blog posts, and books — the four things churches actually do with a transcript, each of which wants a different answer.
A preacher opening a message on Romans 8. Each panel below is the same speech; the only changes between them are removals and rewrites, never additions.
PASTOR: So — um, so turn with me, if you would, turn with me to Romans, uh, Romans chapter eight. Romans chapter eight, verse twenty-eight. [pages turning] And I want you to — I want you to notice something here that I think we, we kind of skip over. We, we read this verse at funerals, right? [laughter] We put it on, on the little card. And, uh, and that's fine, that's — there's nothing wrong with that. But I don't think — I don't think Paul is writing a funeral card here. I think he's writing to people who are, who are actually suffering right now. Right now. Not, not someday.
Fillers, false starts, repeated words, page-turning noise, and audience reaction are all recorded. Nothing is smoothed.
PASTOR: So turn with me, if you would, to Romans chapter eight, verse twenty-eight. And I want you to notice something here that I think we kind of skip over. We read this verse at funerals, right? [laughter] We put it on the little card. And that's fine, there's nothing wrong with that. But I don't think Paul is writing a funeral card here. I think he's writing to people who are actually suffering right now. Right now. Not someday.
Fillers and stammers are gone. Every remaining word is the preacher's own, in his own order. Note that the deliberate repetition of "Right now" survives — it is rhetoric, not stumbling, and cutting it would change the delivery.
Turn with me to Romans 8:28. I want you to notice something here that I think we tend to skip over. We read this verse at funerals. We put it on the little card, and there is nothing wrong with that — but I do not think Paul is writing a funeral card here. I think he is writing to people who are actually suffering right now. Right now. Not someday.
Speaker labels and stage directions drop away. The spoken reference becomes a written citation (Romans 8:28). Contractions and sentence boundaries are normalised for the eye rather than the ear. The argument and the voice are unchanged.
We reach for Romans 8:28 at funerals. We print it on the card, and there is nothing wrong with that. But Paul was not writing a funeral card. He was writing to people who were suffering as they read it — not someday, but that morning.
This is no longer a transcript. It is writing derived from a transcript: reordered, compressed, and recast in the third person. Useful and legitimate — but do not label it a transcript, and do not use it as your caption track.
The first three panels are all transcripts — records of what a person said, differing only in how much debris was swept away. The fourth is not. Once you reorder sentences and change the person and tense, you have written something new using the sermon as a source. That is a perfectly good thing to publish, and it is the right form for a blog post. It just should not be labelled a transcript, and it should never become your caption track.
Every utterance, exactly as produced: fillers, stutters, false starts, repeated words, and bracketed non-speech events. Usually timestamped and speaker-labelled.
Use it when how something was said is itself evidence — depositions, disciplinary matters, oral history where speech patterns are part of the record, and linguistic or qualitative research where hesitation is data.
Every word is still the speaker's, in the speaker's order. What goes is the noise: fillers, stammers, and sentences abandoned before they carried meaning.
Use it when the transcript is a record meant to be read — which covers nearly every sermon archive, podcast transcript, and interview.
Clean verbatim plus light copy-editing for the eye: speaker labels removed, spoken references converted to written citations, sentence boundaries normalised, obvious slips silently corrected.
Use it when the transcript is the reading experience — a sermon archive page a visitor will actually read start to finish rather than search.
Derived writing: reordered, compressed, often recast out of the second person. The sermon is the source material rather than the content.
Use it when you are publishing a blog post, a devotional, a newsletter piece, or a book chapter. See turning a sermon into a blog post and turning a sermon series into a book.
| What you are producing | Style | Why |
|---|---|---|
| Searchable sermon archive | Clean verbatim | Search needs the preacher’s actual words. Editing them out makes the archive miss the phrase someone half-remembers. |
| Sermon transcript page on the website | Clean verbatim or clean read | Either works. Clean read if visitors read it through; clean verbatim if they search it. |
| Closed captions / subtitles | Clean verbatim, closely tracking speech | Captions stand in for the audio itself. Rewriting gives deaf viewers a different sermon. |
| Blog post or newsletter | Edited prose | Spoken structure does not survive on the page. Rewrite, and do not call it a transcript. |
| Book or published collection | Edited prose, heavily | A chapter needs a written argument, not a recording of a spoken one. |
| Discussion or small-group guide | Edited prose | You are extracting questions and points, not reproducing speech. |
| Legal, disciplinary, or oral-history record | True verbatim | Hesitation, repetition and exact phrasing may all matter later. |
| Translation source text | Clean verbatim | Translators need complete, unedited meaning; fillers only add noise. |
Most churches need exactly two artefacts per sermon: one clean verbatim transcript that serves the archive, the search index, and the caption track, and one edited prose piece for whatever gets published. Producing more than two is usually a sign the workflow has drifted.
This is the one place where the choice carries a compliance dimension rather than just a stylistic one, so it is worth stating carefully.
Accessibility guidance treats captions as an equivalent to the audio for someone who cannot hear it. The working principle is that a deaf or hard-of-hearing viewer should receive the same message as everyone else — which means captions follow the speech as delivered. Dropping a meaningless stutter is fine and standard. Silently rewriting the preacher's sentences into tidier ones is not, because the hearing congregation did not get the tidier version.
The practical failure mode is a church that produces a polished blog version of the sermon and then, quite reasonably, thinks to reuse that text as the caption file. It reads beautifully and it is the wrong document. Caption from the clean verbatim transcript instead; keep the polished version for the blog.
Worth noting that no accessibility standard sets a numeric accuracy threshold you can point at — the requirement is effective communication, judged in context. That is a lower bar than perfection and a higher one than “good enough to follow.” More on this in sermon accessibility.
Two of the three steps are mechanical, and one is not.
Strip the fillers.
Remove "um", "uh", "you know", "I mean", and collapse stuttered repeats ("we, we, we read this"). Purely mechanical — our free transcript cleaner does this in one pass.
Repair the false starts.
This is the judgment step and it cannot be automated. When a preacher begins "And I want you to — I want you to notice", the abandoned fragment is noise. But when he begins a thought, breaks off to tell a story, and returns to it four minutes later, that structure is the sermon. A tool cannot tell those apart; you can, in one read.
Decide what stays as rhetoric.
Repetition from the pulpit is frequently deliberate. "Right now. Not someday." is not a stammer, and a cleaner that treats every repeat as an error will flatten the preaching. Read the result aloud once — anything that has lost its cadence was doing rhetorical work.
The free transcript cleaner handles step one on text you paste in — no account, nothing uploaded. Steps two and three are a single read-through, and on a typical sermon they take a few minutes.
True verbatim means every sound a speaker makes is written down: filler words (um, uh), false starts, stutters, repeated words, and non-speech events like coughs or laughter, usually with timestamps and speaker labels. Nothing is tidied. It is the standard for legal depositions, research interviews, and any setting where how something was said carries as much weight as what was said. It is almost never what a church wants for a sermon transcript.
True verbatim keeps every utterance, including "um", "uh", stutters and false starts. Clean verbatim (also called intelligent verbatim) keeps the speaker's exact words and sentence structure but removes the noise: fillers, stammers, and abandoned half-sentences. The wording stays the speaker's; only the debris is dropped. Clean verbatim is the default for sermons, podcasts, and interviews meant to be read.
Clean verbatim for almost everything: website transcripts, sermon archives, and search. Use a lightly edited "clean read" version when the transcript becomes a blog post or book chapter. Use true verbatim only when exact speech matters — a disciplinary or legal matter, an oral history where a person's speech patterns are part of the record, or linguistic research. Captions are a separate case governed by different rules.
Captions should closely follow what is actually said. Accessibility guidance treats captions as an equivalent to the audio for viewers who cannot hear it, so silently rewriting the speaker gives deaf and hard-of-hearing viewers a different sermon than everyone else. Standard practice is to caption the speech as spoken while dropping non-meaningful stutters, and to mark significant non-speech sound in brackets. Do not use a heavily edited prose version as your caption track.
Most AI transcription, including ours, produces something close to clean verbatim by default. The models are trained on written-style text, so they tend to drop the most obvious fillers and repair small stumbles on their own. If you need true verbatim with every "um" preserved, that is a deliberate setting rather than the default, and AI is generally weaker at it than a human transcriptionist.
Expect meaningful shrinkage, though the exact amount depends entirely on the speaker. A preacher who speaks in tight prepared sentences may lose very little; an extemporaneous preacher who circles a point three times before landing it can lose a great deal. The honest way to find out is to clean one of your own transcripts and compare the word counts, since the number is a property of your preacher, not of transcription in general.
For an ordinary sermon transcript on a church website, no — readers assume light cleanup. It matters when you quote someone, particularly a guest speaker: an edited quotation presented as exact speech is a real problem if the edit changed emphasis or removed a qualification. If a transcript will be cited, keep the clean verbatim version as the record and note that the published version is edited.
Partly. Removing filler words and collapsing repeated words is mechanical, and our free transcript cleaner does it in one pass. What cannot be automated is the judgment call: deciding whether an abandoned sentence was a genuine false start or the beginning of a thought the preacher returned to. Run the tool first, then read for sense.
Our transcripts come out close to clean verbatim already — fillers largely gone, the preacher's words intact, scripture references formatted as citations rather than spelled-out numbers. $19/month for unlimited sermons, and your first full sermon is free.