Home › Guides › Captions and transcripts on no budget

Captions and transcripts on no budget

Captions carry speech plus speaker changes and meaningful sound, subtitles assume you can hear, and a transcript is a separate text document. Automatic speech recognition now does most of the work free, but it guesses at names, punctuation and crosstalk. Budget an hour of correction per finished half-hour rather than publishing the raw output.

Flow diagram of 5 steps: Fix the proper nouns first, globally, Mark every speaker change, Repair punctuation and sentence boundaries, Add the non-speech sound that carries meaning, Check reading speed and line length.

Kind Channel is new, nothing has aired, and there is no captioning desk, no transcription budget and no service here that will caption a contribution on anyone’s behalf. Accessibility is the producer’s own work, which is the honest reason this page is about free tools and a manual pass rather than about who to send the file to.

Three different things people call subtitles

The words get used interchangeably and they describe genuinely different deliverables, which matters because producing the wrong one is a common way to do the work and still leave people shut out.

Captions are written for someone who cannot hear the audio. They carry the speech, and they also carry the information the speech alone does not: who is talking when it is not obvious, and any sound that affects understanding — a door going, a phone ringing, the laughter that explains why the next line lands the way it does. Where they are described as SDH, for the deaf and hard of hearing, that is the same idea under a different label. Subtitles in the traditional sense assume you can hear perfectly well and simply do not speak the language, so they translate dialogue and leave the sound design alone. Mixing the two up produces a file that renders a hearing viewer’s translation for a deaf audience, silently dropping everything non-verbal.

A transcript is a different artefact again: a plain text document of what was said, without timing, read outside the player rather than over it. It serves people who would rather read than watch, people on a phone with no headphones, anyone scanning to find whether a piece covers the thing they need, and screen reader and braille users for whom a timed overlay is awkward and a document is not. It is also, incidentally, the only version of your programme a search engine can read in full, which is the subject of the last section here.

The other distinction worth knowing is open versus closed. Open captions are burned into the picture: always visible, impossible to switch off, and guaranteed to survive any platform that strips caption files or any viewer who never opens the settings menu. Closed captions live in a separate sidecar file, usually SRT or WebVTT, which the player renders on demand. Sidecar files are editable after publication, translatable, selectable, and readable by machines. Burned-in captions are none of those things but they always show up. For social video, where autoplay is muted and caption support is inconsistent, burned-in usually wins. For anything with a proper player behind it, a sidecar file is the better artefact, and there is nothing stopping you shipping both.

What free tools actually give you

Automatic speech recognition became genuinely good in the last few years, and the practical constraint on a small production is no longer money. It is the correction pass, and no tool removes that.

The single biggest lever on caption quality is not the software — it is the audio going into it. Recognition accuracy falls off a cliff with room echo, two people talking over each other, a distant microphone or heavy background noise. An hour spent getting clean sound at the recording stage saves more caption correction time than any change of tool, and it improves the programme for everyone listening as well.

Free and low-cost routes to a caption file, and what each one costs you

RouteWhat it givesThe catch
Platform auto-captions on the large video and audio servicesA rough transcript and timings for nothing, generated within minutes of uploadQuality varies by accent and subject, the editing interface is usually poor, and the result is tied to that platform. Always export the file so you own a copy rather than leaving your only captions inside somebody else’s product.
Whisper and the open speech recognition models built on itStrong accuracy, running locally on your own machine, producing SRT or WebVTT directly, with no upload and no per-minute feeIt needs a reasonably capable computer and a willingness to use a command line or a wrapper app. It also invents plausible text during silence and music, so long gaps need checking rather than trusting.
Caption features inside editing softwareTranscription in the same place you are already working, with captions that follow your cuts automaticallyTied to a subscription you may not keep, and the styling defaults are usually wrong — oversized, centred and obscuring faces. Check how the file exports before building a workflow on it.
Paid human transcription servicesHigh accuracy, correct names and punctuation, and no work from youPriced per minute of audio, which is modest for one piece and is a real cost at any regular cadence. Worth it for a piece with dense specialist vocabulary or heavy crosstalk; hard to justify as a routine.
Typing it yourself from scratchComplete control, and a genuinely close listen to your own editRoughly four to six times the runtime for most people, so a twenty-minute piece is an afternoon. Correcting good machine output takes one to two times the runtime instead, which is why almost nobody starts from a blank page any more.

The correction pass, in the order that saves time

Raw recognition output is a draft, and publishing it unchecked is how a programme ends up with a contributor’s name spelled four different ways. The failures are predictable enough that a fixed order of work removes most of them quickly. On a clean twenty-minute interview this is about half an hour; on difficult audio it is longer than the runtime, and that is worth knowing before you promise anything.

Fix the proper nouns first, globally

Names of people, places, organisations and any specialist vocabulary are what recognition gets wrong most, and it gets them wrong consistently. Write the correct spellings in a list before you start, then find and replace each one across the whole file. Doing this first means you are not fixing the same surname thirty times while also reading for sense, and consistency here matters more than almost anything else — a misspelled contributor name is the error people notice and remember.

Mark every speaker change

Machine output usually runs two voices together as one undifferentiated block, which is precisely the information a deaf viewer cannot recover from the picture alone. The conventions are simple: a hyphen at the start of each line in a two-line caption where speakers alternate, or the speaker name in capitals followed by a colon when a new voice arrives and the face is not on screen. Off-screen narration and phone audio need labelling too.

Repair punctuation and sentence boundaries

Recognition punctuates by guessing at pauses, so it produces run-on sentences, commas where full stops belong, and questions that end in a full stop. This is not cosmetic — punctuation is what carries intonation for a reader who cannot hear it, and a question mark is frequently the only thing distinguishing a challenge from an agreement.

Add the non-speech sound that carries meaning

In square brackets, sparingly, and only where the sound does work the picture does not: [laughter], [applause], [door slams], [inaudible] where a word genuinely cannot be recovered. The test is whether a hearing viewer gets something from that sound that a deaf viewer would otherwise miss. Captioning every ambient noise is as unhelpful as captioning none, because it buries the speech in clutter.

Check reading speed and line length

A caption that is technically accurate and gone before it can be read is not accessible. Published broadcast subtitle guidance works to a maximum around 160 to 180 words per minute for adult audiences, which streaming style guides usually express as roughly 17 to 20 characters per second. Keep to two lines, around 37 to 42 characters each, and give every caption at least a second on screen. Where speech outruns those limits, condense rather than truncate: keep the meaning and drop the filler words nobody misses.

Break lines where the sense breaks

Line breaks that split a phrase across two lines make captions measurably harder to read. Break after punctuation, or between clauses, rather than at whatever character count the tool chose. Keep an article with its noun and a preposition with its phrase. This is the step people skip and the one a reader feels most directly.

Watch it through once, muted

The final check is the only one that finds timing drift, captions sitting over burned-in text or a face, and the caption that appears half a second before the line is spoken and gives away a punchline. Muted is the point — it puts you in the position of the person the file is for, and it takes one runtime to do.

Where captions are required, and where they are only right

Obligations here depend on what you are, where you are, and who your work is for, and the honest summary is that a small independent producer online usually sits outside the formal requirements while having every practical reason to caption anyway.

Licensed television broadcasters in the UK work to quotas set by the regulator, expressed as percentages of output that must carry subtitles, signing and audio description, and those rise with the size of the service. Public sector bodies across the UK and the EU have had web accessibility obligations for years, which cover video published on their sites. The European Accessibility Act extended obligations to a range of private-sector services from mid-2025. In the United States, the ADA and Section 508 bite on public entities, places of public accommodation and federal agencies and their contractors. General equality law, including the duty to make reasonable adjustments under the Equality Act in the UK, can apply to a service in ways that are not specific to broadcasting at all.

None of that is legal advice, and the boundaries are exactly where anyone would want a lawyer rather than a guide. The practical position for a community producer is simpler: you are probably not under a quota, and the reason to caption is the obvious one. Around one in six adults in the UK has some degree of hearing loss, and the audience that uses captions is far larger than that, because it includes everyone watching muted in public, everyone in a noisy room, non-native speakers, people with auditory processing differences and a large number of viewers who simply prefer them. Captions are one of the few accessibility measures whose primary beneficiaries are not the group they were designed for.

One honest caveat, because overclaiming here is common. Captions serve deaf and hard-of-hearing audiences and they are not the whole of accessibility. Audio description, which narrates what is visible for blind and partially sighted viewers, is a separate piece of work and a harder one to do well on no budget. Sign language interpretation is separate again and cannot be improvised. A piece with good captions and no audio description is meaningfully more accessible than it was, and it is not accessible to everyone, and saying so plainly is better than implying a tick in a box.

The transcript is the version search engines can read

There is a discovery argument for transcripts that sits alongside the accessibility one, and it is worth understanding properly rather than as a keyword trick.

Search engines index text. A video or an audio file is, to a crawler, largely opaque: it has a title, a description, some structured data if you provided it, and otherwise a wall of content it cannot read. A published transcript turns a thirty-minute programme into several thousand words of indexable text covering every subject the conversation actually touched, in the words your contributors actually used, which are frequently not the words you would have chosen for a title. That is the whole of the mechanism. It is also why the transcript should be real text on the page rather than a downloadable file or an image, and why it should sit alongside the player rather than replacing it.

Two things to avoid, both of which are the difference between helping a reader and gaming a ranking. The first is publishing a raw, uncorrected transcript as page content — it reads as broken text, and a page of garbled machine output helps nobody and does not deserve to rank. The second is stuffing the transcript or the caption file with terms that were not said. Both are the sort of thing search guidelines describe as content made for search engines rather than for people, and both are trivially detectable. A clean transcript, honestly headed, with speaker names and paragraph breaks, is the version that works for a reader and for a crawler at the same time, and there is no version that works for one and not the other.

Practical finishing touches, none of which cost anything: give the transcript its own heading structure so it can be skimmed, timestamp it every few minutes so someone can find the part they want in the player, and name the caption file after the piece rather than leaving it as a string of digits from the exporter. If the piece is published with structured data, the transcript and caption file can be declared there too, which is a small amount of work that makes the accessibility of the piece machine-readable rather than merely present.

Are automatic captions good enough to publish as they are?

Not without a correction pass. Modern speech recognition on clean audio gets most words right, and the errors it makes are concentrated in exactly the places that matter: proper nouns, specialist vocabulary, punctuation, and any moment where two people talk at once. It also produces no speaker labels and no indication of meaningful non-speech sound, which are the parts of a caption a deaf viewer cannot recover from the picture. Accuracy also drops sharply with accents, distant microphones and room echo. Budget roughly one to two times the runtime for correcting machine output on reasonable audio, and treat the raw export as a first draft rather than a deliverable.

What is the difference between captions and subtitles?

Captions are written for someone who cannot hear the audio, so alongside the speech they carry speaker identification and any sound that affects understanding — a knock at the door, laughter, a phone ringing off screen. Subtitles in the traditional sense are written for someone who hears the audio perfectly but does not speak the language, so they translate the dialogue and leave sound effects out entirely. The usage is muddled in practice, particularly in British English where subtitles often means both, and the label SDH signals captions explicitly. The distinction matters at production time: shipping translation-style subtitles as your accessibility provision silently removes everything non-verbal.

Should captions be burned into the picture or supplied as a separate file?

It depends where the piece will be seen, and shipping both is often the right answer. Burned-in captions are part of the image, so they always appear regardless of the player, the platform or the viewer’s settings, which makes them the safer choice for social video watched muted on autoplay. They also cannot be switched off, translated, restyled, resized or read by a machine, and fixing an error means re-exporting the whole video. A sidecar file in SRT or WebVTT can be corrected after publication, offered in several languages, adjusted by viewers who need larger text, and read by search engines. Where a proper player is involved, the sidecar file is the better artefact.

Does publishing a transcript help a programme get found?

Yes, and the mechanism is simply that search engines index text and cannot hear audio. A transcript converts a programme into several thousand words of readable content covering everything the conversation touched, in the contributors’ own phrasing, which often surfaces subjects a title and description never mention. Two conditions attach. It has to be corrected, because publishing raw recognition output puts broken text on the page and helps no one. And it has to be honest, because adding terms that were never said is the kind of search-first content that guidelines specifically describe as made for crawlers rather than people. A clean, accurate transcript is the version that serves both.

What community broadcasting isHow programming decisions get madeBroadcast formats explainedRecording an interview with a phone
Stories That Prove Humanity Is Still Good — A Curated CollectionWhy the World Needs More Good News — And What Happens When We Get ItHow Communities Become Stronger Together — The Evidence and the Method
HomeExplore ideasSubmit an ideaGuidesWho it is forCompare platformsFAQWhy Kind Channel existsAbout