Bad sound costs an audience faster than a bad picture because speech carries the meaning and a degraded voice forces the listener to work for it. Four numbers describe any recording: the noise floor beneath it, the peaks, the average loudness in LUFS, and the true peak. Clipping and room reverb are permanent; level is not.
Flow diagram of 5 steps: Level: fixable, and not worth worrying about, Clipping: permanent, Room reverb: effectively permanent, Steady broadband noise: partly fixable, at a cost, Intermittent noise: usually a retake.
Ask anyone who has published a first piece what the feedback was and it is rarely about the picture. Nobody writes in to say the white balance drifted. They say they could not hear it, or they say nothing at all and close the tab, which is the version you find in the retention graph rather than the comments. The asymmetry is large enough to be worth designing around: an audience will accept a picture that is soft, noisy, slightly green and framed by an amateur, and will abandon a voice that is muffled, echoey or fighting a fridge, usually within the first half-minute.
The reason is that the two channels are not doing the same job. In a talking-heads piece the picture is confirmation — it tells you who is speaking and roughly where they are, and it can lose a great deal of quality while still doing that. The audio carries the entire argument. Degrade the picture by thirty per cent and you still know what was said; degrade the audio by thirty per cent and you have lost the content itself, not the presentation of it.
Underneath that there is a mechanism worth knowing, because it explains why listeners leave without being able to articulate why. Intelligibility in speech lives mostly in the consonants, and consonants are quiet, brief and concentrated in the upper mid-range around two to four kilohertz. Vowels are loud and low and survive almost anything. So a recording swamped by room reflections or a broadband hiss loses its consonants first while still sounding, on a casual listen, like a voice at a reasonable volume. The audience is not consciously hearing a technical fault. They are doing reconstruction work on every sentence, and after thirty seconds of that they are tired and they stop.
This effect has a name in the hearing research literature — listening effort — and the finding that matters for producers is that comprehension and effort are separable. People can score perfectly on what was said and still be exhausted by having said it back to themselves, and fatigue is what drives the decision to leave. It is also why a recording that seemed fine in the room can fail in delivery. In the room your two ears and your brain are performing spatial separation the microphone cannot do, discarding reflections and locking onto the person in front of you. A single microphone hands the audience the unseparated sum and asks them to do the job with one channel.
There is a credibility cost as well, which is less intuitive and well evidenced. Newman and Schwarz published work in Science Communication in 2018 in which listeners heard the same conference talk at different audio qualities, and the group that heard the degraded version rated the research itself as worse and the researcher as less competent. The content was identical. If that holds for a scientist presenting findings, it holds for a community contributor presenting a case for something they care about: thin, hollow audio is read by an audience as evidence about the speaker rather than about the microphone.
Add to that where the piece will actually be heard. A substantial share of listening happens on phone speakers, in cars, on cheap earbuds, on public transport, and while the listener is doing something else. Each of those adds noise and removes low frequencies, so anything already marginal in headphones becomes unusable there. Checking a mix on a phone speaker before publishing is not a courtesy to the minority; for most first productions it is the majority listening condition.
Most audio confusion among first-time producers comes from collapsing several different measurements into the single word volume. They are not the same thing, they are fixed at different stages, and only some of them are still adjustable once the recording exists. Four numbers cover almost everything you need.
The noise floor is the level of everything present when nobody is talking — preamp hiss, ventilation, traffic, computer fans, the electrical hum of the building. It is the number set entirely at the recording stage, and it is the one that decides how much of the rest you can do. What the ear actually reads as clean is the gap between voice and floor, not the absolute level of either. Roughly fifteen decibels of separation is comfortable for a listener with normal hearing in a quiet place; a listener on a bus, or anyone with hearing loss, needs appreciably more, which is the practical argument for getting the microphone close rather than turning the gain up.
The peak level, measured in dBFS, is how close the loudest instant comes to the top of the digital scale. Zero dBFS is a hard ceiling rather than a suggestion. Go past it and the waveform is squared off, which is heard as a harsh crackle and cannot be undone, because the information above the ceiling was never written to the file. Aim peaks at −12 to −18 dBFS while recording. That looks alarmingly low on a meter, and it is correct: a laugh, a cough or a sudden emphatic answer can be ten decibels above a normal speaking level, and that headroom is what absorbs it.
Average loudness, measured in LUFS, is what an audience experiences as loudness across a whole programme, and it is the number that matters at delivery. It is a time-averaged, frequency-weighted measurement built to match perception, which is why two files can peak identically and still sound obviously different in level. A file with sparse loud transients and a quiet body peaks high and measures low; a dense, consistently level voice track peaks the same and measures much louder. Peak meters cannot tell you this, and eyeballing a waveform cannot either.
True peak, in dBTP, is the fourth and the one people meet only when something breaks. It anticipates the peaks that appear after a file is converted or compressed for delivery, which can be higher than anything in the original. This is why delivery specifications ask for a ceiling of −1 dBTP rather than zero: the margin exists so that encoding to AAC or MP3 does not produce distortion that was not in the file you approved.
The four measurements, when each is fixed, and what to aim for
| Measurement | What it describes | When it is decided | Target |
|---|---|---|---|
| Noise floor | Everything audible when nobody is speaking | At the recording, permanently | Voice at least 15 dB above it; more for mobile listening |
| Peak level (dBFS) | How close the loudest instant is to the digital ceiling | At the recording, permanently if exceeded | Peaks at −12 to −18 dBFS while capturing |
| Average loudness (LUFS) | Perceived loudness over the whole programme | At delivery, freely adjustable | Set by the destination — see the next section |
| True peak (dBTP) | Peaks that emerge after encoding for delivery | At export, freely adjustable | −1 dBTP ceiling |
Once a programme is finished it has to be handed over at a specific loudness, and the target depends entirely on where it is going. This is the part that surprises people who learned audio from music production, because the instinct there is to push a mix as loud as it will go. On every major distribution route that instinct is now not just useless but actively harmful.
The reason is loudness normalisation. Broadcasters and streaming platforms measure incoming audio and adjust it to a house target, so a file delivered louder than the target is turned down rather than played louder. What you lose in exchange is dynamic range, because the usual method of getting a file louder is heavy compression and limiting. A programme squashed to maximum and then normalised downward sounds flat, fatiguing and strangely lifeless beside one delivered at target with its dynamics intact. Loudness is no longer a competitive variable. Only the quality of what is inside the target is.
The numbers below are the published targets of the relevant standards and platforms rather than house preferences. EBU R 128, the European broadcast standard, and ATSC A/85 in the United States cover anything going to a television schedule. The podcast and streaming figures are much louder because they are designed for headphones and noisy environments rather than a living room. If a piece is going to more than one destination, master to the strictest requirement it faces and let normalisation handle the rest rather than producing several irreconcilable versions.
Published loudness targets by destination
| Destination | Target | Peak ceiling | Note |
|---|---|---|---|
| European broadcast (EBU R 128) | −23 LUFS | −1 dBTP | Tight tolerance for pre-recorded material; live allows more latitude |
| US broadcast (ATSC A/85) | −24 LKFS | −2 dBTP | LKFS and LUFS are the same measurement under two names |
| Podcast and spoken-word streaming | −16 LUFS stereo, −19 LUFS mono | −1 dBTP | Louder by design, for earbuds and noisy surroundings |
| Music streaming services | Around −14 LUFS | −1 dBTP | Delivering louder than this gains nothing; it is normalised down |
Loudness cannot be judged by ear against a number, and it cannot be read off a peak meter. It needs a meter that implements the standard. Free options exist and are sufficient: ffmpeg will print integrated LUFS and true peak for a file through its loudnorm filter in analysis mode, and will also apply a correction in a second pass. Audacity ships a loudness normalisation effect that accepts a LUFS target. Several free plug-in meters work inside any editor that hosts plug-ins. Measure the finished export rather than the timeline, because the export is the file that gets judged, and measure the whole programme rather than a section — integrated loudness is an average over the full duration, so a two-minute sample of a thirty-minute piece tells you very little.
A single microphone recording one person produces one channel of information. Duplicating it across two channels does not make it stereo; it makes a mono file twice the size, and it invites the mistake of recording two speakers on two channels and delivering that as a stereo mix, which puts one person hard left and the other hard right in headphones. Mix speech to mono, or to a centred stereo file if the delivery spec demands two channels, and remember that the podcast loudness targets differ between mono and stereo by three LU for exactly this reason.
Video runs at 48 kHz, so recording speech at 48 kHz avoids a resampling step and the drift that comes from mixing rates across devices in one project. Set every device in the chain to the same rate before the first take. Twenty-four bit depth is worth taking wherever it is offered, not for resolution reasons anyone can hear but because it makes the conservative recording level above genuinely free: at 24-bit, a track peaking at −18 dBFS has no audible penalty when brought up later, whereas at 16-bit the same move lifts the noise floor with it. Higher rates than 48 kHz buy nothing for a voice and cost storage, battery and processing.
The single most useful division in production audio is between problems that are still solvable when the recording exists and problems that were decided the moment the file was written. Knowing which is which changes what you check before the interview rather than after it.
A recording that is too quiet or too loud overall, provided it never clipped, is a non-problem. Normalisation to a loudness target is a single operation and it is accurate. This is the one everybody frets about on the day and the one that costs nothing to correct, which is why the correct behaviour on set is to record conservatively rather than to chase a healthy-looking meter.
Audio that hit 0 dBFS has had its peaks removed from the file. Declipping tools exist and they interpolate a guess at the missing shape; on brief, isolated peaks the result can be acceptable, and on sustained clipping it is not. There is no version of this that recovers the original. Headroom at the recording stage is the only defence, and it is free.
Reflections arriving within roughly thirty milliseconds of the direct sound fuse with it perceptually, which is why a close microphone in a small soft room sounds natural. Later arrivals read as a hollow smear, and because they are the same voice, arriving fractionally later, no process can separate them from it. De-reverberation tools have improved a great deal and they still work by removing something, so what you buy is a choice between a room and a thinner, slightly synthetic voice. Choose the room before you record.
Hiss, ventilation hum and traffic rumble can be reduced convincingly if the noise is constant and the voice sits well above it, particularly with a clean sample of noise alone. Push the reduction too far and speech acquires a watery, gated quality that is more distracting than the hiss was. A high-pass filter around 80 Hz removes rumble that carries no speech information and is nearly always worth applying. Thirty seconds of room tone recorded on the day is what makes the rest of this possible.
A passing lorry, a slammed door, a phone notification or an aircraft over the middle of the best answer of the day cannot be lifted out from underneath a voice, because the two occupy the same time and much of the same frequency range. The fix is editorial rather than technical: cut around it if the sentence can be lost, or ask the question again. This is the category that most rewards listening on headphones while recording, when a second take still costs thirty seconds.
Where one person sat close to the microphone and the other did not, the near voice can be reduced but the far voice arrives with its own room baked in. You can match the levels and the imbalance remains audible as a change of acoustic every time the conversation turns. Two microphones, or a microphone genuinely close to both speakers, is the only clean answer.
The thump of a p or b striking the microphone head-on is a short burst of low frequency that a high-pass filter or a manual volume dip will usually remove. Harsh s sounds respond to a de-esser. Handling noise from a hand on a case is low-frequency and often survivable, but it frequently triggers a phone microphone into ducking the voice underneath it, and that gain movement is not reversible.
Recordings that stop when a call arrives, when storage fills or when a phone overheats are the most common total losses in amateur production, and none of them are audio problems. They are checks not run before recording: aeroplane mode, storage space, battery, mains power, and splitting long interviews into takes of around twenty minutes.
Everything above assumes one person with a phone or a modest recorder, working alone, which is the situation nearly everyone proposing a first programme is actually in. It is worth being equally plain about the limits.
None of it substitutes for a quiet room and a close microphone. The measurements in this guide are diagnostic; they tell you what you have, and they do not improve it. A recording made two metres away in a tiled kitchen will measure exactly as badly as it sounds, and no delivery target will rescue it. Capture is where the outcome is decided, and the guidance on microphone placement and room choice is the part that does the heavy lifting.
Loudness standards also do not apply where there is no distribution route yet. Delivering a piece at −23 LUFS matters when a broadcaster ingests it against a specification. If a first programme will only ever be heard on a website, the useful target is the spoken-word figure of around −16 LUFS with a −1 dBTP ceiling, and the rest is preparation for a route that may or may not come.
And the honest statement about this platform: Kind Channel is new, the community is small, and nothing has aired. There is no technical desk that will check your levels before you record, no equipment loan scheme, no mastering or repair service, and no post-production team that will rescue a recording made in a bad room. Nobody here will hand a file back to you with a loudness report attached. What that means in practice is that the technical decisions in this guide are yours to make before you record, not corrections someone downstream will apply. A proposal is easier to assess when it says plainly what the audio will be recorded on and where, because those two facts predict most of what the finished programme will sound like — and a person who has already listened to thirty seconds of their empty room on headphones is describing a different kind of plan from one who has not.
It depends on where it is going, and the gap between routes is large. Anything destined for a European television schedule is delivered at −23 LUFS integrated with a true peak ceiling of −1 dBTP under EBU R 128; the United States equivalent, ATSC A/85, asks for −24 LKFS. Spoken-word audio published online or as a podcast is normally much louder, around −16 LUFS for stereo and −19 LUFS for mono, because it is heard on earbuds in noisy places. If there is no confirmed broadcast route, the spoken-word figure is the sensible default. Measure the finished export with a meter that implements the standard rather than judging it by ear.
Some of it, and the distinction is worth learning before you record rather than after. Overall level is trivially corrected, so recording conservatively costs nothing. Steady hiss or ventilation noise can be reduced convincingly, especially with a clean sample of the room alone, though pushing the reduction hard leaves speech sounding watery. Clipping is permanent, because the peaks were never written to the file. Room reverb is effectively permanent, because the reflections are the same voice arriving late and nothing can separate them from it. A lorry passing through the middle of an answer means a retake. The rule of thumb: level problems are cheap and everything acoustic is expensive.
Because peak level and perceived loudness are different measurements. A peak meter shows the loudest single instant; loudness is a time-averaged, frequency-weighted figure across the whole programme, expressed in LUFS. A recording with a few sharp transients — a laugh, a door, an emphatic word — can touch the top of the peak meter while the body of the speech sits far below it, so it measures quiet and is heard as quiet. This is the reason delivery specifications are written in LUFS rather than dBFS, and the reason a waveform cannot be judged by eye. Measure the export with a loudness meter; ffmpeg and Audacity both provide one at no cost.
For a programme built on people talking, yes, and the asymmetry is not subtle. Picture in a talking-heads piece is confirmation of who is speaking and can degrade a long way while still doing that job; audio carries the content itself. Intelligibility lives in the consonants, which are quiet and brief, so reverb and noise remove meaning before they become obviously audible as a fault — leaving an audience doing reconstruction work on every sentence until they tire and leave. There is a credibility effect too: Newman and Schwarz found in Science Communication in 2018 that listeners rated the same conference talk’s research, and its presenter, less favourably when the audio quality was degraded.