Cut structurally before you cut cosmetically. Take out everything before the real start, then your own questions, then false starts and restarts, then whole digressions, and only afterwards tidy the hesitations that survive. A forty-minute interview usually contains about six usable minutes. Clipping, room reverb and two voices sharing one track are recording faults, not editing problems.
Flow diagram of 5 steps: Get a transcript and make the decisions on paper first, Cut the top, which is almost never where the piece starts, Decide what happens to your own questions, Delete whole answers, not parts of them, Remove false starts and restarts inside the answers you kept.
There is a recognisable way a first edit goes wrong, and it is not carelessness. Somebody imports forty-three minutes of interview, plays from the top, hears the interviewee say “um, so, I suppose it started when” and removes the um. Then the next one. Two hours later they are four minutes into the recording, the timeline is a thicket of tiny cuts, and no decision has yet been made about what the piece is actually about. Almost everyone does this once.
It is wasted work for two separate reasons, and the second is the more interesting one. The first is arithmetic. Most of that opening passage is going to be deleted whole, so every hesitation removed from it was removed from material that will not survive the afternoon. A structural cut takes about ten seconds to make and removes four minutes of programme; a hesitation cut takes twenty seconds of careful scrubbing and removes four tenths of a second. Doing the second kind before the first kind inverts the return on every minute spent.
The second reason is that fine work performed before the structure exists is not merely early, it is uninformed. You cannot tell whether a pause is clumsy or useful until you know what sits either side of it in the finished piece. A hesitation in the middle of a long answer reads as clumsy; the same hesitation immediately after a hard cut reads as somebody gathering themselves, and it is often the thing that stops the cut from sounding abrupt. Every editing discipline converges on the same order — decide, then shape, then polish — because polish applied to an undecided thing has nothing to be faithful to.
Worth saying plainly as well: removing every hesitation is not automatically an improvement. Spontaneous speech contains them, listeners do not consciously register most of them, and stripping the lot produces a delivery with an odd machined quality that people notice without being able to name. The stakes rise if the speaker has a stammer, speaks English as a second language, or is simply a slow and considered talker. Editing that pattern out edits away how the person actually sounds, which is a bigger decision than it looks like when it is being made one click at a time. Keep the hesitations that read as thinking; remove the ones that read as a technical fault or that sit awkwardly across a join.
The other thing to settle early is how much you expect to throw away, because first-time editors almost always under-cut rather than over-cut. Conversational speech runs at roughly 140 to 160 words a minute, so a forty-three minute interview is somewhere around six thousand words of transcript, and a six-minute segment is about nine hundred. That is a keep rate in the region of fifteen per cent. If your first assembly is twenty-two minutes long, you have not made a twenty-two minute programme; you have made a rough selects reel, and the work has not started.
Before any of this, make the files safe. Work on a copy, keep the originals untouched somewhere you are not editing from, and do not wipe a camera card or a phone until two copies of its contents exist in two places. Recovering a deleted answer is impossible in a way that redoing an edit never is, and a first-time producer is exactly the person most likely to fill a disk mid-session and start clearing things.
The sequence below runs coarse to fine. Each step assumes the previous one has been done, and each one is cheaper because of it. Resist the urge to jump ahead to a passage you already know you love — you will polish it, and then discover in step five that the answer it belongs to has to go.
Automatic transcription is now effectively free — Whisper runs locally, and most editors have transcript-based editing built in — and it is accurate enough to choose material from even when it is not accurate enough to publish as captions unchecked. Read the whole thing, mark the four to six statements that must survive, and note roughly what order they want to be in. This is the paper edit, and twenty minutes on it saves an evening. Do not skip it because the recording is short; the point is to make the structural decision while you can still see the whole interview at once.
The first few minutes of nearly every interview are throat-clearing: level checks, the slate, the interviewee settling, an answer they give again better twenty minutes later once they have relaxed. Find the first sentence that would make sense to somebody arriving cold and delete everything before it. The equivalent is usually true at the end, where the recording tails off into thanks and chairs moving.
There are three defensible choices and one bad habit. You can keep the questions, which is honest and costs runtime. You can remove them and put the question on screen as text, which is fast and reads cleanly. You can remove them and rely on answers that restate the question — which only works if you asked the interviewee to answer in complete sentences at the time. The bad habit is cutting questions out and leaving answers that start mid-thought with “yeah, exactly, and that was the problem”, which makes an audience work out what was asked.
Go through what is left and remove entire topics that are not in the piece. Whole-answer deletions are the cuts that actually create the programme, and they are emotionally the hardest, because a good answer to a question that has nothing to do with your subject is still a good answer. Keep a separate bin timeline for the ones you cannot bear to lose rather than trying to justify them into the edit.
People routinely begin an answer twice, and the second attempt is usually better organised than the first. Cut to the second. The same applies to the sentence somebody abandons halfway, and to the repeated clause that arrives when they lose the thread and reset. These are structural rather than cosmetic cuts and they remove real runtime — often a minute or more across a six-minute piece.
Now the piece exists end to end, tighten within it: the subordinate clause that repeats the main one, the qualifier stacked on a qualifier, the second example where one was enough. Play the whole thing through before starting this, and again afterwards, because sentence-level trimming is where pace gets accidentally flattened into something relentless with no room to breathe.
Remove hesitations that sit across a join, cluster distractingly, or read as a fault rather than as thought. Cut on the breath rather than tight to the word: leave the inhale attached to the phrase that follows it, because an entrance with no breath in front of it sounds spliced even when nothing else is wrong. Put a short audio crossfade of a couple of frames — roughly 40 to 80 milliseconds — across every audio cut to kill the click that a hard join produces.
Cuts leave gaps of digital silence that the ear hears as a hole, so lay the thirty seconds of room tone you recorded underneath the whole timeline. Then normalise the finished export to a loudness target, measuring the export rather than the clips. Doing it in the other order means re-doing it, because every subsequent cut changes the average.
A cut in sound behaves differently from a cut in picture, and almost every technique for hiding an edit comes down to letting one of them change while the other continues. The audience tracks continuity in the audio far more closely than in the image, which is why a visible jump cut with unbroken sound reads as a stylistic choice while an audible join under a continuous shot reads as a mistake.
The table covers the joins you will need on a first segment. None of them requires paid software; all of them are available in the free version of DaVinci Resolve, in Kdenlive or Shotcut, or in Audacity for audio-only work.
Joins in a talking-heads segment, what the audience registers, and what to do about it
| Join | What the audience notices | What to do |
|---|---|---|
| Hard cut inside one answer, same shot | A jump: the head and hands snap to a new position on the frame after the cut | Either lean into it as an obvious, rhythmic style, or cover it — a cutaway, a caption, or a slow push on the frame either side so the two positions no longer match exactly. |
| Cutting tight to the first word | An entrance that sounds clipped and slightly aggressive, with no air in front of it | Move the cut point earlier so the speaker’s inhale stays attached to the phrase. The breath is the sound that makes a join read as speech rather than as assembly. |
| J-cut and L-cut: sound leads or trails the picture | Nothing, which is the point — the scene changes while a voice carries across the boundary | Extend the audio a second or so past the picture cut, or start it early under the outgoing shot. The single most useful habit for making a first edit feel deliberate. |
| Audio join with no crossfade | A faint click or tick at the cut, often mistaken for a file problem | Apply a crossfade of a couple of frames at every audio edit. It is inaudible as a fade and it removes the discontinuity that produces the click. |
| A gap of true silence after removing a cough | The floor drops out for a fraction of a second and the room disappears, which sounds more edited than the cough did | Lay recorded room tone under the full timeline so gaps fall back to the same background rather than to nothing. |
| Cross-dissolve between two talking-head shots | A soft, slightly dated transition that signals time passing | Use sparingly and deliberately — as a marker of a real jump in time or topic, not as a way to soften a cut you were unsure about. |
| Fade to black at the end of a thought | A full stop, and often an exit: some of the audience will leave at a mid-programme fade | Keep fades for the actual end. Mid-piece, a hard cut with sound continuing holds attention where a fade releases it. |
Every edit is a claim that the shortened version is a fair account of what somebody said. Most of the time that claim is uncontroversial: nobody thinks a six-minute segment is the whole conversation. The problem is that the same tools that remove forty minutes of digression can also, with no extra effort and no obvious moment of decision, produce a sentence the person never said.
The practical test is simple to state and worth applying literally. Play the finished segment in your head as though the interviewee were sitting next to you. If there is a join you would want to explain, that join is the problem — not because they have a veto over the edit, which they do not, but because your discomfort is an accurate detector of a cut that has changed the meaning rather than the length.
Removing your questions, cutting whole topics, dropping repetitions and false starts, tightening a rambling sentence into the shorter version of itself, and putting two answers to closely related questions next to each other — all of these are ordinary and expected. So is removing a digression that leads nowhere, even if the speaker was proud of it. The unifying test is that none of these alters what the person was asserting, only how much of it is present and how densely it is packed.
The documentary term for a quote assembled from fragments recorded at different moments is a frankenbite, and it is the clearest thing on the wrong side of the line: half a sentence from minute four joined to the end of a sentence from minute thirty-one, producing a statement that reads as one utterance and was never one. The smaller and far more common version is the lost qualifier. Somebody says “that worked for us, though we had a council grant and most groups will not”, and the second clause goes because the segment was running long. What is left is a claim the speaker deliberately hedged and would not stand behind. Hedges are the first thing to disappear under time pressure and they are frequently the most load-bearing part of the sentence.
Moving material is not forbidden and cannot be, since a paper edit is largely an argument about order. What changes the meaning is moving an answer away from the question that framed it. An answer given about one particular week in 2019 becomes a general claim when it is lifted out and placed after a question about how things are now. If you reorder, carry the framing with it — in the speaker’s own words where possible, on screen as a caption otherwise — rather than assuming the audience will supply a context they were never given.
A small production has no compliance desk and no lawyer to ask, so the fallback is procedural rather than legal. Send the person the specific quotes you have used and a plain description of the context, and ask whether anything reads as unfair. This is not handing over editorial control, and it is worth saying so explicitly when you ask: you are checking accuracy, not seeking approval. If a cut still feels wrong after that, keep the longer version. Runtime is a cheaper thing to spend than the trust of the first few people who agreed to be filmed by someone with no track record.
A great deal of advice about editing quietly implies that problems are deferrable, and some genuinely are. Level is adjustable at any point. Pace is adjustable. A weak opening can be replaced by a better sentence from later in the interview. Others are not deferrable at all, and knowing which is which changes decisions made on the day rather than in the edit.
Clipping is permanent. When a recording hits the top of the scale the waveform is flattened and the information that was above the ceiling is not stored anywhere; a declipper guesses at the missing curve, and on a loud laugh or an emphatic word the guess sounds like a different kind of distortion rather than like the original. Room reverb is the same category of problem. Once reflections are baked into the same track as the voice there is no process that removes them without taking the texture of the speaker along with them, and the de-reverb tools that have improved most in recent years are still choosing which artefact you prefer.
Two people on one microphone is the one that surprises people most, because separation tools exist and are genuinely impressive on music. Separating two human voices recorded in the same room, overlapping, at similar levels is a much harder problem, and the output tends to be usable as a rescue and not as a programme. If two speakers talk over each other on a shared track, the overlap is gone — which is a reason to record two tracks, and a reason to ask people not to agree audibly while somebody else is speaking.
Intermittent noise is the quiet killer. A constant hum can be filtered and a steady floor can be lived with, but a lorry, an aircraft or a slammed door on top of the best sentence of the day is not removable, because the noise occupies the same frequencies and the same moment as the words. That is why it is worth asking for one more take of an answer when something passed through it, rather than assuming it will be fixable later. Missing room tone belongs in the same bracket: it costs thirty seconds to record and there is no way to synthesise a convincing substitute afterwards.
On the picture side, autofocus that hunted during the take, exposure that pumped as somebody gestured, and a phone image made in a dim room are all fixed at the point of recording. Grading changes colour and contrast; it does not restore detail that was never captured, and the smeared, faintly plastic look of heavy noise reduction on skin is not a grading problem.
It is also worth being plain about what this platform does and does not provide. Kind Channel is new. Nothing has aired, the community is small, and there is no post-production facility, no edit suite to book, no assistant editor who will sync your rushes and no colourist waiting at the end of the process. A first segment is something you will cut yourself, on a laptop, in free software, and the guidance above assumes exactly that. The practical consequence for anyone thinking about proposing something is that the edit is part of the commitment rather than a stage that gets handled elsewhere — and a proposal is easier to assess when it says honestly who is going to do it and roughly how long the finished piece is meant to run.
Let the material set the length rather than deciding a target in advance. For a single interview, something between four and eight minutes is usually where the strong material runs out, and a piece that ends while the audience is still interested is better than one padded to hit a round number. A useful check is arithmetic: conversational speech runs at roughly 140 to 160 words a minute, so a six-minute segment is about nine hundred words of transcript. If your assembly is three times that, the structural cutting is not finished.
No. Remove the ones that sit across a join, cluster distractingly, or read as a technical fault, and leave the ones that read as somebody thinking. Stripping all of them produces a clipped, machined delivery that listeners notice without being able to name, and the stakes are higher when the speaker has a stammer, is talking in a second language, or is simply deliberate in how they speak — in those cases aggressive removal edits away how the person actually sounds. It is also the slowest possible use of editing time, so it belongs at the very end of the process rather than the start.
No. The free version of DaVinci Resolve covers everything a talking-heads segment needs, including transcript-based editing, and Kdenlive and Shotcut are fully open source alternatives. Audacity handles audio-only work. Automatic transcription is effectively free as well, since Whisper runs on an ordinary laptop. What free tools do not remove is the need to check the licence terms of anything web-based before publishing, because some consumer editing apps claim broad rights over uploaded material, and some bundle music that is not cleared for use outside their own platform.
Yes, as long as the answer still carries its own framing. Problems appear when a question supplied the context and its removal turns a specific statement into a general one — an answer about one particular year becoming an apparent claim about the present, for example. Three workable approaches: keep the questions in, put each question on screen as text, or ask the interviewee at the time to answer in complete sentences that restate the subject. What to avoid is an answer that opens mid-thought with something like “yes, exactly, and that was the whole problem”, which leaves the audience reconstructing what was asked instead of listening.