A language model reliably catches structural problems: unclear claims, missing context, buried arguments, inconsistent framing. It cannot verify facts, judge whether something is true to your experience, or reliably assess unusual work, which it tends to read as unclear. Treat its notes as a careful first reader, not a verdict.
Diagram of the 5 areas this page covers: What it is genuinely good at; What it cannot do; The unusual-work problem; How to use the notes without deferring to them; The trap: optimising for the feedback.
A model is an extremely consistent reader of structure. It will notice that your second paragraph assumes something you never established, that the thing your title promises does not arrive until two thirds of the way down, or that you have used one term to mean two different things. Human readers notice these too, but inconsistently and usually only when they are already annoyed.
It is also tireless and unembarrassed, which matters more than it sounds. Most people will not tell you your opening is weak, and almost nobody will tell you twice. A model will say the same thing on the fourth draft without any social cost to either of you.
The most common gap in human feedback is not accuracy, it is candour. A model has no relationship to protect, so it will say the boring true thing — this is unclear, this is unsupported, this repeats — that a friend will skip.
It cannot verify anything. If you write that a scheme served four hundred people, the model has no way to check and will treat the number as given. Every factual claim in the work is your responsibility alone, and confident-sounding feedback on a piece full of errors can create a false sense that it has been checked.
It cannot judge whether something is true to your experience. When you write about your own life, the model can assess whether the writing is clear but not whether it is honest, and those are different qualities that only you can evaluate.
And it cannot reliably assess work that does not resemble what it has seen. This is the failure that matters most for creative work.
A model predicts based on patterns in existing text. Work that follows a recognisable pattern is easy for it to evaluate; work that departs from one often reads to it as confused rather than novel. A deliberately fragmented structure, an unfamiliar cultural reference frame, or an argument that inverts a common assumption can all come back as low clarity.
So a poor clarity note on unconventional work is genuinely ambiguous. It might mean the writing is unclear. It might mean the writing is unfamiliar. Distinguishing those two is a judgement the model cannot make and you can, which is the strongest single argument for never giving it a veto.
How to read the notes
| If the note says… | It is usually reliable when… | Be sceptical when… |
|---|---|---|
| This is unclear | The passage has several possible readings | The idea is unconventional rather than muddled |
| This is unoriginal | The claim is genuinely common | The framing is new but the subject is familiar |
| This lacks support | You asserted something without evidence | The evidence is lived experience you did state |
| This will not engage | The opening delays the point | The audience is narrow and you know it exists |
| This is not feasible | It needs resources nobody has | It only needs resources the model cannot know about |
Read the reasoning, not the score. A number tells you nothing actionable; the explanation behind it either identifies a real problem or reveals that the model misread you. Both are useful and you can only tell which by reading the argument.
Then decide deliberately which notes to act on, and write down why you are ignoring the others. That last step sounds bureaucratic and is the whole discipline: it forces you to distinguish between a note you disagree with for a reason and a note you are avoiding because acting on it would be work.
Once you can see what a model rewards, it becomes tempting to write for it. This produces work that scores well and interests nobody — clear, well-supported, conventionally framed, and completely forgettable. The scores go up as the work gets worse, which is the most dangerous shape a metric can have.
The defence is to treat the feedback as diagnostic rather than as a target. It is good at telling you where a reader might stumble. It has no view worth having on whether your idea is worth making.
It can tell whether your writing is clear, consistent and structurally sound, which is part of good but not the whole of it. It cannot judge whether the work is true, necessary or worth anyone attention, and those are usually what determine whether something matters.
Because unfamiliar work often reads to a model as unclear rather than as new. It evaluates against patterns in existing text, so a genuine departure from those patterns can be penalised for exactly the quality that makes it interesting.
Only where the note identifies a real problem a reader would also hit. Rewriting to raise a score produces work that satisfies the rubric and nobody else, and the score rising can disguise the work getting duller.
No. It has no way to verify claims and will generally treat what you assert as given. Accuracy remains entirely your responsibility, and polished feedback on an inaccurate piece can be misleading precisely because it sounds thorough.