A model scores how closely a text resembles the writing it was trained to treat as good, so bias appears wherever resemblance and quality come apart. The reliable penalties fall on short drafts, spoken and vernacular registers, unfamiliar subjects, and English written by second-language speakers. None of them are judgements about the idea.
Flow diagram of 4 steps: The paraphrase test, The register swap, The order swap, The published-neighbour check.
A language model has never read a piece of writing and found it moving. What it has done is absorb an enormous quantity of text in which some writing was presented, arranged or labelled as better than other writing, and learn the surface properties that came with that label. When it scores a draft, it is answering a question about resemblance: how much does this look like the things I have learned to treat as good. Most of the time resemblance and quality travel together, which is why the notes are useful at all. Bias is the name for what happens in the gap between them.
That gap has three separate sources and they are worth keeping apart, because they behave differently. The first is the training text itself, in which some varieties of English, some subjects and some registers are represented thousands of times more heavily than others. The second is the preference tuning that follows, where human raters chose between model outputs and, in aggregate, rewarded answers that were longer, more fluent and more confident than the alternatives. The third is the request you make, which can put a thumb on the scale through nothing more than the order in which two things are presented.
None of this is the model holding an opinion about you. It is a statistical tendency with a direction you can predict in advance, and that is precisely what makes it manageable. A random error averages out across drafts and is not worth defending against. A systematic one pushes the same kind of writer down every single time, and if you happen to be that kind of writer, no amount of revision inside the register will fix it, because the register is what is being penalised.
The useful distinction is not between a fair score and an unfair one but between an error that is unpredictable and one that has a direction. If a model marks your clarity down because a sentence really did have two readings, that is noise you can fix. If it marks clarity down every time you write in the voice you actually speak in, that is a bias, and the correct response is to identify it once and then stop treating that column as information about your writing.
Some of these are well established in the research literature on using models as judges, and some are visible to anyone who runs the same paragraph through a scoring tool twice with small changes. What they have in common is a direction that does not depend on the piece being assessed. You can predict which way each one will push before you submit anything.
Biases with a predictable direction
| Bias | What it rewards | What it costs a writer |
|---|---|---|
| Length | More words, more sections, more qualification | A tight 300-word piece reads as underdeveloped beside a padded 700-word one |
| Register | Standard edited English, mid-formal, lightly hedged | Spoken, vernacular and regional writing comes back marked as unclear |
| Fluency as a proxy for thought | Smooth sentences and clean transitions | A rough draft with a real idea loses to a polished one with none |
| Familiarity of subject | Topics heavily covered in the training text | Local, recent or minority-community subjects read as niche or unsupported |
| Standardness of English | Writing by people schooled in the dominant variety | Second-language and dialect writers absorb a penalty unrelated to their idea |
| Position | Whichever item was read first in a comparison | A judgement between two drafts can flip when you swap the order |
| Self-resemblance | Text that already sounds model-written | Editing towards the machine register raises the score and flattens the voice |
The clearest public demonstration is not about scoring but about detection: a 2023 Stanford study found that AI-detection tools flagged a majority of TOEFL essays written by non-native English speakers as machine-generated, while classifying essays by US-born students almost perfectly. The detectors were keying on limited lexical variety and predictable word choice, which are properties of writing in a second language as much as properties of writing by a machine. Scoring tools are not detectors, but they read the same surface signals, and a system that mistakes ordinary second-language prose for machine output will also mistake it for weak writing.
The category is worth policing, because treating every unwelcome score as bias is a comfortable way to stop learning anything from feedback. Three things get misfiled here regularly.
The first is a genuine ambiguity that you can see once it is pointed out. If a sentence permits two readings and the model took the one you did not intend, it has not misjudged you — it has demonstrated that the text allows a reading you did not want. That is the single most useful thing an automated pass produces, and it is the note most often dismissed as the machine failing to understand.
The second is disagreement with the criteria rather than with their application. If you think originality should count for more than social impact, or that feasibility has no business being scored on an idea nobody has costed yet, that is an argument about how the rubric was designed. It is a legitimate argument and it is not a claim that the scoring was applied unfairly to you. Those two complaints need different remedies and conflating them makes both harder to act on.
The third is a criterion measuring something real that you have chosen not to prioritise. A piece that deliberately withholds its subject until the final paragraph will lose points on engagement, and the score is not wrong about what it observed. You made a trade knowingly. The number simply records the cost of it, and the right response is to accept the cost rather than to dispute the observation.
Each of these isolates one variable. Run them when a score does not match your own read of the work, and stop as soon as one of them answers the question.
Ask the model to summarise your argument back to you before you look at any score. If the summary is accurate and clarity still came back low, the draft is not unclear — something about the surface is being read as difficulty. If the summary is wrong, you have a real problem and the register question does not arise.
Rewrite one paragraph into flat standard prose, changing no claim, no fact and no structure — only the voice. Submit both. A score that moves is measuring your prose style; a score that holds is measuring your argument. This is the most direct test of the whole set and it takes about ten minutes.
Any time you ask a model to compare two versions, ask twice with the order reversed. If the winner changes, the comparison carried no information and you should not act on either answer. This costs one extra request and catches a failure that is otherwise completely invisible in the reply.
Take a piece of published writing you admire that sits in the same register as yours and run it through the same scoring. If a well-regarded piece scores badly for the same reasons yours did, the penalty attaches to the register rather than to your execution of it.
If the tests point at bias, keep the register and act only on the notes the paraphrase test corroborated. But there is an honest complication worth stating: a register a model mis-scores is sometimes also a register that costs you readers, and sometimes it is not. The tests above tell you what the machine is reacting to; they cannot tell you whether a human audience would react the same way. That question needs people, and it is one of the few places where a small number of real readers beats any amount of automated assessment.
Every bias described above is present in the scoring on this site. It would be straightforwardly dishonest to write a page about systematic bias in model evaluation and imply that our own implementation has been exempted from it.
What is real, and checkable: each submission is scored on five criteria from 0 to 10 — clarity, originality, social impact, engagement potential and feasibility — and combined into a figure out of 100 with fixed published weights of 20, 20, 30, 20 and 10 per cent. The written reasoning is shown in plain language alongside every score, and it is visible to you and to anyone reading the idea in the public feed. The notes block nothing. You can revise and resubmit as often as you like. Community votes, cast by people, decide what gets made.
What we do not run: any fairness audit at all. There is no measurement of how scores are distributed across different kinds of writer, no calibration set, no test of whether the clarity criterion is doing work independent of register, and no adjustment applied to any score. That absence is currently rational rather than principled. Kind Channel is new, the community is small and nothing has aired, so there is no distribution to audit — a fairness statistic computed on this volume would be noise wearing the costume of assurance, and publishing one would be worse than publishing nothing.
So the residual exposure is real and worth naming precisely. A good idea submitted in strong vernacular, or written by somebody working in their second language, will probably lose points on clarity and engagement that the idea itself does not deserve, and we have no correction for that today. The two protections that do exist are structural rather than measured: the reasoning is published next to the number, so a mis-scoring is visible rather than buried, and the number has no authority over the outcome. Whether the rubric penalises particular registers is a question we expect to be able to answer once there are enough submissions to look at, and not before.
The evidence says yes, and the mechanism is well understood. A 2023 Stanford study found AI-detection tools flagged a majority of TOEFL essays by non-native English speakers as machine-generated while judging US-born students almost perfectly, because both tools key on limited lexical variety and predictable word choice. Scoring systems read the same surface signals, so ordinary second-language prose tends to lose points on clarity that the underlying thinking does not deserve.
Because length was rewarded during the preference tuning that shaped the model. Human raters comparing outputs tended to pick the fuller, more qualified answer, and that tendency is now baked into how models assess text as well as how they produce it. The practical effect is that a tight 300-word piece can read as underdeveloped beside a padded 700-word one covering less ground. Cutting a draft to its strongest form is often penalised by scoring and rewarded by readers.
It will raise the score, which is exactly the reason to be careful. Editing towards the register a model rewards moves your writing towards the register the model itself produces, and the usual result is prose that is smoother, more even and noticeably less like you. A better sequence is to test whether the register is what is being penalised — rewrite one paragraph into flat standard prose and see whether the score moves — then decide deliberately, rather than converging on the machine voice by default.
No, and it would be misleading to imply otherwise. There is no fairness audit, no calibration set and no adjustment applied to any score, because the platform is new and the volume of submissions is far too small for such a measurement to mean anything. What exists instead is structural: the written reasoning is published beside every score so a bad judgement is visible rather than hidden, the score blocks no submission, and community votes decide what actually gets made.