Home › Guides › Bias in AI evaluation

Bias in AI evaluation

A model scores how closely a text resembles the writing it was trained to treat as good, so bias appears wherever resemblance and quality come apart. The reliable penalties fall on short drafts, spoken and vernacular registers, unfamiliar subjects, and English written by second-language speakers. None of them are judgements about the idea.

Flow diagram of 4 steps: The paraphrase test, The register swap, The order swap, The published-neighbour check.

A model scores resemblance, not quality

A language model has never read a piece of writing and found it moving. What it has done is absorb an enormous quantity of text in which some writing was presented, arranged or labelled as better than other writing, and learn the surface properties that came with that label. When it scores a draft, it is answering a question about resemblance: how much does this look like the things I have learned to treat as good. Most of the time resemblance and quality travel together, which is why the notes are useful at all. Bias is the name for what happens in the gap between them.

That gap has three separate sources and they are worth keeping apart, because they behave differently. The first is the training text itself, in which some varieties of English, some subjects and some registers are represented thousands of times more heavily than others. The second is the preference tuning that follows, where human raters chose between model outputs and, in aggregate, rewarded answers that were longer, more fluent and more confident than the alternatives. The third is the request you make, which can put a thumb on the scale through nothing more than the order in which two things are presented.

None of this is the model holding an opinion about you. It is a statistical tendency with a direction you can predict in advance, and that is precisely what makes it manageable. A random error averages out across drafts and is not worth defending against. A systematic one pushes the same kind of writer down every single time, and if you happen to be that kind of writer, no amount of revision inside the register will fix it, because the register is what is being penalised.

Systematic is the word doing the work

The useful distinction is not between a fair score and an unfair one but between an error that is unpredictable and one that has a direction. If a model marks your clarity down because a sentence really did have two readings, that is noise you can fix. If it marks clarity down every time you write in the voice you actually speak in, that is a bias, and the correct response is to identify it once and then stop treating that column as information about your writing.

The biases that show up reliably

Some of these are well established in the research literature on using models as judges, and some are visible to anyone who runs the same paragraph through a scoring tool twice with small changes. What they have in common is a direction that does not depend on the piece being assessed. You can predict which way each one will push before you submit anything.

Biases with a predictable direction

BiasWhat it rewardsWhat it costs a writer
LengthMore words, more sections, more qualificationA tight 300-word piece reads as underdeveloped beside a padded 700-word one
RegisterStandard edited English, mid-formal, lightly hedgedSpoken, vernacular and regional writing comes back marked as unclear
Fluency as a proxy for thoughtSmooth sentences and clean transitionsA rough draft with a real idea loses to a polished one with none
Familiarity of subjectTopics heavily covered in the training textLocal, recent or minority-community subjects read as niche or unsupported
Standardness of EnglishWriting by people schooled in the dominant varietySecond-language and dialect writers absorb a penalty unrelated to their idea
PositionWhichever item was read first in a comparisonA judgement between two drafts can flip when you swap the order
Self-resemblanceText that already sounds model-writtenEditing towards the machine register raises the score and flattens the voice

The evidence on non-standard English is the strongest of these

The clearest public demonstration is not about scoring but about detection: a 2023 Stanford study found that AI-detection tools flagged a majority of TOEFL essays written by non-native English speakers as machine-generated, while classifying essays by US-born students almost perfectly. The detectors were keying on limited lexical variety and predictable word choice, which are properties of writing in a second language as much as properties of writing by a machine. Scoring tools are not detectors, but they read the same surface signals, and a system that mistakes ordinary second-language prose for machine output will also mistake it for weak writing.

What is not bias, even when it feels like it

The category is worth policing, because treating every unwelcome score as bias is a comfortable way to stop learning anything from feedback. Three things get misfiled here regularly.

The first is a genuine ambiguity that you can see once it is pointed out. If a sentence permits two readings and the model took the one you did not intend, it has not misjudged you — it has demonstrated that the text allows a reading you did not want. That is the single most useful thing an automated pass produces, and it is the note most often dismissed as the machine failing to understand.

The second is disagreement with the criteria rather than with their application. If you think originality should count for more than social impact, or that feasibility has no business being scored on an idea nobody has costed yet, that is an argument about how the rubric was designed. It is a legitimate argument and it is not a claim that the scoring was applied unfairly to you. Those two complaints need different remedies and conflating them makes both harder to act on.

The third is a criterion measuring something real that you have chosen not to prioritise. A piece that deliberately withholds its subject until the final paragraph will lose points on engagement, and the score is not wrong about what it observed. You made a trade knowingly. The number simply records the cost of it, and the right response is to accept the cost rather than to dispute the observation.

Four tests that separate bias from a defect

Each of these isolates one variable. Run them when a score does not match your own read of the work, and stop as soon as one of them answers the question.

The paraphrase test

Ask the model to summarise your argument back to you before you look at any score. If the summary is accurate and clarity still came back low, the draft is not unclear — something about the surface is being read as difficulty. If the summary is wrong, you have a real problem and the register question does not arise.

The register swap

Rewrite one paragraph into flat standard prose, changing no claim, no fact and no structure — only the voice. Submit both. A score that moves is measuring your prose style; a score that holds is measuring your argument. This is the most direct test of the whole set and it takes about ten minutes.

The order swap

Any time you ask a model to compare two versions, ask twice with the order reversed. If the winner changes, the comparison carried no information and you should not act on either answer. This costs one extra request and catches a failure that is otherwise completely invisible in the reply.

The published-neighbour check

Take a piece of published writing you admire that sits in the same register as yours and run it through the same scoring. If a well-regarded piece scores badly for the same reasons yours did, the penalty attaches to the register rather than to your execution of it.

What to do once you know which it is

If the tests point at bias, keep the register and act only on the notes the paraphrase test corroborated. But there is an honest complication worth stating: a register a model mis-scores is sometimes also a register that costs you readers, and sometimes it is not. The tests above tell you what the machine is reacting to; they cannot tell you whether a human audience would react the same way. That question needs people, and it is one of the few places where a small number of real readers beats any amount of automated assessment.

What Kind Channel does, and what we cannot claim

Every bias described above is present in the scoring on this site. It would be straightforwardly dishonest to write a page about systematic bias in model evaluation and imply that our own implementation has been exempted from it.

What is real, and checkable: each submission is scored on five criteria from 0 to 10 — clarity, originality, social impact, engagement potential and feasibility — and combined into a figure out of 100 with fixed published weights of 20, 20, 30, 20 and 10 per cent. The written reasoning is shown in plain language alongside every score, and it is visible to you and to anyone reading the idea in the public feed. The notes block nothing. You can revise and resubmit as often as you like. Community votes, cast by people, decide what gets made.

What we do not run: any fairness audit at all. There is no measurement of how scores are distributed across different kinds of writer, no calibration set, no test of whether the clarity criterion is doing work independent of register, and no adjustment applied to any score. That absence is currently rational rather than principled. Kind Channel is new, the community is small and nothing has aired, so there is no distribution to audit — a fairness statistic computed on this volume would be noise wearing the costume of assurance, and publishing one would be worse than publishing nothing.

So the residual exposure is real and worth naming precisely. A good idea submitted in strong vernacular, or written by somebody working in their second language, will probably lose points on clarity and engagement that the idea itself does not deserve, and we have no correction for that today. The two protections that do exist are structural rather than measured: the reasoning is published next to the number, so a mis-scoring is visible rather than buried, and the number has no authority over the outcome. Whether the rubric penalises particular registers is a question we expect to be able to answer once there are enough submissions to look at, and not before.

Do AI scoring tools penalise people writing in a second language?

The evidence says yes, and the mechanism is well understood. A 2023 Stanford study found AI-detection tools flagged a majority of TOEFL essays by non-native English speakers as machine-generated while judging US-born students almost perfectly, because both tools key on limited lexical variety and predictable word choice. Scoring systems read the same surface signals, so ordinary second-language prose tends to lose points on clarity that the underlying thinking does not deserve.

Why does a longer draft usually score better than a shorter one?

Because length was rewarded during the preference tuning that shaped the model. Human raters comparing outputs tended to pick the fuller, more qualified answer, and that tendency is now baked into how models assess text as well as how they produce it. The practical effect is that a tight 300-word piece can read as underdeveloped beside a padded 700-word one covering less ground. Cutting a draft to its strongest form is often penalised by scoring and rewarded by readers.

Should I rewrite my work in a more standard style to score higher?

It will raise the score, which is exactly the reason to be careful. Editing towards the register a model rewards moves your writing towards the register the model itself produces, and the usual result is prose that is smoother, more even and noticeably less like you. A better sequence is to test whether the register is what is being penalised — rewrite one paragraph into flat standard prose and see whether the score moves — then decide deliberately, rather than converging on the machine voice by default.

Does Kind Channel test its own scoring for bias?

No, and it would be misleading to imply otherwise. There is no fairness audit, no calibration set and no adjustment applied to any score, because the platform is new and the volume of submissions is far too small for such a measurement to mean anything. What exists instead is structural: the written reasoning is published beside every score so a bad judgement is visible rather than hidden, the score blocks no submission, and community votes decide what actually gets made.

AI feedback on creative workWhat is AI editorial feedback?AI as editor, not generatorHow scoring rubrics work
Real Stories of Kindness That Restored Our Faith in Humanity50 Everyday Acts of Kindness Anyone Can Do Starting TodayKindness in the Workplace — A Business Strategy, Not Just a Value
HomeExplore ideasSubmit an ideaGuidesWho it is forCompare platformsFAQWhy Kind Channel existsAbout