Home › Guides › How scoring rubrics work

How scoring rubrics work

A rubric turns one judgement into several smaller ones, each scored separately, then recombines them with fixed weights. That makes the reasoning visible and the arithmetic auditable, but it does not make the underlying judgements measurements. A 7 is a considered opinion expressed as a number, and it is only useful if you can read what produced it.

Flow diagram of 4 steps: Anything whose value is one unrepeatable detail, Humour, tone and timing, Work whose form is the argument, Anything requiring local knowledge to value.

A rubric is a way of splitting up a judgement, not a way of measuring

Ask somebody whether an idea is any good and you get one answer, usually a feeling, and no way to argue with it. A rubric breaks that single verdict into a handful of narrower questions, scores each one on its own, and then adds the results back together with fixed weights. The gain is entirely in the decomposition: instead of "this is a six", you get "this is clear and unoriginal", which is a sentence you can actually act on. Rubrics were invented for marking essays at scale, and the reason they spread is that they make an assessor explain themselves.

What decomposition does not do is turn an opinion into a measurement, and this is where most people misread scores. A ruler works because length exists independently of the ruler. Nothing equivalent is happening when a reader — human or model — puts 7 next to originality. The number is a compressed opinion, and every property a real measurement has is missing: it is not reproducible, the intervals are not equal, and there is no unit. Two assessors can produce 6 and 8 on the same paragraph without either being wrong, because there is nothing they were both approximating.

That is not an argument against scoring. It is an argument about what to do with the output. A rubric score is worth reading as an index into the reasoning — it tells you which of the narrower questions went badly, so you know which paragraph of explanation to read first. Read as a verdict, it tells you almost nothing you can use.

The scale is ordinal, so the gaps are not equal

Moving from 3 to 4 on clarity usually means a genuine defect got fixed. Moving from 8 to 9 often means nothing at all — the top of any 0-10 scale is crowded, and assessors are reluctant to spend the last point. Scores on creative work also cluster hard around 7 and 8, which means the visible range is roughly four points wide, not eleven. Treat a one-point difference near the top as noise and a one-point difference near the bottom as information.

Reproducibility is the property most people assume and no rubric has

Submit identical text twice to a language model and the component scores can differ, because the output is sampled rather than computed. Human assessors are worse: marker agreement on essays is a well-documented problem, which is precisely why rubrics exist. So the honest reading of any single score is that it is one draw from a distribution. If a number matters enough to argue about, the argument should be about the written reasoning, which is far more stable than the digit attached to it.

The arithmetic behind a combined score, using the one on this site

Kind Channel scores each submission on five criteria, each 0-10, then combines them into a figure out of 100 with fixed weights: clarity 20 per cent, originality 20, social impact 30, engagement potential 20, feasibility 10. There is nothing clever about the formula — it is a weighted sum — and publishing it is more useful than describing it, because once you can see the weights you can see what the rubric is for.

Those weights are an editorial position written as arithmetic, which is the part worth dwelling on. Social impact carries the largest share, so the rubric is saying that whether something addresses a real need matters more than whether it is watchable. Feasibility carries the smallest, so it is also saying that difficulty should not kill an idea at the first stage — a thing that is hard to make can still be worth putting in front of people, and production problems are for later. Every rubric encodes a set of priorities like this. Most do not tell you what they are.

The five criteria and what each weight is really saying

CriterionWeightWhat it asksWhat it cannot see
Clarity20%Can a reader tell what this is and what it would contain?Whether an unfamiliar structure is confused or deliberate
Originality20%Has this angle been done, or is the framing genuinely new?Whether it is new to your community, which is often what matters
Social impact30%Does it address a real human or social need?Whether that need is acute where you live, or already met
Engagement potential20%Would somebody not already interested watch this?Who your audience actually is, and whether a small one is enough
Feasibility10%Could this realistically be produced as a programme?What access, contacts, footage or permissions you already hold

What the weights do to two specific submissions

A clear, unoriginal idea about a real need — clarity 8, originality 5, impact 8, engagement 7, feasibility 6 — comes out at 70. Its mirror image, an original idea that is hard to follow — clarity 5, originality 9, impact 7, engagement 6, feasibility 5 — comes out at 66. Four points apart, which is the rubric telling you it cannot really separate them. The useful information is not the gap; it is that one of them needs a rewrite and the other needs a fresher angle, and no combined figure will ever say that.

A 10 per cent weight means an impossible idea can still score 90

Score a submission 10 on clarity, originality, impact and engagement and 0 on feasibility, and the formula returns 90. That is a real consequence of the weighting, not a bug: the arithmetic genuinely does not care whether the thing can be made. It is worth knowing because it tells you where the actual constraint lives. Feasibility is nearly free at the scoring stage and expensive later, when somebody has to work out whether the footage exists and who would agree to appear.

Why combined scores compress, and why the criteria overlap

Averaging pulls everything towards the middle. That is a property of averaging rather than a flaw in any particular rubric, but it has a specific effect on creative work: it flattens the distinctive submission. A piece that is remarkable on one dimension and mediocre on the rest — which describes a great deal of interesting work — averages out to unremarkable, while a piece that is competent everywhere and interesting nowhere sails through. If you are picking things to make, the composite is systematically biased towards the second kind, and the component scores are the only place the first kind shows up.

The second problem is that the criteria are not independent, and weighted sums assume they are. Clarity, engagement potential and, to a large extent, originality all move with how well the thing is written. A well-expressed idea scores better on all three at once, so the composite quietly counts writing quality something like three times and social need once. The stated weight on impact is 30 per cent; the effective weight is lower than that, because two of the other four criteria are partly measuring prose. This is the single most common defect in scoring rubrics of every kind, and it is invisible unless you go looking for correlations between the columns.

Both problems have the same practical answer, which is to read across the criteria rather than down to the total. A submission at 66 with a 9 for originality is a more interesting object than a submission at 74 with 7s across the board, and the composite ranks them the other way round.

Where rubrics stop working on creative work

A rubric can only score what its criteria name, and it applies them to a proposal rather than to a finished thing. Both limits bite hardest on exactly the work most worth making.

Anything whose value is one unrepeatable detail

The whole point of some pieces is a single moment — a thing somebody said, a scene nobody could have arranged. That does not decompose. It scores as a thin proposal because there is nothing to describe except the moment, and describing it well is a separate skill from having it.

Humour, tone and timing

No criterion on any general rubric covers whether something is funny, and funny does not survive summarisation. A comic idea explained in a submission box reads flat almost regardless of how good it is, so it will score as low engagement for reasons that have nothing to do with the programme it would become.

Work whose form is the argument

If the structure is deliberately fragmented or the piece withholds its subject on purpose, clarity scoring reads the intention as a defect. The score is not wrong about what it saw; it is answering a question about legibility when the relevant question was about effect.

Anything requiring local knowledge to value

Social impact assumes the assessor can tell whether a need is real. For a specific street, estate or town, it cannot. It will guess from how common the general problem is, which is a different question and frequently gives the wrong answer for the place you actually mean.

Publishing the weights makes gaming possible, and that is still the right trade

Anybody who reads this page now knows to weight their submission towards social need and not to worry much about production difficulty. That is a real cost of transparency. It is worth paying because the alternative — a hidden formula — does not stop people optimising, it just means only the ones who guess correctly benefit, and nobody can tell whether the rubric is reasonable. The defence against gaming is not secrecy. It is that a community vote sits after the scoring stage and is not impressed by a submission engineered to look impactful.

What the numbers on this site do, and what they do not

Concretely: the five component scores and a combined figure out of 100 are returned with written reasoning, and all of it is visible to you and to anyone reading your idea in the public feed. The combined figure attaches a label — 60 and above reads as accepted, 35 to 59 as needing work, below 35 as not a fit — and it is worth being precise about what those thresholds mean, because they are not demanding. Five straight sixes clears 60 exactly. Falling under 35 takes averages below four across every criterion, which in practice means a one-line or off-topic submission rather than a weak idea.

The label is a label, not a gate. A low score does not remove your idea from the feed, does not stop it collecting votes, and does not prevent it being scheduled — community votes decide what gets made, and the model casts none. You can also revise and resubmit as often as you like. So the honest description of the number is that it is advice with a badge attached, and the badge is the least useful part of it.

Two limits to state plainly. The component scores and the combined figure are produced together by the model rather than the total being recalculated from the parts afterwards, so the two can drift apart slightly; if they disagree, the components are the thing to trust, and adding up your own five scores against the published weights is a check you can run in a few seconds. And there is no calibration behind any of this yet. Kind Channel is new, the community is small and nothing has aired, so nobody can tell you what a 72 predicts about how an idea will fare here, because there is no record to compare it against. What can be said without qualification is the design: the weights are published above, the reasoning is shown, the model blocks nothing, and people vote.

What does a score of 7 out of 10 on creative work actually mean?

It means an assessor put that idea in the upper-middle band on one specific question, and very little more. A 0-10 scale applied to creative work is ordinal rather than measured: the gaps between points are not equal, scores cluster around 7 and 8 so the usable range is narrower than it looks, and the same text scored twice can come back one point different. The number is best used to find which written note to read first.

Why does a weighted total hide more than the individual criteria show?

Because averaging pulls everything towards the middle, and it does that hardest to work that is exceptional on one dimension and ordinary elsewhere. A submission that is genuinely original but badly explained ends up next to a submission that is competent and dull, and the total cannot tell you which is which. The component scores can: one needs a rewrite, the other needs a better idea. Read across the criteria rather than down to the total.

Can I raise my score by writing what the rubric wants?

Partly, and it will not get you as far as it looks. Weighting a submission towards whichever criterion carries the most weight does move the number, which is a real cost of publishing the formula. What it does not move is whether people vote for the idea, and on Kind Channel the vote decides what gets made. A submission engineered to read as high-impact tends to read as generic to a human, which is the failure mode the score cannot detect and a reader spots immediately.

Do the five criteria measure five separate things?

Not cleanly, and this is the standard weakness of weighted rubrics everywhere. Clarity, engagement potential and to some degree originality all improve when the writing improves, so a well-written submission gains on three criteria at once and the composite ends up counting prose quality several times over. The practical effect is that the stated weight on social need is higher than its real influence on the total.

AI feedback on creative workWhat is AI editorial feedback?AI as editor, not generatorWhen to disclose that AI helped
50 Everyday Acts of Kindness Anyone Can Do Starting TodayKindness in the Workplace — A Business Strategy, Not Just a ValuePositive News Stories That Will Restore Your Hope This Week
HomeExplore ideasSubmit an ideaGuidesWho it is forCompare platformsFAQWhy Kind Channel existsAbout