How the score works
Twenty-two measurements. Four of them count.
If a number is going to tell you that you score 51 and the firm you keep losing work to scores 78, you are entitled to know what the number is made of before you believe a word of it.
So here is the whole of it, short of the source code: what is measured, what is measured and deliberately ignored, where the scale came from, and the four things this score is not.
The four that count
What is actually being measured
Every measure below earned its place the same way. It had to separate six real businesses, across five sectors, from a control written deliberately to look competent and mean nothing. Anything that failed that was dropped, however sensible it sounded.
The marks add to 100. The weights are not a secret and never have been: what makes this hard to copy is knowing what to throw away, not knowing what to add up.
Own words
35 marks
How much of the writing came out of your category's phrasebook. The stock phrases that turn up on every competitor's site, and the ones ChatGPT, Claude, Copilot and Gemini all reach for when they start from a blank page, because the phrasebook is what they have read most of.
Measured as a density rather than a count, so a long page cannot hide behind its own length. The single most reliable signal there is, and the one that costs most sites their marks.
Rhythm
30 marks
The spread of your sentence lengths. Not the average, which turned out to be useless. The spread.
Writing that clusters around one length reads as wallpaper whoever produced it. Genuine voices measured between 7.7 and 9.7 on that spread. The control sat at 4.3, and in six cases the two ranges never overlapped once.
Landing
25 marks
The proportion of sentences short enough to hit. Real voices ran between 15% and 29%. The control managed nought per cent, and it is worth sitting with that for a second: a page of entirely reasonable business English contained not one sentence of five words or fewer.
Roughly one in six is the target, placed where the argument turns.
Evidence
10 marks
Figures, prices, results, named specifics. Weighted lightest of the four because it is the easiest to fix in an afternoon, and the easiest to pad.
Seven tenths of the mark requires an actual number. Names alone will not do it, because a page can be thick with brands and places and still contain no evidence at all.
One thing is deliberately not published here: the phrasebook itself, meaning the list of terms that count against the vocabulary mark. Publish that and it becomes a list of words to avoid, which is not remotely the same thing as sounding like yourself. The measure would stop working the day it became a checklist.
Reading the number
What each band means
- 90 and above, established and distinctive. A voice worth specifying and protecting before somebody else writes on your behalf and loses it.
- 70 to 89, sound writing under category language. The voice is there. Polish rather than rebuild.
- 40 to 69, a real voice, buried. The writing is genuinely there. The category vocabulary is sitting on top of it. This is where most businesses land, and it is the band a specification fixes fastest.
- 10 to 39, little distinctive prose. Structural work first. There may not be enough writing on the page to judge.
- Below 10, no voice present. Nothing here could not have been written by anyone.
Read the breakdown before the total. Two businesses on 68 can need completely different work: one is drowning in category vocabulary and writes with a good ear, the other has original things to say in sentences that all run to the same length.
The eighteen that do not count
A measure that flatters a fake is not a test
This is the part most writing tools get wrong, and never find out about.
To build the scale, a generic control was written on purpose: fluent, tidy, professional, saying nothing. The sort of thing that comes back when you ask an assistant for website copy and accept the first draft. Then everything anyone might sensibly measure was run against it alongside six real businesses, to see which measures could tell them apart.
Several could not. Worse, several got it backwards.
- The control used more contractions than two of the three genuine voices.
- Its sentences were longer on average than all of them.
- It used no long dashes at all, exactly like every real voice measured, which quietly disposes of the internet's favourite tell.
Score any of those and you build a test that ranks the fake above the genuine article and looks confident doing it.
Eighteen further measurements are taken, shown to you, and carry nought marks. They describe writing. They do not detect anything.
Pronoun habits, colons, semicolons, question rate, hedging, how much of the page is prose rather than furniture: all of it is genuinely useful when a person is reading your writing properly, and all of it is worthless as a test. So it is reported and never scored, and it is labelled as reported so you know which is which.
What you get from that separation: a number that cannot be gamed by tidying your punctuation, and a breakdown that still tells you the interesting things about how you write.
Four honest limits
What the number is not
It is not a quality score. It has no idea whether what you said is true, useful, or worth anybody's time. A well-argued page and a confident lie can score the same. Distinctiveness is the only question it asks.
It is not an AI detector, and nothing that claims to be one works. The test cannot tell whether you typed the words, dictated them, or prompted ChatGPT for them, and it never tries. Neither can the person reading your website. A human being is perfectly capable of writing something anyone could have written, and frequently does. The tools are simply faster at it. The only question that matters commercially is whether the finished thing could have come from anybody, and that is the question being scored.
It is not a grammar or readability check. Those exist, they are cheap, and they will happily certify prose that says nothing.
It is not an opinion, and no model is consulted. The same text produces the same number today, next Tuesday, and in a year. Nothing is sent to ChatGPT or Claude at any point in the scoring, because a judgement that drifts is not a measurement, and a score you cannot reproduce is not evidence of anything.
Which is precisely why it can be used to compare two businesses. Reproducibility is the whole basis of the comparison.
Where the scale came from
Six businesses, five sectors, and one deliberate fake
Three of the six are named on this site and you can read them yourself. The other three are covered by confidentiality and are named nowhere, which is the same undertaking every client gets. All six were chosen to be awkward: two sit in adjacent markets, one was picked specifically for having weak published copy, and two were written by the same hand, to check that the method was measuring the business rather than the writer.
Those last two came out opposite. One says "I" and never "we". The other says "we" and never once says "I". Had the method been quietly measuring its author, they would have matched.
Reproducible
The same passage returns the same score every time it is run. There is a regression test that fails the build if any recorded score moves by so much as a tenth.
Recorded, not patched
One threshold has been moved since the scale was set. An early setting collapsed the entire middle band, which is most real businesses. Widening it moved one subject from 62 to 84 and left every other subject untouched. That change is written into the code with its reason, rather than quietly applied.
Honest about its own status
The breakdown is validated across six subjects. The total sits on a scale still being calibrated, and it is presented as indicative every time it appears, including in the paid reports. Anybody claiming external validation for a number like this is selling you something.
When it declines to answer
It refuses more often than you would expect
Below eighty words of continuous prose, no score is given at all. Below twenty-five sentences, rhythm and landing are flagged as noisy rather than presented as solid, because a well-written 120-word passage can legitimately contain no five-word sentence and there is no honest way to mark it down for that.
Headings, bullets, navigation and form labels are stripped before anything is counted. An early version read a brochure page as 32% short sentences when the page contained almost no sentences at all, only headings. Counting furniture measures the wrapping.
The point of all three rules is the same: a confident score on a paragraph would be exactly the sin this tool exists to name.
Comparing yourself to somebody else
Why one business can fairly be scored against another
Because nothing is interpreted. Your public pages and theirs go through identical code on the same day, on the same scale, and the subject gets no benefit of the doubt. Nobody is asked for anything, nothing is submitted, and neither party can influence the result by explaining itself.
Two things that number does not mean, and it would be dishonest not to say so on the same page.
A higher score is not a bigger business. A firm can score 51 and out-sell one scoring 78, on price, on reach, on being first in the search results, or on being genuinely better at the actual work. Distinctiveness is one advantage among several.
A gap is a diagnosis, not a verdict. The useful output of a comparison is never the ranking. It is the list of phrases all of you share, the ones that are yours alone, and the single best sentence each of you has already written and buried.
You end up knowing exactly why a buyer cannot tell the two of you apart, phrase by phrase, and which of your own habits to keep.