How it works

Why two detectors give a different number for the same text

Every detector measures something slightly different and was trained on different texts, so two numbers for one text are not an error — they are two different views. A score does not say what share of the text a model wrote; it says how closely the text resembles what that tool knows as machine writing. When tools disagree sharply, that is information in itself: the text is borderline and percentages will not settle it. DetekceGPT therefore shows seven detectors side by side and never merges them into one number.

Each one measures a different property

One tool counts how predictable the next word is, another tracks the variation in sentence length, a third looks for turns of phrase typical of models, and a fourth compares the text with your earlier work. These are not more or less accurate answers to one question — they are answers to four different questions.

That is why a text can come out high on one detector and low on another: it has a machine rhythm but no typical phrases, or the other way round. In DetekceGPT you see it directly — AI phrases, Machine feel and Style deviation are three separate numbers and each is read on its own.

Same signal, different scale

Even when two tools track a similar thing, each turns the result into percentages its own way. Each has its own threshold for flagging a sentence and its own set of texts it was calibrated on. Sixty percent in one tool therefore does not mean the same as sixty in another.

Language matters too. Most detectors were built on English and behave differently on other languages, so the spread between tools tends to be wider on a Czech or German text than on an English one.

What to do with the disagreement

Do not average. The mean of two numbers that measure different things is meaningless, and worse, it hides the most interesting part — that the tools disagreed. Read the sentences each detector flagged instead and check whether they point at the same places.

If they agree on the same passages, you have a stronger basis. If each points somewhere else, treat it as a sign that the text is borderline and base the decision on something else — on how far it departs from the author's earlier texts, or on a conversation about the content.

FAQ

Which detector is the right one?

None of them. Accuracy cannot be verified from outside, because you do not know the texts the tool learned from. Try it on five texts whose authorship you are sure about.

Is it worth running a text through several tools?

As a probe yes, as proof no. More numbers do not mean more certainty, only more views — and they can diverge.

Why does DetekceGPT not give one final number?

Because it would have to merge quantities that measure different properties. One number looks more certain, but it would hide the disagreement between detectors, which is the most useful part.

Try it on your own text

Paste a text or upload a file and look at the score and at the specific sentences that came out suspicious.

Run a detection

More articles