Skip to content

You wrote it. A detector says you did not.

This page is written by people who build one of these tools. That is the reason to read it and also the reason to check it: we sell a detector, and we sell a humanizer, so we have an interest in you believing detection works and an interest in you believing it does not. What follows is what our own measurements say, published in full, including the parts that make our product look bad.

A score is not evidence

These tools read writing style. They do not know who typed anything, they have no record of your document, and they cannot see a draft history. What they measure is whether prose has the statistical shape that machine-written text often has — and plenty of human writing has that shape too.

On our own benchmark of 48 labelled samples, at the default threshold of 50:

  • 1 of 24 pieces of human writing were flagged as AI — a 4% false positive rate.
  • 23 of 24 pieces of AI writing were missed — a 96% false negative rate.

Read the second number again. A tool that misses 96% of the AI text put in front of it is not a tool that can tell anyone what you did. The full methodology and corpus are here.

The human sample we got wrong is worth naming, because it is the kind of writing students are taught to produce: Statistical-ecology abstract, US. Long subordinate clauses, "However"/"In these settings" transitions. It scored 50. Formal, careful, evenly-paced prose is the writing our detector is most likely to mistake for a machine's — which means the better your academic register, the more exposed you are. We ranked every register in the benchmark by the score it got, and a human academic abstract sits above four of the five kinds of AI writing in it.

Different detectors disagree about the same paragraph

There is no single answer that these tools converge on. Each vendor trains on different text, sets its threshold in a different place, and publishes a different number for its own accuracy. The same paragraph can come back as a confident accusation from one and "likely human" from another, and neither result is checkable from outside: their terms of service forbid the automated testing that would let anyone — including us — measure them independently.

That is worth saying plainly to whoever raised the case. A number from one vendor is one vendor's opinion, produced by a method nobody outside that company has been allowed to audit.

What actually settles it: how the draft was written

A detector looks at the finished text. Your defence is everything the finished text cannot show — and most of it already exists without your having done anything.

  • Version history. Google Docs keeps a complete revision timeline under File → Version history → See version history. Word does the same through OneDrive or SharePoint. It shows the document growing over hours or days, with the false starts and the paragraph you deleted and rewrote.
  • Your notes and sources. Reading notes, a photographed library book, a highlighted PDF, the tab you kept open. Work that came from somewhere leaves a trail behind it.
  • The messy middle. Earlier drafts, the outline, the paragraph order you changed your mind about. Nobody who pasted a finished essay has these.
  • Ask to talk about the content. Someone who wrote an argument can explain why they chose it, what they rejected, and where it is weakest. This is the oldest and still the most reliable check there is, and it is one no detector can perform.

If you are reading this before an accusation rather than after one: the single most useful habit is to write in a tool that keeps history, and to leave that history switched on.

What we would not do

We sell a humanizer. We are not going to suggest you run an accusation through it. Rewriting work you have already submitted does not answer the question that was asked, it destroys the drafting evidence that would have answered it, and if it were discovered it would convert a disputed score into something much harder to explain.

We also make no claim that our humanizer defeats any third-party detector, here or anywhere else. We have never been able to measure that — testing against those tools is prohibited by their terms — so anyone quoting you a pass rate, us included, would be guessing.

If you want to see what a detector sees

Ours names the specific phrases it reacted to and why — hedging filler, a formulaic connector, a stock opener — rather than handing you a number with nothing behind it. It is free, unlimited, and needs no account. It will not prove you wrote something. Nothing of this kind will.