AI Detector / Accuracy
How Accurate Is This AI Detector?
Every detector is sometimes wrong. Here is exactly how often we are, and where.
4%
False positive rate
1 of 24 human samples flagged as AI
96%
False negative rate
23 of 24 AI samples passed as human
-1.1
Separation
Mean AI score 21.0 minus mean human score 22.0
48
Corpus size
Labelled adversarial samples, threshold 50
Measured on our adversarial benchmark · last regenerated 2026-08-26 · updated with every scoring change
Read these numbers for what they are. 48 samples is a small corpus. We publish it anyway because a small measured number beats a large invented one — but a single additional false positive would move that 4% by 4 points. These are lab-benchmark results on deliberately hostile text, not a guarantee about your document.
What this tool can and cannot tell you
It can tell you this
Whether your draft still reads as machine-written: the stock transitions, the uniform sentence rhythm, the em-dash habit, the tidy rule-of-three lists. On our benchmark it did not flag a single one of the 5 texts written by non-native English speakers, which is the accusation we care most about never making.
It cannot tell you this
Whether a person wrote it. Ask any model to write casually, or to vary its sentence lengths, and the tells disappear along with our ability to see them — every such sample in our benchmark scored in single digits. Treat a low score as "no obvious tells", never as "a human wrote this".
This is a limit of the method, not a setting we can turn up. Our engine is statistical and runs on our own servers with no language model involved — that is what makes it free and unlimited, and it is also why it sees style and nothing else.
Where it holds and where it fails
The two headline rates average over very different situations. Broken out by the kind of writing, from a corpus of 24 human texts published before 2020 and 24 AI texts with recorded model and prompt:
| Kind of writing | Samples | Mean score | We got wrong |
|---|---|---|---|
| Academic abstract (human) | 4 | 35.8 | 1 of 4 wrongly flagged |
| Esl nonnative (human) | 5 | 25 | 0 of 5 — none accused |
| Government bureaucratic (human) | 4 | 19.8 | 0 of 4 — none accused |
| Technical documentation (human) | 4 | 21.8 | 0 of 4 — none accused |
| Casual first person (human) | 4 | 12.8 | 0 of 4 — none accused |
| Business marketing (human) | 3 | 14.7 | 0 of 3 — none accused |
| Instructed casual (AI) | 5 | 9 | 5 of 5 missed |
| Instructed varied rhythm (AI) | 5 | 9.4 | 5 of 5 missed |
| Humanizer output (AI) | 3 | 9.7 | 3 of 3 missed |
| Domain specific (AI) | 4 | 19.8 | 4 of 4 missed |
| Default assistant (AI) | 7 | 43.3 | 6 of 7 missed |
"Humanizer output" is text from our own AI Humanizer. Our detector does not catch it. We would rather you learn that here than discover it yourself.
Use these numbers
Everything here is free to quote, reproduce and check, including the parts that make us look bad — a number nobody can verify is worth nothing. The raw benchmark output is one file, and the corpus records every sample's source and licence so you can re-run it against your own detector.
Raw data
detector-accuracy.json — every sample, its score, our call, and whether we got it right. Regenerated with each scoring change.
Citation
Coda One, AI Detector Accuracy Report, 2026-08-26. codaone.ai/ai-detector/accuracy — CC BY 4.0.
If you re-run it and get different results, we want to know — that is the only way a published number stays worth publishing. Tell us what you found.
What is wrong with our own corpus
Every benchmark has a shape that flatters or punishes it. These are the limitations recorded by whoever built the corpus, reproduced as written rather than summarised by us:
How we know the human samples are human
Every sample was published between 2010 and 2019, i.e. before any general-purpose LLM was publicly available. Human authorship here is a date fact, not a stylistic judgement. Each sample records how the date was verified, preferring identifiers that cannot be back-dated: DOIs (checked against Crossref), PMC ids and Europe PMC firstPublicationDate, Federal Register document numbers, MediaWiki revision ids, git tags, and Stack Exchange server-assigned creation timestamps.
On the non-native English samples
READ THIS BEFORE PUBLISHING ANY ESL NUMBER. The five "esl-nonnative" samples are NOT learner-corpus essays. No learner corpus with a license compatible with commercial use could be sourced: PELIC is CC BY-NC-SA (NC excludes us), the Cambridge Learner Corpus FCE set, ICLE, ICNALE, NUCLE and the BEA-2019 W&I+LOCNESS data are all behind restrictive or registration-gated licenses, and the ETS TOEFL11 corpus is a paid LDC product. WHO and FAO publications were also rejected as CC BY-NC-SA. Rather than substitute something unlicensed, these five are CC BY / CC BY-SA scholarly prose written by authors whose institutional affiliations are all in non-anglophone countries (Indonesia, Turkey, Iran, China), selected because visible L2-transfer features survived peer review uncorrected. That is evidence of non-native authorship, not proof of any individual author's first language, and the register is academic rather than the undergraduate essay that produces our worst false positives. Treat any FPR computed on this slice as a floor, not as the ESL-essay false-positive rate.
Which models the AI samples came from
Read this before quoting any FNR computed from this file. WHICH MODELS ARE ACTUALLY IN HERE (24 samples): - claude-opus-5 (Anthropic), generated 2026-08-26 for this benchmark: 18 samples. - OpenAI gpt-4o-mini, as the engine of the Codaone production humanizer, rewriting claude-opus-5 text: 3 samples. Mixed lineage — Claude wrote it, GPT rewrote it. - OpenAI ChatGPT, December 2022 web release (gpt-3.5 era), via the HC3 dataset: 3 samples. WHICH ARE NOT: Google Gemini. Meta Llama. Mistral. DeepSeek. Qwen. xAI Grok. Cohere. Every current-generation OpenAI model (GPT-4o, GPT-5 family) except gpt-4o-mini in its humanizer role. Every current-generation Claude except opus-5. Nothing from a consumer "undetectable AI" service (StealthGPT, Undetectable.ai, Phrasly) — those are the strongest evasion tools in the wild and we have zero samples of their output. THEREFORE: an FNR computed over this corpus is an FNR *for Claude Opus 5 output, for our own humanizer's output, and for three-year-old ChatGPT output*. It is not an estimate of how often we miss AI text in general, and it must not be published as one. If the number is quoted on /ai-detector/accuracy, this sentence has to travel with it. A SECOND BIAS, LESS OBVIOUS: 18 of the samples were written by the same model that assembled this corpus, while it knew it was building a detector benchmark. That is a demand-characteristic risk in both directions — the model may have written more stereotypically "AI" prose for the default-register samples, and may have tried harder than a naive user would on the evasion samples. The prompts are recorded verbatim so a third party can rerun them on any model and check. Doing that on a non-Claude model is the cheapest available fix for both biases. THIRD, THE REGISTER MIX IS NOT A USAGE DISTRIBUTION: {"default-assistant":7,"instructed-casual":5,"instructed-varied-rhythm":5,"humanizer-output":3,"domain-specific":4}. It is deliberately weighted toward evasion (13 of 24 samples are instructed-casual, instructed-varied-rhythm, or humanizer output) because §1.4 of AI-DETECTOR-COMPETITIVE-2026-08-25 measured that all of our false negatives come from disguised AI. A corpus weighted this way will report a WORSE FNR than a corpus of default-register text would. That is intentional and honest; it is not comparable to a competitor's headline accuracy number, which is measured on whatever mix flatters them. FOURTH, ON THE HUMANIZER SAMPLES: the free anonymous quota is 3/day/IP and all three were used on 2026-08-26, so there are exactly three and no reruns. They are the highest-value rows in the file — a false negative on any of them means our own paid product defeats our own free product, which is the central question the accuracy page exists to answer honestly.
On how few voices some registers have
Registers are unevenly hard to source under an open license. Government and technical-documentation samples come from a small number of institutional voices (three Federal Register rules, two Kubernetes docs pages, three English Wikipedia revisions), so within-register variance understates the real world.
What each error costs, and to whom
False positive — the expensive error
A false positive accuses a human of using AI. In a classroom that is an integrity case; for a freelancer it is a withheld payment; for an ESL writer it is both, plus the suggestion that their English is not their own. This error has a victim.
False negative — the cheap error
A false negative lets AI text pass as human. Nobody is wrongly accused, but the tool failed to do the job you came for. Ours is 96% — the number is that bad because our engine reads style, and AI that has been asked to avoid AI-sounding style leaves nothing for it to read. We would rather publish that sentence than a number that flatters us.
Every sample, every score
The full benchmark run behind the numbers above — including the ones we got wrong. Scores at or above 50 count as an AI call.
| Truth | Score | Our call | Sample |
|---|---|---|---|
| Human | 24 | correct | esl-nonnative: Indonesian authors, education research abstract. L1-transfer markers: "The subject of this research are", "the interaction that occurs in students", "This interaction is built with heterogeneous student skills." Katarina Tri Utaminingtyas, Rachmadina Eka Herdianti, Inti Hayatul Fitria & Anton Prayitno, "Small Groups: Student Productive Interactions in Learning Cooperative (Case Study of Mathematics Learning at Junior High School in Pakis, Malang)", Educational Process: International Journal 6(2), 2017 — https://doi.org/10.22521/edupij.2017.62.3 |
| Human | 15 | correct | esl-nonnative: Turkish nursing academics, structured abstract. L2 markers: "The scale was carried out to the students", "While analyzing the research data;", "Fisher' Exact". Ebru Özen Bekar, Dilek Konuk Şener, Çetin Yılmaz & Şengül Cangür, "The Evaluation of Professional Self-esteem of Nurses and Social Workers Before and After Graduation", Sağlık ve Hemşirelik Yönetimi Dergisi (Journal of Health and Nursing Management) 4(3), 2017 — https://doi.org/10.5222/SHYD.2017.050 |
| Human | 35 | correct | esl-nonnative: Iranian-Persian L1 authors at Turkish/Iranian institutions; discursive, argumentative register — the closest thing in this set to a student essay. Sahar Pouya & Homa Irani Behbahani, "Landscape visual assessment: A case of Iran-Iraq war memorial garden", Turkish Journal of Forestry / Türkiye Ormancılık Dergisi 18(4), 2017 — https://doi.org/10.18182/tjf.294916 |
| Human | 9 | correct | esl-nonnative: Iranian clinical-trial abstract. Article omission throughout ("in treatment of patients", "in case they had"), a classic L1-Persian marker. Shakiba M, Moazen-Zadeh E, Noorbala AA, Jafarinia M, Divsalar P, Kashani L, "Saffron (Crocus sativus) versus duloxetine for treatment of patients with fibromyalgia: A randomized double-blind clinical trial", Avicenna Journal of Phytomedicine, 2018 — https://europepmc.org/article/PMC/PMC6235666 |
| Human | 42 | correct | esl-nonnative: Chinese-hospital author team, medical case series. Semicolon-splice and "which may demand an additional approach to the ongoing practice" are non-native constructions that survived peer review. Pandey S, Li L, Deng XY, Cui DM, Gao L, "Outcome Following the Treatment of Ventriculitis Caused by Multi/Extensive Drug Resistance Gram Negative Bacilli; Acinetobacter baumannii and Klebsiella pneumonia", Frontiers in Neurology 9:1174, 2019 — https://doi.org/10.3389/fneur.2018.01174 |
| Human | 42 | correct | academic-abstract: Review abstract, UK. Uniform long sentences, heavy formal connectors ("Furthermore", "Overall") — the canonical AI-lookalike register. Barton AJ, Hill J, Pollard AJ, Blohmke CJ, "Transcriptomics in Human Challenge Models", Frontiers in Immunology 8:1839, 2017 (Oxford Vaccine Group, University of Oxford) — https://doi.org/10.3389/fimmu.2017.01839 |
| Human | 9 | correct | academic-abstract: Structured systematic-review abstract with labelled sections. Extremely uniform sentence length and near-zero first-person voice. Cairns AE, Pealing L, Duffy JMN, Roberts N, Tucker KL, Leeson P, "Postpartum management of hypertensive disorders of pregnancy: a systematic review", BMJ Open 7(11):e018696, 2017 (University of Oxford) — https://doi.org/10.1136/bmjopen-2017-018696 |
| Human | 42 | correct | academic-abstract: Plant-genetics abstract, US. Dense nominalisation and hedged claims; the kind of prose humans write that scores as "too uniform". PLOS ONE 13(12):e0207723, 2018 — MSU-DOE Plant Research Laboratory, Michigan State University — https://doi.org/10.1371/journal.pone.0207723 |
| Human | 50 | false positive | academic-abstract: Statistical-ecology abstract, US. Long subordinate clauses, "However"/"In these settings" transitions. PLOS ONE 13(12):e0204150, 2018 — Department of Forestry, Michigan State University — https://doi.org/10.1371/journal.pone.0204150 |
| Human | 30 | correct | government-bureaucratic: FAA final rule summary. "Additionally", "Finally", "These actions are necessary to" — the exact connector profile our detector punishes. Federal Aviation Administration, "Regulatory Relief: Aviation Training Devices; Pilot Certification, Training, and Pilot Schools; and Other Provisions", final rule, Federal Register document 2018-12800 — https://www.federalregister.gov/documents/2018/06/27/2018-12800 |
| Human | 21 | correct | government-bureaucratic: FDA final rule summary — one 100-word sentence built out of semicolon-separated clauses. Human, and maximally machine-like. Food and Drug Administration, "Food Labeling: Revision of the Nutrition and Supplement Facts Labels", final rule, Federal Register document 2016-11867 — https://www.federalregister.gov/documents/2016/05/27/2016-11867 |
| Human | 15 | correct | government-bureaucratic: Department of Education interim final rule. Statutory cross-references and self-referential procedural language. Department of Education, "Student Assistance General Provisions, Federal Perkins Loan Program, Federal Family Education Loan Program, William D. Ford Federal Direct Loan Program, and Teacher Education Assistance for College and Higher Education Grant Program", interim final rule, Federal Register document 2017-22851 — https://www.federalregister.gov/documents/2017/10/24/2017-22851 |
| Human | 13 | correct | government-bureaucratic: CDC surveillance report opening. Numbered citations, passive voice, agency-as-actor ("CDC analyzed data from..."). Cree RA, Bitsko RH, Robinson LR, Holbrook JR, Danielson ML, Smith C, et al., "Health Care, Family, and Community Factors Associated with Mental, Behavioral, and Developmental Disorders and Poverty Among Children Aged 2-8 Years — United States, 2016", MMWR Morbidity and Mortality Weekly Report 67(50), 2018 — https://europepmc.org/article/PMC/PMC6342550 |
| Human | 15 | correct | technical-documentation: Encyclopedic network-protocol description. Flat declaratives, list-like enumeration, zero first person. Wikipedia contributors, "Transmission Control Protocol", English Wikipedia, revision 843439827 (2018-05-29T05:04:10Z) — https://en.wikipedia.org/w/index.php?title=Transmission_Control_Protocol&oldid=843439827 |
| Human | 42 | correct | technical-documentation: Cryptography explainer. Long conditional sentences and "For this to work it must be" constructions. Wikipedia contributors, "Public-key cryptography", English Wikipedia, revision 736786168 (2016-08-29T20:54:56Z) — https://en.wikipedia.org/w/index.php?title=Public-key_cryptography&oldid=736786168 |
| Human | 18 | correct | technical-documentation: Product documentation, procedural voice. Repetitive parallel clauses ("does not kill... does not create...") read as templated. Kubernetes documentation, "Deployments" (content/en/docs/concepts/workloads/controllers/deployment.md), kubernetes/website at tag release-1.16 — https://github.com/kubernetes/website/blob/release-1.16/content/en/docs/concepts/workloads/controllers/deployment.md |
| Human | 12 | correct | technical-documentation: Networking internals documentation. Comparative technical claims with hedging, written by contributors for whom English varies. Kubernetes documentation, "Service" (content/en/docs/concepts/services-networking/service.md), kubernetes/website at tag release-1.16 — https://github.com/kubernetes/website/blob/release-1.16/content/en/docs/concepts/services-networking/service.md |
| Human | 16 | correct | casual-first-person: Blunt forum answer. Sentence fragments, an em dash, a typo ("polystychrene"), and a joke — high burstiness. Stack Exchange user "Criggie", answer 51225 on Bicycles Stack Exchange, 2017-12-01 — https://bicycles.stackexchange.com/a/51225 |
| Human | 9 | correct | casual-first-person: Advice post with imperatives, a parenthetical aside and an editorialising last line. Wildly uneven sentence lengths. Stack Exchange user "keshlam", answer 82857 on The Workplace Stack Exchange, 2017-01-12 — https://workplace.stackexchange.com/a/82857 |
| Human | 12 | correct | casual-first-person: Reflective first-person explanation of social convention, with quoted dialogue and a self-deprecating closing parenthesis. Stack Exchange user "Max", answer 38180 on Travel Stack Exchange, 2014-11-03 — https://travel.stackexchange.com/a/38180 |
| Human | 14 | correct | casual-first-person: Wikipedia talk-page comment. Enumerated grievance, slang ("crabon"), signed and timestamped by the editor. Wikipedia editor Dennis Bratland, comment dated 22:30, 20 July 2014 (UTC) on Talk:Bicycle; captured in English Wikipedia revision 664757797 (2015-05-30T21:01:42Z) — https://en.wikipedia.org/w/index.php?title=Talk%3ABicycle&oldid=664757797 |
| Human | 17 | correct | business-marketing: Enforcement press release. Announcement-lede structure, superlatives ("record", "by far the largest"), third-person institutional voice. Federal Trade Commission, press release "Google and YouTube Will Pay Record $170 Million for Alleged Violations of Children's Privacy Law", September 4, 2019 — https://www.ftc.gov/news-events/news/press-releases/2019/09/google-youtube-will-pay-record-170-million-alleged-violations-childrens-privacy-law |
| Human | 16 | correct | business-marketing: Press release with an executive quote — the register that reads most like generated corporate copy. Federal Trade Commission, press release "Uber Agrees to Expanded Settlement with FTC Related to Privacy, Security Claims", April 12, 2018 (quote from Acting FTC Chairman Maureen K. Ohlhausen) — https://www.ftc.gov/news-events/news/press-releases/2018/04/uber-agrees-expanded-settlement-ftc-related-privacy-security-claims |
| Human | 11 | correct | business-marketing: Product release announcement. Feature-benefit sentences and capability claims — vendor marketing prose written by engineers. Kubernetes 1.16 Release Team, "Kubernetes 1.16: Custom Resources, Overhauled Metrics, and Volume Extensions", Kubernetes blog, 2019-09-18 — https://kubernetes.io/blog/2019/09/18/kubernetes-1-16-release-announcement/ |
| AI | 42 | missed — passed as human | default-assistant: Classic essay register: abstract subject, "Furthermore", uniform long sentences. The easy case — if we miss this we have nothing. Also the input to humanizer sample #1. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 23 | missed — passed as human | default-assistant: Default register with one em-dash. Included deliberately: commit ba4f0aef demoted em-dashes from conviction evidence to corroboration, and this sample is the regression guard for that decision on the AI side. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 70 | correct | default-assistant: SEO-blog intro register — the highest-volume real-world use of an LLM and the text most likely to be pasted into a detector by an editor. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 42 | missed — passed as human | default-assistant: Default register, informational/advisory. Also the input to humanizer sample #2, so the pair isolates what the humanizer actually changes. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 9 | missed — passed as human | instructed-casual: Texting register: lowercase, no terminal punctuation on the last line, fragments. Direct replication of the §1.4 miss that scored 15. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 9 | missed — passed as human | instructed-casual: Reddit-comment imitation — the §1.4 miss that scored 9, and the single cheapest evasion a real user can perform (one sentence of prompt). claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 9 | missed — passed as human | instructed-casual: Fake user review in a chatty voice — the commercial evasion case (review farms), and a register where sentence-capitalization is preserved so the detector cannot key on lowercase alone. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 9 | missed — passed as human | instructed-casual: Self-deprecating personal confession — tests whether emotional first-person content plus low lexical formality is enough to push the score under threshold. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 9 | missed — passed as human | instructed-casual: Personal-blog voice with the transition words explicitly banned in the prompt — isolates how much of our AI signal is carried by "moreover / additionally" alone. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 9 | missed — passed as human | instructed-varied-rhythm: Extreme length variance (2-word sentences next to 45-word sentences). This is the direct attack on a CV-of-sentence-length feature. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 9 | missed — passed as human | instructed-varied-rhythm: Varied rhythm carrying an anecdote with reported speech — the register a student actually gets when they ask for "a personal essay that does not sound like AI". claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 9 | missed — passed as human | instructed-varied-rhythm: Argumentative op-ed with varied rhythm and a concrete verifiable claim (Buffalo 2017). Tests whether specific facts and dates read as human to us. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 11 | missed — passed as human | instructed-varied-rhythm: Literary reflective memoir with varied rhythm — sits directly on top of the human Woolf/Fitzgerald samples in the human half. If this scores like they do, the score is not separating authorship, it is separating genre. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 9 | missed — passed as human | instructed-varied-rhythm: Technical explainer written with varied rhythm and a second-person analogy — the hardest combination for us, because it is also exactly how a good human technical writer writes. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 9 | missed — passed as human | humanizer-output: Our production humanizer applied to default-assistant sample #1 (remote work). API metrics on the call: aiScoreBefore 24, aiScoreAfterPass1 26, aiScoreAfter 8, iterations 2, 120→141 words. Note the humanizer's own estimator already scored the untouched Claude essay at only 24 — that estimator disagreeing with detect.js is its own finding. Two-stage pipeline. Stage 1: claude-opus-5 (Anthropic) produced the input text (default-assistant sample #1, remote-work essay) on 2026-08-26. Stage 2: that text was POSTed verbatim to the Codaone production humanizer (POST https://www.codaone.ai/api/tools/humanize, Origin: https://www.codaone.ai, body {text, mode:"standard"}, anonymous/free plan) on 2026-08-26. The API reported model "gpt-4o-mini" (OpenAI), fallback:false. The value of the response's "humanized" field is stored below verbatim, including its paragraph breaks. — https://www.codaone.ai/api/tools/humanize |
| AI | 9 | missed — passed as human | humanizer-output: Our production humanizer applied to default-assistant sample #4 (small-business cybersecurity). API metrics: aiScoreBefore 24, aiScoreAfterPass1 18, aiScoreAfter 14, iterations 2, 114→152 words. Two-stage pipeline. Stage 1: claude-opus-5 (Anthropic) produced the input text (default-assistant sample #4, small-business cybersecurity paragraph) on 2026-08-26. Stage 2: that text was POSTed verbatim to the Codaone production humanizer (POST https://www.codaone.ai/api/tools/humanize, Origin: https://www.codaone.ai, body {text, mode:"standard"}, anonymous/free plan) on 2026-08-26. The API reported model "gpt-4o-mini" (OpenAI), fallback:false. The value of the response's "humanized" field is stored below verbatim, including its paragraph breaks. — https://www.codaone.ai/api/tools/humanize |
| AI | 11 | missed — passed as human | humanizer-output: Our production humanizer applied to a formal academic abstract (domain-specific sample #1, lightly shortened to fit the free word budget). API metrics: aiScoreBefore 16, aiScoreAfter 20, iterations 1, 123→141 words. This is the one call where our estimator scored the output HIGHER than the input (16→20) — the humanizer made formal text read more like generic AI, not less. Two-stage pipeline. Stage 1: claude-opus-5 (Anthropic) produced the input text (domain-specific sample #1, an academic abstract, with the "(n = 4,812)" parenthetical and the final clause about student effort removed so the input fit the 300-word free budget cleanly) on 2026-08-26. Stage 2: that text was POSTed verbatim to the Codaone production humanizer (POST https://www.codaone.ai/api/tools/humanize, Origin: https://www.codaone.ai, body {text, mode:"standard"}, anonymous/free plan) on 2026-08-26. The API reported model "gpt-4o-mini" (OpenAI), fallback:false. The value of the response's "humanized" field is stored below verbatim, including its paragraph breaks. — https://www.codaone.ai/api/tools/humanize |
| AI | 9 | missed — passed as human | domain-specific: AI-written empirical abstract. Pairs with the false-positive side: if human academic prose scores high AND this scores high, the feature is "academic", not "AI". Also the stage-1 input for humanizer sample #3. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 42 | missed — passed as human | domain-specific: AI-written contract-law analysis. Direct counterpart to the human Marbury v. Madison sample in the human half — same register, opposite label. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 14 | missed — passed as human | domain-specific: AI-written API reference. Counterpart to the human Wright brothers patent specification in the human half: both are uniform procedural prose with near-zero sentence-length variance, which is precisely where a variance-based score has no information. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 14 | missed — passed as human | domain-specific: AI-written internal corporate memo. The register a real employee most plausibly delegates to an LLM, and the register a manager most plausibly runs through a detector. claude-opus-5 (Anthropic), generated 2026-08-26. The model was run as a Claude Code agent tasked with building this corpus; it issued the recorded prompt to itself and wrote the completion in the same session. Output copied verbatim into this file with no human drafting, editing, or trimming. |
| AI | 42 | missed — passed as human | default-assistant: CROSS-FAMILY CONTROL (OpenAI, not Claude). Encyclopedic explainer register. Three years older than our own samples, so it also probes whether we only detect current-generation phrasing. OpenAI ChatGPT, December 2022 web release (gpt-3.5 era; exact checkpoint not disclosed by OpenAI). Collected by Guo et al., "How Close is ChatGPT to Human Experts?" (arXiv:2301.07597) into the HC3 dataset, config "wiki_csai", row_idx 11 (record id 11), field chatgpt_answers[0]. Retrieved 2026-08-26 via the Hugging Face datasets-server rows API. — https://huggingface.co/datasets/Hello-SimpleAI/HC3 |
| AI | 42 | missed — passed as human | default-assistant: CROSS-FAMILY CONTROL (OpenAI, not Claude). Financial explainer, domain register. Contains a real generation artifact — a missing space at "potential investors.Stock splits" — which is preserved verbatim rather than cleaned up. OpenAI ChatGPT, December 2022 web release (gpt-3.5 era; exact checkpoint not disclosed by OpenAI). Collected by Guo et al., "How Close is ChatGPT to Human Experts?" (arXiv:2301.07597) into the HC3 dataset, config "finance", row_idx 15 (record id 15), field chatgpt_answers[0]. Retrieved 2026-08-26 via the Hugging Face datasets-server rows API. — https://huggingface.co/datasets/Hello-SimpleAI/HC3 |
| AI | 42 | missed — passed as human | default-assistant: CROSS-FAMILY CONTROL (OpenAI, not Claude). ELI5 prompt, but note the answer is still in default assistant register — the casual PROMPT did not produce a casual REGISTER. That contrast with our instructed-casual block is the point of including it. OpenAI ChatGPT, December 2022 web release (gpt-3.5 era; exact checkpoint not disclosed by OpenAI). Collected by Guo et al., "How Close is ChatGPT to Human Experts?" (arXiv:2301.07597) into the HC3 dataset, config "reddit_eli5", row_idx 10 (record id 10), field chatgpt_answers[0]. Retrieved 2026-08-26 via the Hugging Face datasets-server rows API. — https://huggingface.co/datasets/Hello-SimpleAI/HC3 |
Where we know we get it wrong
Seven categories of text where this detector — and, in most cases, every statistical detector — systematically errs. For each: what happens, why, and what we do about it.
ESL and non-native English writing
False positive riskWhy it fails: Statistical detection reads simpler vocabulary and regular sentence rhythm as machine-like. This is the best-documented bias in the entire detector category, and it lands on the people with the most at stake. On our internal adversarial set, a real ESL student essay scored 42.
What we do: Mid-range scores get a hedged verdict ("Mixed or AI-assisted", "Possibly AI-generated") and low confidence — never a flat accusation. If you write in English as a second language and got flagged, that context outweighs our number.
Formal, legal, and corporate human prose
False positive riskWhy it fails: Connectors like "Furthermore" and "In conclusion" are also the strongest LLM tells, so a regulation, a contract, or a press release can out-score actual AI text on the pattern model. On our internal set, a human-written regulatory passage scored 68 and a 2019 press release scored 62.
What we do: In this band the verdict language stays at "Possibly" with low confidence, and the per-sentence view shows exactly which phrases drove the score — so a human can judge whether they came from a lawyer or a language model.
Uniform technical documentation
False positive riskWhy it fails: Good documentation is deliberately uniform: same sentence shape, same structure, restrained vocabulary. A variance-based model reads that discipline as generation. A human-written scheduler doc scored 50 ("Mixed") on our internal set.
What we do: Uniform prose caps out at hedged verdicts with low confidence. If you are checking documentation, weigh the per-sentence breakdown over the headline score.
Punctuation-styled writing (em-dash-heavy prose)
False positive risk (fixed, still watched)Why it fails: We used to over-weight punctuation habits as an AI tell: an em-dash-heavy human essay scored 80 on an earlier version of the scoring. Plenty of humans write like that; so do LLMs.
What we do: We downgraded punctuation from primary evidence to corroboration. The same essay now scores 13 on the benchmark below, and it stays in the corpus permanently so a regression here fails the benchmark before it ships.
AI written to sound casual or literary
False negative riskWhy it fails: Prompt an LLM to write casually — or run its output through a paraphraser — and the formal tells our pattern model keys on are stripped away. Both disguised-AI samples in the benchmark below slipped through (scored 45 and 9). This is the main way to beat us, and every other statistical detector.
What we do: Mostly, we cannot fix this within our approach, and we say so. Treat a low score on text you already suspect as weak evidence of anything. Uneven per-sentence scores on a polished piece are sometimes the residue of partial rewriting — a lead worth following, not a verdict.
Short samples
Unreliable both waysWhy it fails: Statistics computed over a handful of sentences are close to noise. Sentence-length variance, vocabulary richness, and starter diversity all need material to measure.
What we do: Under 30 words we refuse to give a verdict at all. Just above that line, read the score as a hint, not a finding.
Non-English text
No meaningful signalWhy it fails: The scoring is built for English. Other Latin-script languages get degraded analysis; Chinese, Japanese, and Korean text does not get meaningful statistical analysis at all.
What we do: Do not use this tool on non-English text. A score produced there is not evidence of anything, and we would rather tell you that here than let a number imply otherwise.
Methodology
Where the samples come from
Every human sample is genuine, attributable human prose — public-domain works with the source printed next to each result above (Darwin 1859, Marshall's Marbury v. Madison opinion 1803, the Wright brothers' 1906 patent, Twain 1869, Woolf 1915, Fitzgerald 1925), chosen to cover the registers detectors most often get wrong. Every AI sample is genuinely machine-generated, with the generating model and prompt intent recorded. A benchmark whose "human" text was written by a machine would measure nothing; ours is checkable line by line.
What runs
The benchmark feeds every sample through the same scoring engine that answers the live tool's detection API. There is no separate "benchmark mode" — what we measure is what you get.
The corpus
48 labelled English samples, adversarial in both directions: human writing that looks AI-ish to a statistical model (an academic abstract, legal prose, technical documentation, em-dash-heavy essay writing) and AI text prompted to sound casual and human. Uniform-but-human prose is over-represented on purpose, because that is exactly what a variance-based detector gets wrong. Each sample records why it is in the set, so a regression tells us which property broke — not just which string.
The threshold
Scores at or above 50 count as an AI call. False positive rate is the share of human samples at or over that line; false negative rate is the share of AI samples under it. Separation — mean AI score minus mean human score — matters more than either: if it is small, the score is noise no matter where the threshold sits.
Regeneration
Every change to the scoring reruns this benchmark, and the run rewrites the data file this page is built from. The numbers above cannot drift away from the code, because they are produced by it.
The promise
We will never quietly change these numbers. If a scoring change makes them worse, this page publishes worse numbers. The benchmark data lives in version control next to the scoring code, so the history of every published figure is inspectable.
Why publish this at all
AI detectors are evidence, not proof. The measurable signals in a piece of text — sentence rhythm, vocabulary spread, phrase patterns — overlap between careful human writers and language models, and no amount of engineering makes that overlap disappear. Anyone claiming 99% accuracy is selling certainty they cannot deliver, because the certainty does not exist to sell.
We would rather show you a small, honest benchmark than a big, unverifiable claim. The numbers on this page are less flattering than the ones you will see advertised elsewhere, and that is precisely why you can trust them: they were measured on hostile inputs by the same code that scores your text, and they update automatically whether they improve or not.
Use the detector the way we use it ourselves: as one signal among several, weakest exactly where this page says it is weakest, and never as the sole basis for a decision about a person.
Frequently Asked Questions
How accurate is this AI detector really?
Can an AI detector prove someone used AI?
Why publish your error rates at all?
Will these numbers change?
More AI Tools: AI Humanizer · AI Rewriter · AI Summarizer · Plagiarism Checker