Unlike most popular AI detectors, our tool doesn’t produce a general document score. It checks every sentence, deciding the probability of it being AI-generated, and flags parts with AI traces in the report. How much AI presence is too much is up to you to judge.

“20% AI-generated” result leaves you wondering whether it is a whole machine-generated paragraph or just small edits across the document. Highlighted lines give you a conversation-starting point and clear evidence to make data-driven decisions.
How accurate is the result?
99.23% AI detection accuracy
We checked 33,699 sentences in 709 documents generated by seven frontier AI models and correctly classified 99.23% of them as AI, measured per sentence.
1.04% false positives (sentence level)
Of 280,781 sentences written by people, we misclassified 1.04% as AI.
0.00% false positives (document level)
With a sentence-by-sentence check approach, a misclassified line doesn’t condemn the whole paper. This way, at the document level, our false-positive rate drops to zero. None of the 2,805 human-written documents were wrongly classified as AI.
Why two false-positive figures
Was part of an essay generated by ChatGPT, or was it the detector catching formulaic-sounding extracts in a handful of sentences? With a general metric per document, it is hard to tell.
Occasionally, every detector mismarks the sentences that resemble robotic patterns as potentially AI. On their own, they don’t prove the document’s origin. What matters is the whole paper’s shape, the consistency of highlighted text, and its structure. Usually, an AI-generated essay won’t contain scattered machine-sounding phrases. It will have flagged whole chunks of writing, if not the majority of the text.
So, we don’t hurry into classifying the whole paper. We analyze every sentence and highlight those containing AI traces. You set the threshold for what proportion of flagged content is alarming, and we give you evidence to decide whether it crosses your limit.
- Initially, we classify every sentence. Hence, the primary metrics we provide are per sentence.
- Our competitors provide one verdict per document, and their metrics are per document. To compare like for like, we translated our sentence measure into a document-level figure.
- At the document level, our error rate drops to zero, as with our approach, one misclassified sentence doesn’t condemn the whole paper.
The threshold you set
Checking thesis papers and editing a social media post are two different tasks that require different approaches to measurement. Sentence-by-sentence analysis allows us to provide a flexible threshold rather than a fixed operating point.
You choose which setting fits your tasks today. High sensitivity is recommended for early intervention, and a more conservative setting is appropriate when the stakes are high and the consequences serious. The same scan supports both, with no re-processing needed.
Which AI models we detect
When trained to recognize one model, the detector fails to catch the other provider’s output. Here is evidence that we are not just a ChatGPT detector. The total spread across seven flagship models is 1.56 points, which means we classify them with 98.31%-99.87% accuracy.
As new models emerge and are continually upgraded, we retrain our detector to stay up to date. You can follow our recent upgrades here.
How we checked
To minimize false positives, the detector must be trained and tested on purely human-written text. AI presence in it affects the accuracy rate and test results. Here is how we avoided that.
Human text corpus from pre-AI era
Our main corpus is the British Academic Written English (BAWE) collection, the real university coursework gathered between 2004 and 2007. It is 15 years before ChatGPT was launched; hence, it physically cannot contain AI traces. BAWE contains 2,688 documents (277,033 sentences), to which we added our own internal human corpus of 117 documents (3,748 sentences). 280,781 purely human-written sentences in total were used to test our AI checker.
AI text corpus from frontier models
Seven LLM models, roughly 100 documents from each, 33,699 sentences in total. Each text used for testing is a pure AI output, with no paraphrasing or manual editing.
Per-sentence analysis
Each document submitted for checking is split into sentences. The detector classifies each of them, and the test results are counted at the sentence level. The document-level score is calculated based on sentence analysis, and the final verdict is defined by the threshold you set, from 5% to 50%.
Data transparency
Every rate we publish is reported with a 95% Wilson confidence interval. We reveal our sample size, 280,781 sentences, making the 1.04% false-positive rate figure meaningful rather than demonstrative.
Versioned and updated
AI models are being upgraded, and so are our AI detector and this report. Every figure and piece of data on this page is marked with a model version and the date when the test was run. We keep improving the tool, our test methods, and this report regularly.
How this compares
Most detectors on the market claim 98-99.98% accuracy. However, those numbers, including ours, are marketing figures, based on the vendors’ own tests.
What distinguishes our study is that we publish how many human-written content samples we tested, a metric none of the competitors disclose. This matters because AI presence in human-claimed texts distorts false-positive rate figures, making them look better than they are and reducing test precision.
Accuracy rate by itself is not representative unless it is backed by a research method and transparent study figures. This is why we publish our corpora, sample sizes, confidence intervals, and threshold curve.
What we haven’t tested yet
There are some gaps in the 2026 results we openly warn you about and plan to fill in the next tests.
- Paraphrased and humanized text. All the figures above apply to unedited AI output. We have not yet published the data for text run through tools designed to disguise AI use, so we won’t claim parity with competitors until revealing the numbers.
- Hybrid content. Mixed documents containing part human-written, part AI text are the most common real-world case and exactly what our sentence-level analysis is designed for. However, we haven’t benchmarked it yet, and this is the next test we run.
- Independent tests. All published figures are internal, not verified by independent study – yet. We are working with researchers to change it, and the offer below stays open.
Free access for researchers
We invite researchers and independent evaluators to open testing. Reach out to get free access to our AI detector, with no review of your results before publication.