An AI detector can return a confident score in seconds. That score is harder to judge without knowing the writing's history. We tested two paragraphs preserved in a Wikipedia revision from before ChatGPT launched and recorded what six detectors said about each one.
How We Tested
Both excerpts appeared in an archived November 29, 2022 revision of Wikipedia's NATO article, one day before ChatGPT's public launch. That revision gives us a date independent of any detector's estimate.
We submitted each excerpt separately to ZeroGPT, GPTZero, Grammarly, Originality AI, Copyleaks, and QuillBot. Every result links to its screenshot.
The excerpts share a source and revision date. They are also similar in length, at 119 and 124 words, making their results easy to compare side by side.
Example 1: The NATO Introduction
This 119-word paragraph appeared at the start of that revision:
The North Atlantic Treaty Organization (NATO, /ˈneɪtoʊ/), also called the North Atlantic Alliance, is an intergovernmental military alliance between 30 member states – 28 European and two North American. Established in the aftermath of World War II, the organization implemented the North Atlantic Treaty, signed in Washington, D.C., on 4 April 1949. NATO is a collective security system: its independent member states agree to defend each other against attacks by third parties. During the Cold War, NATO operated as a check on the perceived threat posed by the Soviet Union. The alliance remained in place after the dissolution of the Soviet Union and has been involved in military operations in the Balkans, the Middle East, South Asia, and Africa.
| Detector | AI result | Evidence |
|---|---|---|
| ZeroGPT | 100% AI | View screenshot |
| GPTZero | 60% AI | View screenshot |
| Grammarly | 0% AI | View screenshot |
| Originality AI | ≤15% AI | View screenshot |
| Copyleaks | 0% AI | View screenshot |
| QuillBot | 24% AI | View screenshot |
ZeroGPT labeled it 100% AI. GPTZero returned 60% AI, although its verdict described the classification as uncertain. The text was already present in the November 2022 revision before ChatGPT's release.
Example 2: Another Paragraph From the Same Article
This 124-word excerpt appeared in that revision's Kosovo intervention section:
The US, the UK, and most other NATO countries opposed efforts to require the UN Security Council to approve NATO military strikes, such as the action against Serbia in 1999, while France and some others claimed that the alliance needed UN approval. The US/UK side claimed that this would undermine the authority of the alliance, and they noted that Russia and China would have exercised their Security Council vetoes to block the strike on Yugoslavia, and could do the same in future conflicts where NATO intervention was required, thus nullifying the entire potency and purpose of the organization. Recognizing the post-Cold War military environment, NATO adopted the Alliance Strategic Concept during its Washington summit in April 1999 that emphasized conflict prevention and crisis management.
| Detector | AI result | Evidence |
|---|---|---|
| ZeroGPT | 0% AI | View screenshot |
| GPTZero | 11% AI | View screenshot |
| Grammarly | 0% AI | View screenshot |
| Originality AI | ≤15% AI | View screenshot |
| Copyleaks | 0% AI | View screenshot |
| QuillBot | 0% AI | View screenshot |
All six tools gave this text low AI scores, a sharp contrast with the first paragraph from the same revision.
Why Human Writing Can Be Flagged
AI detectors are predictive models. They assess patterns in the text and estimate whether those patterns resemble AI-generated writing. They do not check a document's publication history or observe how it was written.
A classifier can learn from examples labeled as human or AI writing, then apply those learned distinctions to new text. But human and AI writing are not two perfectly separate categories of language. Both can use orderly sentences, familiar phrasing, and consistent structure. A model can therefore make a confident prediction about the wrong source.
Human writing can share patterns that a model associates with AI. When that leads a detector to label human writing as AI-generated, it is a false positive.

This is different from a plagiarism checker, which can point to a matching source. An AI detector does not find the conversation or document that created the text. Its output is an inference from the writing itself, even when an archived revision provides stronger evidence about when that writing appeared.
AI detection remains useful for identifying text worth reviewing. Its score is a prediction based on the writing, not a record of how it was created.
What the Test Shows
The first paragraph received high AI scores from two detectors, while the second scored low across all six. Both appeared in a Wikipedia revision dated one day before ChatGPT launched. The first was most likely written by a person, making those high scores likely false positives.
The takeaway is that a high AI score can be wrong and should be checked against the writing's history.
Are AI Detectors Still Reliable?
Yes. Established AI detectors are generally reliable at identifying clearly AI-generated writing, particularly when they have enough text to analyze. A study of AI-text detectors found that several commercial tools performed well across its test material, while also finding meaningful differences between products and conditions.
Accuracy does not mean perfection. A detector can correctly identify AI-written text and still produce a false positive on human writing, as this test illustrates. If you want to check writing yourself, Verva's AI detector can provide a useful signal.
FAQ
Can AI detectors flag text written before ChatGPT?
Yes. ZeroGPT labeled the first paragraph 100% AI, and GPTZero gave it a 60% AI score, even though the text appeared in a Wikipedia revision from the day before ChatGPT launched.
Why do AI detectors sometimes flag human writing?
Detectors classify writing by its patterns, not by checking who wrote it. Human prose can share patterns a model associates with AI-generated text, producing a false positive. That is why a flagged result should be checked against other evidence, such as dated revisions, drafts, or document history. Our review of universities that disabled or rejected AI detectors looks at how some institutions address this risk.
Can AI detectors accurately identify AI-generated writing?
Yes. Established detectors can identify clear, unedited AI-generated text, especially when they have enough writing to analyze. Results vary by tool and input, so a score is most useful alongside other evidence.
Can editing AI-generated text change a detector's result?
Yes. Editing can change a detector's score, but replacing a few words may not be enough to change the result. Verva's AI humanizer goes further: it rewrites AI-generated text with more natural wording, sentence structure, and flow, helping it sound human and bypass AI detectors while preserving its meaning.



