What AI checkers actually do, and why they often get it wrong
An AI checker is a tool that analyzes text and tries to guess whether a human or an AI model wrote it. Most work by looking for statistical patterns — word choice, sentence length, repetition, how ideas connect — and comparing them to what they've learned about AI-generated text. The honest answer is that none of them are reliable, and some are barely better than a coin flip.
The core problem is that AI writing and human writing overlap too much. A human can write in a simple, repetitive style. An AI can write in a complex, varied style. A tool trained on one AI model (like GPT-3) may fail completely on text from a different model (like Claude). And as AI models improve, they get better at mimicking human patterns, which makes detection harder over time.
If you're trying to figure out whether something was AI-written because you need to know for a specific reason — a school assignment, a job application, content you're publishing — the tool result alone won't be enough. You'll need to think about the context and read the text yourself.
Key Takeaways
- AI checkers look for statistical patterns in writing, but those patterns overlap significantly between human and AI text, making them inherently unreliable.
- Different AI models produce different writing styles, so a tool trained on one model may fail on text from another.
- A tool that says text is 85% AI-written might be wrong — the confidence score is not a may provide, and the tool itself may have been trained on limited data.
- Reading the text yourself and checking facts is more useful than relying on a checker's score alone.
- If you need to enforce a policy against AI writing, you'll need human review, not just a tool result.
How these tools try to detect AI writing
Most AI checkers use one of two approaches. The first is statistical analysis: they measure things like average word length, sentence length, how often certain words repeat, and how predictable the next word is. AI models often produce text with lower entropy (more predictable word choices) and less variation in sentence structure than humans do.
The second approach is machine learning classification: the tool is trained on a dataset of known human text and known AI text, then learns to recognize patterns that separate the two. Tools like GPTZero, Originality.AI, and Turnitin's AI detection all use versions of this method.
The problem with both approaches is that they're trained on specific datasets from specific time periods. If the tool was trained mostly on GPT-3 output, it may not recognize GPT-4 or Claude. If it was trained on formal writing, it may flag casual human writing as AI. And as AI models get better at mimicking human patterns, the patterns the tool learned become outdated.
Why these tools fail, and how often
Independent testing has found that popular AI checkers produce false positives (saying human text is AI) and false negatives (missing actual AI text) at rates that make them unreliable for high-stakes decisions. A 2024 study by researchers at UC Santa Barbara tested several checkers on essays written by humans and by GPT-4, and found that most tools had accuracy rates between 50% and 70% — barely better than guessing.
Some tools perform worse on certain types of writing. Text that is very simple, very formal, or very repetitive is more likely to be flagged as AI even if a human wrote it. Text that is creative, conversational, or full of personal anecdotes is more likely to pass as human even if an AI wrote it.
The confidence scores these tools show — "92% likely to be AI" — are also misleading. A high score does not mean the tool is 92% sure. It means the tool's internal model assigned a high probability to the AI class. That probability is based on patterns the tool learned, not on ground truth. A tool can be very confident and still be wrong.
What you can actually learn by reading the text
If you need to know whether something was AI-written, reading it carefully is more useful than running it through a checker. Look for specific things: Does the text have factual errors? Does it cite sources, and are those sources real? Does it make claims that are too broad or too confident? Does it show knowledge of recent events, or does it seem to cut off at a training date?
AI models often make up facts, especially about recent events, specific people, or niche topics. They also tend to hedge less than humans do — they're more likely to state something as fact when a human would say "probably" or "in most cases." And they rarely cite sources, or they cite sources that sound real but don't exist.
Another sign is structure. AI-written text often follows a very predictable outline: introduction, three main points, conclusion. Human writing, especially in casual contexts, is messier. It circles back, goes on tangents, contradicts itself, then clarifies.
When a checker result actually matters
If you're a teacher trying to catch cheating, a checker result alone is not enough evidence. You need to talk to the student, ask them to explain their ideas, and check whether the text matches what you know about their writing level. A student who usually writes at a 6th-grade level but suddenly turns in a college-level essay is suspicious — but the suspicion should lead to a conversation, not an accusation based on a tool.
If you're a publisher or content platform trying to label AI-generated content, you need a policy that accounts for the tool's limitations. You might require human review of any text that scores above a certain threshold, or you might require the author to disclose whether they used AI, rather than relying on detection alone.
If you're using a checker to screen job applications or academic submissions, be aware that you may be rejecting legitimate human work and accepting AI-written work. The tool is a signal, not a verdict.
Popular AI checkers and what they claim
GPTZero was one of the first widely used checkers. It looks at "burstiness" (variation in sentence length) and "perplexity" (how predictable the text is). It's free to use for small amounts of text, but it has been criticized for high false positive rates on human-written text.
Originality.AI combines statistical analysis with machine learning and claims to detect AI from multiple models. It charges per scan and is marketed toward content creators and publishers. Independent testing has found it more reliable than some competitors, but still not accurate enough for high-stakes use.
Turnitin added AI detection to its plagiarism checker in 2023. Because Turnitin already has access to millions of student papers, it can train its model on real human writing. But Turnitin has also been criticized for false positives, particularly on non-native English speakers and on text written by students with certain writing styles.
Safeassign (owned by Blackboard) and Copyscape also offer AI detection, though less information is publicly available about how they work or how accurate they are.
What to do if you need to know for sure
If the stakes are high — you're making a hiring decision, enforcing an academic integrity policy, or deciding whether to publish something — do not rely on a checker alone. Here's a better approach:
- Run the text through a checker if you want a starting signal, but treat the result as one data point, not a conclusion.
- Read the text yourself and look for the signs listed above: factual errors, made-up sources, missing hedging, suspiciously perfect structure.
- If you suspect AI writing, ask the author directly. In many cases, they'll tell you. If they deny it and you have reason to doubt them, ask follow-up questions about the content — a human author can usually explain their reasoning, sources, and process.
- If you're setting a policy for an organization, decide what you actually care about. Is it that the work is original? That it's thoughtful? That it discloses AI use? Your policy should match what you're trying to prevent, not just "no AI."
Frequently Asked Questions
Can AI checkers tell the difference between ChatGPT and Claude?
Not reliably. Most checkers are trained on one or two AI models, so they work better on those models and worse on others. A tool trained mostly on ChatGPT output may miss Claude-written text entirely. As new AI models are released, existing checkers become less accurate on the new models.
What if a checker says a text is 95% AI but I'm pretty sure a human wrote it?
The checker is probably wrong. A high confidence score does not mean the tool is correct — it means the tool's internal model assigned a high probability to the AI class. Read the text yourself, check the facts, and ask the author if you need to. A tool result is not proof.
Do these checkers work on text in languages other than English?
Most are trained primarily on English text, so they're less accurate on other languages. Some have versions for other languages, but they're usually less reliable than the English versions. If you need to detect AI writing in another language, you'll need a tool specifically trained on that language.
If I use an AI checker on my own writing, will it flag me as AI?
It might, depending on your writing style. If you write in a simple, direct style with short sentences and common words, a checker is more likely to flag you. If you write in a complex, varied style with longer sentences, you're less likely to be flagged. This is one reason these tools are unreliable — they penalize certain human writing styles.
Are there any AI checkers that are actually accurate?
No tool is accurate enough to use alone for high-stakes decisions. Some are better than others — Originality.AI and Turnitin tend to perform better in independent testing than some free tools — but all of them have false positive and false negative rates that are too high for reliable detection. Human review is still necessary.