HomeBlog › Model coverage

Do AI Detectors Work on GPT-5, Claude and Gemini?

The AI Detector Editorial · August 22, 2026

Detectors are generally model-agnostic: they look at statistical properties of text rather than fingerprints of a specific system. That means they do work on newer models — but less reliably than on older ones, and the trend is downward.

Title graphic asking whether AI detectors work on GPT-5, Claude and Gemini, noting that newer models are harder to detect.
Detection gets harder with every model generation — but not uniformly.

Detectors are not model-specific

A common misconception is that a detector recognises a particular model the way antivirus recognises a virus signature. It does not. There is no watermark to find in ordinary generated text, and no lookup table of known outputs.

What a detector measures is whether the text carries the statistical properties typical of machine generation — predictable word choice, even sentence rhythm, low variation. Those properties are shared across systems, which is why a detector trained before a model existed can still flag its output.

It also explains the limitation: as models produce text that is statistically closer to human writing, the signal every detector depends on gets weaker for all of them at once.

Why each generation is harder

Four numbered steps explaining why newer AI models are harder to detect: training targets human-likeness, sampling adds variation, detectors lag new releases, and prompting changes output.
Four compounding reasons detection gets harder over time.

The fourth point is the one most people underestimate. A default-prompted response tends to have a recognisable shape — measured, balanced, evenly structured. Ask the same model to write in a specific voice, or to a tight word count, or in an unusual format, and the statistical profile shifts considerably.

This means the same model can produce output that scores very differently depending on how it was asked. The prompt matters as much as the model.

What still shows up

Detection is harder, not dead. Signal that tends to persist across generations:

None of these is conclusive. All of them can appear in human writing, especially in formal or institutional registers.

Practical consequence: if a detector returns a low AI probability, that does not establish the text was human-written — only that it does not carry the patterns the detector looks for. Absence of evidence is weak evidence of absence here.

What about mixed and edited text?

Most real-world text is not purely one or the other. A person drafts with a model and rewrites, or writes themselves and asks for a polish. This is the hardest case for detection, and the most common.

Editing tends to move a passage toward the middle of the scale. A human pass over generated text raises perplexity and burstiness; a model polish over human text lowers them. A mid-range score frequently means exactly this — mixed authorship — rather than a confident finding either way.

How to interpret results across models

Test any passage

Run text from any model through our free detector and see how it scores.

Open the AI Detector →

Frequently asked questions

Do AI detectors work on GPT-5?

Generally yes, but less reliably than on older models. Detectors measure statistical properties of text rather than model-specific fingerprints, so they still apply — but newer models produce text closer to human patterns, which weakens the signal.

Can a detector tell which AI model wrote something?

No. Detectors estimate whether text resembles generated writing in general. They do not identify the specific system, and any tool claiming to identify the model is overstating what the method supports.

Why do newer AI models score lower on detectors?

They are trained to produce text people prefer, which means text that reads more naturally. Sampling settings add variation, and detectors are tuned on models that already existed when they were built.

Does editing AI text change the detection score?

Yes. A human editing pass typically raises perplexity and burstiness, moving the score toward the middle. This is why mid-range results often indicate mixed authorship rather than a confident finding.

Does a low AI score prove text was written by a human?

No. It means the text does not carry the patterns the detector looks for. That is weaker than proof of human authorship.

Related: How Accurate Are AI Detectors? · Perplexity and Burstiness · AI Detector