Do AI Detectors Work on GPT-5, Claude and Gemini?
Detectors are generally model-agnostic: they look at statistical properties of text rather than fingerprints of a specific system. That means they do work on newer models — but less reliably than on older ones, and the trend is downward.

Detectors are not model-specific
A common misconception is that a detector recognises a particular model the way antivirus recognises a virus signature. It does not. There is no watermark to find in ordinary generated text, and no lookup table of known outputs.
What a detector measures is whether the text carries the statistical properties typical of machine generation — predictable word choice, even sentence rhythm, low variation. Those properties are shared across systems, which is why a detector trained before a model existed can still flag its output.
It also explains the limitation: as models produce text that is statistically closer to human writing, the signal every detector depends on gets weaker for all of them at once.
Why each generation is harder

The fourth point is the one most people underestimate. A default-prompted response tends to have a recognisable shape — measured, balanced, evenly structured. Ask the same model to write in a specific voice, or to a tight word count, or in an unusual format, and the statistical profile shifts considerably.
This means the same model can produce output that scores very differently depending on how it was asked. The prompt matters as much as the model.
What still shows up
Detection is harder, not dead. Signal that tends to persist across generations:
- Sustained evenness over long passages. Human writers drift; their attention and energy vary across a piece. Generated text tends to hold its register.
- Absence of specific, verifiable detail. Generated prose gravitates toward the general unless pushed for specifics.
- Structural symmetry. Three examples per point, balanced paragraphs, a neat closing summary.
- Hedged, non-committal positions where a human writer with expertise would take a side.
None of these is conclusive. All of them can appear in human writing, especially in formal or institutional registers.
What about mixed and edited text?
Most real-world text is not purely one or the other. A person drafts with a model and rewrites, or writes themselves and asks for a polish. This is the hardest case for detection, and the most common.
Editing tends to move a passage toward the middle of the scale. A human pass over generated text raises perplexity and burstiness; a model polish over human text lowers them. A mid-range score frequently means exactly this — mixed authorship — rather than a confident finding either way.
How to interpret results across models
- Use longer samples. Length is the single biggest factor in stability, regardless of which system produced the text.
- Do not infer the model from the score. Detectors do not identify which system wrote something, and any tool claiming to is overstating.
- Expect newer output to score lower. A modest AI probability on recent-model text is not the same as a modest probability on older output.
- Treat mid-range as mixed or unknown, not as half-generated.
Test any passage
Run text from any model through our free detector and see how it scores.
Open the AI Detector →Frequently asked questions
Do AI detectors work on GPT-5?
Generally yes, but less reliably than on older models. Detectors measure statistical properties of text rather than model-specific fingerprints, so they still apply — but newer models produce text closer to human patterns, which weakens the signal.
Can a detector tell which AI model wrote something?
No. Detectors estimate whether text resembles generated writing in general. They do not identify the specific system, and any tool claiming to identify the model is overstating what the method supports.
Why do newer AI models score lower on detectors?
They are trained to produce text people prefer, which means text that reads more naturally. Sampling settings add variation, and detectors are tuned on models that already existed when they were built.
Does editing AI text change the detection score?
Yes. A human editing pass typically raises perplexity and burstiness, moving the score toward the middle. This is why mid-range results often indicate mixed authorship rather than a confident finding.
Does a low AI score prove text was written by a human?
No. It means the text does not carry the patterns the detector looks for. That is weaker than proof of human authorship.
Related: How Accurate Are AI Detectors? · Perplexity and Burstiness · AI Detector