Research
Can humans detect AI text?
Six studies, two opposite conclusions. Untrained readers score at chance on frontier models; trained experts in familiar genres reach 60–70%. We walk through each, including the one that contradicts the headline and gets buried in detector-vendor marketing.
The short answer
Untrained readers identify modern frontier-LLM text at chance, about 50%. Trained domain experts reading familiar genres (essays they grade, abstracts they review) reach 60–70%. No peer-reviewed study has shown reliable above-80% human accuracy on current frontier output. The genre and the rater matter more than any stable "human ability."
The studies below are the ones most cited in the debate, including several from vendor-side arguments that humans can't reliably detect AI. We also include the one those roundups tend to leave out: the educator study where instructors hit 70%.
The studies
Clark et al. 2021, "All That's Human Is Not Gold" (ACL)
arXiv:2107.00061 · 255 stories, ~2,111 news articles, 50 recipes, GPT-2 and GPT-3 vs human writing. Untrained evaluators hit 57.9% on GPT-2 and 49.9% on GPT-3, literally chance. Training (with instructions, annotated examples, paired examples) lifted accuracy to ~55%, not significantly different from untrained. This is the strongest peer-reviewed anchor for the "humans at chance" claim because ACL is a top-tier NLP venue and the sample is genre-diverse.
Jones & Bergen 2024, Turing test on GPT-4 (NAACL)
arXiv:2405.08007 · 500 participants, five-minute conversational Turing tests. GPT-4 was judged human 54% of the time; real humans 67%; ELIZA 22%. GPT-4's pass rate was not statistically different from chance (z = 0.14, p = 0.48). Caveat: this is interactive messaging, not pasted-text forensics; the dynamics are different.
The educator study, instructors at 70%
International Journal for Educational Integrity (2023), 140 instructors + 145 students, ChatGPT-generated college essays. Instructors hit 70% accuracy; students 60%. This is the study that contradicts the "humans can't detect" headline. Trained raters reading a familiar genre (essays they grade weekly) beat any general "human ability" framing.
Surgical abstracts (Arthroscopy journal)
Academic reviewers, ChatGPT-generated abstracts in shoulder and elbow surgery (Arthroscopy, 2023): 62% accuracy with a 38% false-positive rate (real human abstracts flagged as AI). Very narrow domain.
Plastic surgery personal statements
Canadian Society of Plastic Surgeons (2023), GPT-4, 11 AI + 11 human personal statements rated by 2 retired surgeons with 20+ years of experience: 65.9% accuracy. Tiny N, basically a case study.
Dental research abstracts (Nature Scientific Reports, March 2026)
150 abstracts, six early-career academics, ChatGPT 3.5 (now three generations old). Accuracy ranged from 44% to 76%, extreme inter-rater variance, on a model nobody's asked to detect for years.
What the literature actually says, taken together
- Untrained × frontier model × neutral genre ≈ chance (Clark, Jones).
- Trained × frontier model × familiar genre ≈ 60–70% (educators on essays, reviewers on abstracts). Above chance, but a long way from reliable.
- None of the published studies measure ability to quantify AI-likeness on a continuous scale; they're all binary forced-choice. That gap is why we built a continuous score with per-feature explanations.
Reading vendor roundups carefully
The detector industry has a commercial interest in the framing "humans can't reliably detect AI": it's the case for buying a tool. The studies most often cited are real, but vendor write-ups lead with the worst-for-humans results (Clark 2021, Jones 2024) and bury the educator study at 70%. Often missing is Liang et al. 2023, the Stanford paper showing detectors mislabel 61% of non-native-English human essays as AI. That study points the other way in the same debate, and it's why the limits of AI detection matter before you trust any single tool.
What this means for the Slop Score
We're not in the "is this AI?" business. The reason: the question doesn't cleanly resolve, and the products that try to answer it carry an accuracy debt that compounds every time a court reads their methodology. The slop register is measurable independently of authorship — humans produce it, careful AI rewrites avoid it. We score the register and show the spans; the writer decides what to do with them.