deslopify.me

All of deslopify.me

Research

The studies behind the Slop Score

Below are plain-language summaries of the academic and industry literature that shaped how this tool measures stylistic slop. Every claim links to its source; where studies disagree, both are linked.

Every page here links primary sources and flags where a source is self-interested (a detector vendor publishing "humans can't detect AI" is selling you something). When two sources disagree, both are linked.

What slop sounds like

Our editorial definition. Slop is a stylistic register — voiceless, hedge-heavy, interchangeable prose — not a verdict on AI authorship. What it sounds like at 20, at 60, at 90.

Can humans detect AI text?

Six studies, two opposite conclusions. Untrained readers hit chance (~50%) on modern frontier models; trained domain experts hit 60–70% on familiar genres. We walk through each study, including one whose finding cuts the opposite way from how it's commonly cited.

The limits of AI detection

Non-native English bias (Liang 2023 at Stanford: 61% false-positive rate on TOEFL essays). Paraphrase evasion (Krishna 2023 at NeurIPS). Memorized canonical text. A theoretical bound by Sadasivan et al. Why we built a diagnostic, not a verdict.

Blind readers prefer AI marketing copy

Zhang & Gosline 2023 (MIT Sloan / Berkeley Haas, 1,203 participants): in blind tests, AI-written ads beat human-expert ads on satisfaction and willingness-to-pay. Readers don't avoid AI prose; they prefer humans only when told which is which. Why slop ≠ low quality.

How much of the web is AI now?

Industry studies converge near 50% of new articles primarily AI-written and 74% of new pages touched by AI in some form, with only a few percent fully AI. Numbers disagree because every measurement comes from a detector; direction doesn't.

Open benchmarks

How the Slop Score performs against the CI-pinned calibration anchors, the Zhang & Gosline 4-paradigm experimental anchors, and the paraphrase-resilience deltas (−17 surface vs −33 substantive). Plus what we still can't measure.

Why we publish this

The AI-detector category headlines accuracy claims in the high 90s that collapse on independent test sets (Liang 2023; Sadasivan et al.). The honest answer to "how good is this score?" is that calibration holds on stylistic register and breaks on adversarial paraphrase — and we show you exactly where. That's what these pages are for.

Where you find an error or a study we should add, the waitlist email reaches us. We re-cite when the literature changes.