Calibration set
The texts we pin our score against
A score is only worth trusting if it lands where it should. The reference texts below are pinned in our test suite. Every change to engine weights or dictionaries has to keep them inside their band, or it doesn't ship. We publish them in full so you can check the bands against your own ear.
None of our competitors publish their calibration set. We do it for two reasons. First, it's the only honest answer to "how do I know your score is meaningful?": you can read the anchor text, look at the band, and decide whether our notion of slop matches yours. Second, it's the gate that keeps the engines from drifting toward whichever stylistic feature was last bolted on. The known limitations page covers where the score breaks down; this page covers where it's pinned.
The bands, at a glance
Default-prompted AI marketing copy
id: ai_marketing
80–100
LinkedIn thought-leader post (AI)
id: linkedin_slop
70–95
Corporate jargon memo (human, slop-adjacent)
id: corporate_memo
48–72
Plain-spoken argumentative essay (human)
id: pg_essay
0–28
Literary personal essay (human)
id: literary_personal
0–22
Children's book paragraph (human)
id: childrens_book
0–20
The texts, in full
Each anchor is a representative passage written in its style, not a copyrighted excerpt. The bands are deliberately wide enough to absorb minor scoring drift but narrow enough that any engine change that shifts the shape of how slop is measured will break a band and surface the change in CI.
Default-prompted AI marketing copy
Expected band: 80–100
What you get from a major LLM when you give it a generic marketing brief and click send. Tricolons stacked, em-dashes everywhere, every cliché in the corpus.
In today's fast-paced digital world, businesses must navigate the ever-evolving landscape of customer engagement. Our comprehensive platform leverages cutting-edge technology to deliver seamless integration across every touchpoint. It's not just a tool — it's a testament to what's possible when innovation meets execution. We delve deeper into the data so you don't have to, unlocking actionable insights that foster growth. Whether you're a startup or an enterprise, our robust, multifaceted solution elevates your workflow and empowers your team to thrive. At the intersection of simplicity and power, we help you embark on a transformative journey. Ultimately, success in the modern era hinges on agility, scalability, and vision. Our holistic approach ensures that you remain ahead of the curve. In conclusion, there has never been a better time to harness the power of intelligent automation and unlock your organization's full potential.
LinkedIn thought-leader post (AI)
Expected band: 70–95
The thought-leader register: contrast-pair anaphora ("Most people X. But the best Y."), staccato one-line paragraphs, big closing claim. Reads confident, says nothing.
Most people treat their calendar like a to-do list. But the highest performers treat it like a strategy document. Here's what that actually means in practice: Your calendar is not where work goes. It's where priorities go. Most people optimize for busy. The best optimize for leverage. That's the difference between motion and progress. Time blocking is less about discipline and more about clarity. A good block usually combines a theme, a deadline, and a single outcome. The biggest takeaway from all of this: Productivity is evolving from doing more into deciding better. The people who win won't just manage their time. They'll design it.
Corporate jargon memo (human, slop-adjacent)
Expected band: 48–72
Human-written, but slop-adjacent. Heavy jargon, hedged action items, no specific names or numbers. Should land in the middle, voiceless without being AI.
Team — following Q3 planning, we are realigning priorities to maximize stakeholder value. Each workstream lead should circle back with their OKRs by Friday so we can socialize the roadmap before the board review. Per the steering committee, any headcount asks need a business case attached. Let's leverage the momentum from the offsite and keep cross-functional dependencies visible on the tracker. If you are blocked, flag it early rather than at the eleventh hour. Reply-all with open items and we will triage async ahead of Monday's standup. Appreciate everyone's hustle on this.
Plain-spoken argumentative essay (human)
Expected band: 0–28
Plain-spoken, opinionated, specific. Takes a position ("default alive beats default fundable"), names a concrete observation, refuses the both-sides closer.
The mistake most founders make is believing that fundraising is the hard part. It isn't. The hard part is building something people actually want, and you can usually tell within a week of launching whether you have. I watched a friend raise four million dollars for an idea no one needed. The money didn't fix anything; it just let him be wrong for longer, with more employees. Default alive beats default fundable every single time. So spend the first cheque finding out whether anyone comes back on a Tuesday. If they do, raising later is easy. If they don't, no deck will save you.
Literary personal essay (human)
Expected band: 0–22
A specific person remembering a specific thing in a specific way. Concrete nouns ("brass hook," "chipped blue dish"), a stance about what the detail meant.
My father kept his ties on a brass hook by the door, and every morning he chose one the way other men choose their words. The blue one meant a good day was still possible. The gray one meant he had already given up but would go to work anyway, because that is what you did. I learned to read his whole week off that hook before I could read a clock. On Sundays the hook was empty, and the house held its breath, and we ate eggs in a silence that felt, to a child, almost like peace.
Children's book paragraph (human)
Expected band: 0–20
Short, voicey, image-rich. Should score very low. This is the shape of writing that sounds like one person, not the average of everyone.
The little pig was named Wilbur. He lived in a warm corner of the barn, close to the cows and the patient sheep. In the morning he drank his milk from a chipped blue dish, and in the long afternoons he watched the sky change colors over the orchard. The barn smelled of hay and of apples and of the good sweet breath of sleeping animals. Wilbur thought it was the finest place in all the world, and on most days he was right.
How the gate works
The deterministic calibration test scores every anchor using a mocked epistemic sub-score (the value listed on each anchor) and checks the composite lands inside the band. A live variant calls the real Claude judge for the epistemic layer and verifies the same thing end-to-end. Engine changes that would push an anchor out of band fail CI and don't merge.
We don't treat the bands as sacred, when the meaning of slop shifts (a new model family produces a new register, a new pattern enters the dictionary), the right response is to add anchors and adjust bands, not to fudge a score back inside. We publish the next planned additions on the what-slop-sounds-like page.