The Fiction Slop Index
Which AI writes fiction that doesn't read like AI? We rank the major AI models on how human their prose reads, measured against the 758-entry AI Cliché Corpus and the fiction-craft tells other leaderboards ignore. Higher means it reads like a person wrote it.
Preview snapshot. These are the four mechanical axes only. The Arena (blind human A/B voting, 35% of the full Human Prose Score) and the complete 18-scenario fiction set are next. Scores here come from a fixed batch of generated passages, not cherry-picked.
As of July 2026, Gemini 3.1 Pro lead the Fiction Slop Index on mechanical human-likeness, scoring 85 out of 100. GLM-5.2 ranks lowest of the 9 models tested.
Mechanical Prose Score
Snapshot · 9 models · 36 passages each
| # | Model | Prose Score0–100 blended | Tells20% | Templat.15% | Rhythm15% | Show/Tell15% |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro Google | 85 | 72 | 91 | 94 | 86 |
| 2 | GPT-5.5 OpenAI | 81 | 64 | 93 | 99 | 75 |
| 3 | GPT-5.6 Sol OpenAI | 81 | 57 | 98 | 100 | 75 |
| 4 | GPT-5.6 Luna OpenAI | 80 | 58 | 96 | 99 | 74 |
| 5 | Claude Sonnet 5 Anthropic | 78 | 59 | 89 | 98 | 71 |
| 6 | Claude Fable 5 Anthropic | 77 | 59 | 84 | 99 | 73 |
| 7 | GPT-5.6 Terra OpenAI | 77 | 48 | 93 | 99 | 77 |
| 8 | DeepSeek V4 Pro DeepSeek | 76 | 58 | 81 | 98 | 75 |
| 9 | GLM-5.2 Z.ai | 74 | 52 | 78 | 100 | 73 |
The four axes
Tells
20%Corpus density across 758 AI writing tells, incl. fiction-specific clichés.
Templating
15%Repeated openers, framing frames, symbolic endings, detachable epigrams.
Rhythm
15%Sentence-length burstiness vs. the AI-default metronomic cadence.
Show / Tell
15%Filter verbs, emotion-naming, and telling labels: naming a feeling instead of showing it.
The full score adds a fifth component the leaderboard above does not yet include: Arena (35%), blind A/B crowd voting on real passages. Until it ships, the four mechanical axes are blended at their relative weights.
The Tells axis runs on the open 758-entry AI Cliché Corpus, built from published frequency studies and direct model testing.
No LLM judge in the score
The Prose Score uses only human fiction baselines and crowd votes. No model grades itself, which is what keeps the ranking honest.
Slop-free isn't the same as good
A separate Craft Quality board keeps the blind judge scores for voice, originality, and resonance. A model can read perfectly human and still be dull, which is why we keep the two scores apart.
Frequently asked questions
What is the Fiction Slop Index?+
The Fiction Slop Index ranks large language models on how human their creative writing reads. Each model's fiction is scored on four mechanical axes (AI tells, templating, sentence rhythm, and show-don't-tell), plus a planned Arena of blind human votes. A higher score means the prose reads more like a person wrote it.
Which AI model writes the most human-like fiction?+
As of July 2026, Gemini 3.1 Pro lead the Fiction Slop Index on mechanical human-likeness, scoring 85 out of 100. GLM-5.2 ranks lowest of the 9 models tested. These are preview results from the mechanical axes only; the crowd-voted Arena is not yet included.
What is AI slop in fiction writing?+
In fiction, AI slop is the cluster of recognizable machine-writing habits: overused clichés and vocabulary, repeated sentence templates and symbolic endings, a uniform metronomic rhythm, and telling instead of showing (naming an emotion or filtering the scene through 'she saw', 'he felt'). The Fiction Slop Index measures the density of these patterns.
How is the Fiction Slop Index scored?+
Each passage is scored by deterministic detectors, not by another AI acting as judge. Tells come from density against the 758-entry AI Cliché Corpus; templating from repeated openers, framing phrases, symbolic endings and epigrams; rhythm from sentence-length burstiness; and show-don't-tell from filter verbs, emotion-naming, and telling labels. Axis scores are inverted to a 0-100 human-likeness scale and blended.
Does a low slop score mean the fiction is good?+
No. Slop-free is not the same as good. A model can avoid every AI tell and still write dull, inert prose. The Fiction Slop Index measures how human the writing reads, not how good the story is; a separate Craft Quality board tracks voice, originality, and emotional resonance.
How is this different from other AI writing leaderboards?+
General slop leaderboards grade functional writing: email, essays, social posts, chat. The Fiction Slop Index is the only one built for creative writing, so it catches fiction-specific craft failures like body-language clichés, said-bookisms, and telling instead of showing that functional-prose tests never look at.
Score your own writing against the same corpus
The Slop Checker runs the exact Tells axis behind this ranking on any passage you paste, free and entirely in your browser.