vol. iii · no. 32 · tuesday, the 4th of augustink on paper · also, software · a lamplit evening
the real leaderboards

which ai actually writes well?

ranked against real benchmarks, not our opinion

most "best AI for writing" pages guess, or make up numbers. we pulled the actual leaderboards instead: where our models rank out of the field, on the benchmarks that map to how you actually use Quarkle. where no real benchmark exists, we say so.

339
models scanned across 4 leaderboards to find ours
8
models Quarkle offers, all charted
4,535
real usage picks analyzed, not survey data
0
numbers we made up
ranked by real benchmark, not vibes

six questions with real answers.

in Quarkle today
legacy, still available
tops the board, not offered in Quarkle
not yet benchmarked

whose sentences actually sound good?

source: EQ-Bench Creative Writing v3 · fetched Aug 3, 2026 · of 116 models tested

Claude Opus 5
85.3
#1/116
GPT-5.4
84.5
#3/116
Kimi K2.6NSFW
83.3
#9/116
GPT-5.6 Terra
82.8
#15/116
Claude Sonnet 4.6
82.5
#19/116
Claude Sonnet 5
82.3
#20/116
Gemini 3.1 Pro
80.2
#44/116
Kimi K2.5NSFW
79.7
#51/116
Grok 4.3NSFW
not on this leaderboard yet

LLM-judged prose rubric (voice, plot craft, coherence). Judged by Claude Sonnet 4.6, worth knowing since Claude models are partly being scored by a Claude judge. Rank is by Rubric Score specifically, not the site's own default Elo sort, so it lines up with the bars above.

view as table
ModelRubric ScoreEloRank
Claude Opus 585.352429.5#1
GPT-5.484.451885.9#3
Kimi K2.683.351709.6#9
GPT-5.6 Terra82.801886.5#15
Claude Sonnet 4.682.501872.6#19
Claude Sonnet 582.351781.8#20
Gemini 3.1 Pro80.201459.5#44
Kimi K2.579.701576.9#51
Grok 4.3---

whose plans and plots hold up?

source: EQ-Bench Longform Creative Writing · fetched Aug 3, 2026 · of 125 models tested

Claude Opus 5
86.3
#1/125
Claude Sonnet 4.6
79.9
#6/125
Kimi K2.6NSFW
78.5
#8/125
GPT-5.4
78.3
#9/125
Claude Sonnet 5
78.3
#10/125
GPT-5.6 Terra
78.0
#12/125
Kimi K2.5NSFW
74.9
#17/125
Gemini 3.1 Pro
68.2
#37/125
Grok 4.3NSFW
not on this leaderboard yet

This benchmark's own method: brainstorm a plot from a minimal prompt, revise the plan, then write it out over 8 rounds of ~1,000 words. Closest direct test of brainstorming we found. Same judge caveat applies.

view as table
ModelScoreRank
Claude Opus 586.30#1
Claude Sonnet 4.679.90#6
Kimi K2.678.50#8
GPT-5.478.30#9
Claude Sonnet 578.30#10
GPT-5.6 Terra78.00#12
Kimi K2.574.90#17
Gemini 3.1 Pro68.20#37
Grok 4.3--

who overuses AI clichés the least?

source: EQ-Bench Creative Writing v3: Slop Score · fetched Aug 3, 2026 · of 116 models tested

Claude Opus 5
6.6
#1/116
Claude Sonnet 4.6
9.9
#3/116
Claude Sonnet 5
11.6
#9/116
GPT-5.4
12.2
#14/116
GPT-5.6 Terra
12.4
#15/116
Kimi K2.6NSFW
13.3
#19/116
Kimi K2.5NSFW
17.8
#33/116
Gemini 3.1 Pro
30.0
#66/116
Grok 4.3NSFW
not on this leaderboard yet

Tallies words and constructions LLMs overuse ("not just X, but Y") against real human writing. Lower is better here: a shorter bar means fewer AI-isms, not a worse score.

view as table
ModelSlop ScoreRank
Claude Opus 56.59#1
Claude Sonnet 4.69.90#3
Claude Sonnet 511.60#9
GPT-5.412.20#14
GPT-5.6 Terra12.40#15
Kimi K2.613.30#19
Kimi K2.517.77#33
Gemini 3.1 Pro30.03#66
Grok 4.3--

can it tell good writing from bad?

source: EQ-Bench Judgemark v4 · fetched Aug 3, 2026 · of 36 models tested

Claude Opus 4.6
90.7
#1/36
Claude Sonnet 4.6
82.2
#4/36
Gemini 3.1 Pro
78.7
#6/36
GPT-5.4
72.1
#11/36
Claude Sonnet 5
70.9
#12/36
Kimi K2.6NSFW
57.3
#18/36
Grok 4.3NSFW
49.6
#23/36
GPT-5.6 Terra
not on this leaderboard yet
Kimi K2.5NSFW
not on this leaderboard yet

Judgemark doesn't test editorial notes. It tests whether a model's own scoring of writing samples actually separates strong writing from weak. That's necessary for critique, not the same thing as giving useful notes, but it's the closest real benchmark we found, and the only one here with a Grok 4.3 result.

view as table
ModelJudgemark ScoreRank
Claude Opus 4.690.73#1
Claude Sonnet 4.682.15#4
Gemini 3.1 Pro78.69#6
GPT-5.472.08#11
Claude Sonnet 570.89#12
Kimi K2.657.26#18
Grok 4.349.57#23
GPT-5.6 Terra--
Kimi K2.5--

whose quality holds up over a long chapter?

source: EQ-Bench Longform Creative Writing: Degradation · fetched Aug 3, 2026 · of 125 models tested

Claude Sonnet 4.6
0.000
#1/125
Kimi K2.6NSFW
0.011
#10/125
Claude Sonnet 5
0.014
#11/125
GPT-5.4
0.019
#14/125
GPT-5.6 Terra
0.036
#15/125
Kimi K2.5NSFW
0.070
#22/125
Gemini 3.1 Pro
0.123
#36/125
Grok 4.3NSFW
not on this leaderboard yet

Measures how much quality drops from a model's first chapter to its last, across an 8-chapter run. 0 means no measurable drop-off. Lower is better here: a shorter bar means it holds up better over length, not a worse score. No "not offered" reference bar on this one: Claude Sonnet 4.6 is already #1 in the full field.

view as table
ModelDegradation ScoreRank
Claude Sonnet 4.60.000#1
Kimi K2.60.011#10
Claude Sonnet 50.014#11
GPT-5.40.019#14
GPT-5.6 Terra0.036#15
Kimi K2.50.070#22
Gemini 3.1 Pro0.123#36
Grok 4.3--

does comprehension hold up at 192,000 tokens?

source: Fiction.liveBench: Long Context Deep Comprehension · fetched Aug 3, 2026

0 tokens
100.0
400 tokens
100.0
1k tokens
100.0
2k tokens
100.0
4k tokens
94.4
8k tokens
88.9
16k tokens
86.1
32k tokens
88.9
60k tokens
88.9
120k tokens
78.1
192k tokens
87.5

Kimi K2.5 is the only one of our 8 models this eval has tested. Same story, comprehension questions asked at increasing context lengths. Real curve, not a cross-model ranking: this is the shape of the problem, for the one model we could verify.

view as table
Context lengthComprehension ScoreRank
0 tokens100.00-
400 tokens100.00-
1k tokens100.00-
2k tokens100.00-
4k tokens94.40-
8k tokens88.90-
16k tokens86.10-
32k tokens88.90-
60k tokens88.90-
120k tokens78.10-
192k tokens87.50-
not a benchmark, actual usage

what quarkle writers actually pick.

62% of the time someone hand-picks a model, it's Claude Sonnet 4.6 or Kimi K2.6.

source: Quarkle production usage, chat & brainstorm · Jul 24–Aug 3, 2026 · 4,535 explicit picks across 164 conversations

Claude Sonnet 4.6 36.6%
Kimi K2.6 25%NSFW
Grok 4.3 13.6%NSFW
Claude Sonnet 5 9%
Kimi K2.5 8.6%NSFW
GPT-5.4 3.7%
Gemini 3.1 Pro 2.7%
GPT-5.6 Terra 0.8%

excludes "Auto" turns · Sonnet 5 / GPT-5.6 Terra shipped Aug 1, thin data

how we built this

Every number on this page came from a live leaderboard we pulled ourselves, not a summary, not a cached quote: EQ-Bench's Creative Writing v3, Longform, and Judgemark v4 boards plus fiction.liveBench's long-context eval, fetched Aug 3, 2026 and linked above each chart. A few caveats worth keeping: the writing-quality boards are judged by a Claude model (Sonnet 4.6), so Claude's own scores there deserve a raised eyebrow; Judgemark measures something different, it scores a model's own judging ability, not another judge's opinion of it, so that caveat doesn't apply there; Grok 4.3 is missing from most boards, appearing only on Judgemark so far; and fiction.liveBench has only tested one of our 8 models. Gaps in the data, not evidence anything writes worse. Check the links yourself; they're the same ones we used. The grayed-out bar on some charts is the actual #1 model on that specific board, not currently offered in Quarkle. We're showing it so this stays an honest leaderboard, not just a lineup of our own models.

so, which model should you use?

honestly? whichever one fits the scene.

The data above doesn't point at one clean winner, so we didn't build Quarkle around one. Pick a model per chapter, switch mid-book, or leave it on auto and let it choose. Free to try, no card required.