which ai actually writes well?
ranked against real benchmarks, not our opinion
most "best AI for writing" pages guess, or make up numbers. we pulled the actual leaderboards instead: where our models rank out of the field, on the benchmarks that map to how you actually use Quarkle. where no real benchmark exists, we say so.
six questions with real answers.
whose sentences actually sound good?
source: EQ-Bench Creative Writing v3 · fetched Aug 3, 2026 · of 116 models tested
LLM-judged prose rubric (voice, plot craft, coherence). Judged by Claude Sonnet 4.6, worth knowing since Claude models are partly being scored by a Claude judge. Rank is by Rubric Score specifically, not the site's own default Elo sort, so it lines up with the bars above.
view as table
| Model | Rubric Score | Elo | Rank |
|---|---|---|---|
| Claude Opus 5 | 85.35 | 2429.5 | #1 |
| GPT-5.4 | 84.45 | 1885.9 | #3 |
| Kimi K2.6 | 83.35 | 1709.6 | #9 |
| GPT-5.6 Terra | 82.80 | 1886.5 | #15 |
| Claude Sonnet 4.6 | 82.50 | 1872.6 | #19 |
| Claude Sonnet 5 | 82.35 | 1781.8 | #20 |
| Gemini 3.1 Pro | 80.20 | 1459.5 | #44 |
| Kimi K2.5 | 79.70 | 1576.9 | #51 |
| Grok 4.3 | - | - | - |
whose plans and plots hold up?
source: EQ-Bench Longform Creative Writing · fetched Aug 3, 2026 · of 125 models tested
This benchmark's own method: brainstorm a plot from a minimal prompt, revise the plan, then write it out over 8 rounds of ~1,000 words. Closest direct test of brainstorming we found. Same judge caveat applies.
view as table
| Model | Score | Rank |
|---|---|---|
| Claude Opus 5 | 86.30 | #1 |
| Claude Sonnet 4.6 | 79.90 | #6 |
| Kimi K2.6 | 78.50 | #8 |
| GPT-5.4 | 78.30 | #9 |
| Claude Sonnet 5 | 78.30 | #10 |
| GPT-5.6 Terra | 78.00 | #12 |
| Kimi K2.5 | 74.90 | #17 |
| Gemini 3.1 Pro | 68.20 | #37 |
| Grok 4.3 | - | - |
who overuses AI clichés the least?
source: EQ-Bench Creative Writing v3: Slop Score · fetched Aug 3, 2026 · of 116 models tested
Tallies words and constructions LLMs overuse ("not just X, but Y") against real human writing. Lower is better here: a shorter bar means fewer AI-isms, not a worse score.
view as table
| Model | Slop Score | Rank |
|---|---|---|
| Claude Opus 5 | 6.59 | #1 |
| Claude Sonnet 4.6 | 9.90 | #3 |
| Claude Sonnet 5 | 11.60 | #9 |
| GPT-5.4 | 12.20 | #14 |
| GPT-5.6 Terra | 12.40 | #15 |
| Kimi K2.6 | 13.30 | #19 |
| Kimi K2.5 | 17.77 | #33 |
| Gemini 3.1 Pro | 30.03 | #66 |
| Grok 4.3 | - | - |
can it tell good writing from bad?
source: EQ-Bench Judgemark v4 · fetched Aug 3, 2026 · of 36 models tested
Judgemark doesn't test editorial notes. It tests whether a model's own scoring of writing samples actually separates strong writing from weak. That's necessary for critique, not the same thing as giving useful notes, but it's the closest real benchmark we found, and the only one here with a Grok 4.3 result.
view as table
| Model | Judgemark Score | Rank |
|---|---|---|
| Claude Opus 4.6 | 90.73 | #1 |
| Claude Sonnet 4.6 | 82.15 | #4 |
| Gemini 3.1 Pro | 78.69 | #6 |
| GPT-5.4 | 72.08 | #11 |
| Claude Sonnet 5 | 70.89 | #12 |
| Kimi K2.6 | 57.26 | #18 |
| Grok 4.3 | 49.57 | #23 |
| GPT-5.6 Terra | - | - |
| Kimi K2.5 | - | - |
whose quality holds up over a long chapter?
source: EQ-Bench Longform Creative Writing: Degradation · fetched Aug 3, 2026 · of 125 models tested
Measures how much quality drops from a model's first chapter to its last, across an 8-chapter run. 0 means no measurable drop-off. Lower is better here: a shorter bar means it holds up better over length, not a worse score. No "not offered" reference bar on this one: Claude Sonnet 4.6 is already #1 in the full field.
view as table
| Model | Degradation Score | Rank |
|---|---|---|
| Claude Sonnet 4.6 | 0.000 | #1 |
| Kimi K2.6 | 0.011 | #10 |
| Claude Sonnet 5 | 0.014 | #11 |
| GPT-5.4 | 0.019 | #14 |
| GPT-5.6 Terra | 0.036 | #15 |
| Kimi K2.5 | 0.070 | #22 |
| Gemini 3.1 Pro | 0.123 | #36 |
| Grok 4.3 | - | - |
does comprehension hold up at 192,000 tokens?
source: Fiction.liveBench: Long Context Deep Comprehension · fetched Aug 3, 2026
Kimi K2.5 is the only one of our 8 models this eval has tested. Same story, comprehension questions asked at increasing context lengths. Real curve, not a cross-model ranking: this is the shape of the problem, for the one model we could verify.
view as table
| Context length | Comprehension Score | Rank |
|---|---|---|
| 0 tokens | 100.00 | - |
| 400 tokens | 100.00 | - |
| 1k tokens | 100.00 | - |
| 2k tokens | 100.00 | - |
| 4k tokens | 94.40 | - |
| 8k tokens | 88.90 | - |
| 16k tokens | 86.10 | - |
| 32k tokens | 88.90 | - |
| 60k tokens | 88.90 | - |
| 120k tokens | 78.10 | - |
| 192k tokens | 87.50 | - |
what quarkle writers actually pick.
62% of the time someone hand-picks a model, it's Claude Sonnet 4.6 or Kimi K2.6.
source: Quarkle production usage, chat & brainstorm · Jul 24–Aug 3, 2026 · 4,535 explicit picks across 164 conversations
excludes "Auto" turns · Sonnet 5 / GPT-5.6 Terra shipped Aug 1, thin data
Every number on this page came from a live leaderboard we pulled ourselves, not a summary, not a cached quote: EQ-Bench's Creative Writing v3, Longform, and Judgemark v4 boards plus fiction.liveBench's long-context eval, fetched Aug 3, 2026 and linked above each chart. A few caveats worth keeping: the writing-quality boards are judged by a Claude model (Sonnet 4.6), so Claude's own scores there deserve a raised eyebrow; Judgemark measures something different, it scores a model's own judging ability, not another judge's opinion of it, so that caveat doesn't apply there; Grok 4.3 is missing from most boards, appearing only on Judgemark so far; and fiction.liveBench has only tested one of our 8 models. Gaps in the data, not evidence anything writes worse. Check the links yourself; they're the same ones we used. The grayed-out bar on some charts is the actual #1 model on that specific board, not currently offered in Quarkle. We're showing it so this stays an honest leaderboard, not just a lineup of our own models.
honestly? whichever one fits the scene.
The data above doesn't point at one clean winner, so we didn't build Quarkle around one. Pick a model per chapter, switch mid-book, or leave it on auto and let it choose. Free to try, no card required.