AI Education

Best AI Model for Writing: Why Four Leaderboards Disagree

Illustration: a glowing brain made of text floating above an open book and a vintage typewriter

“Best AI model for writing” has even less of a clean answer than the coding question. For coding, vendors at least publish benchmark scores, even if they disagree on which benchmark. For writing, six of the seven vendors’ own announcement pages make no writing-specific claim at all: the words “write” and “creative” on those pages refer to code, presentations or agents. The only evidence is independent: one board built on human votes and three built on language-model judges. This guide reads all four side by side, and the headline result is that they do not agree.

Scope, stated plainly: this compares the seven models this site tracks. Newer models exist and sit above them on the same boards, including GPT-6 Astra, GPT-6.1 Sol, Claude Opus 5 and 5.5, Gemini 4 and Grok 4.7. On Arena’s Creative Writing board the top entry is a Gemini 4 variant at 1519, and none of the seven leads any board. All leaderboard numbers below were read directly from the live pages on 3 October 2026.

What each leaderboard actually measures

  • Arena Creative Writing (arena.ai, formerly LMArena) uses human votes. People see two anonymous responses to the same prompt and pick one; the board converts about 1.37 million votes across 411 models into an Elo-style score. We use the default view, which has style control on: it adjusts for response length and formatting so that longer or more heavily formatted answers do not win on style alone. The page is dated 2 October 2026.
  • EQ-Bench Creative Writing v3 scores short pieces with a language model as the judge. Rubric scores come from Claude Sonnet 4; the Elo ranking uses Sonnet 4.6 as judge. The page prints no update date, and its table has no rank column, so the positions below are our count of the sorted table (141 entries).
  • EQ-Bench Longform Creative Writing tests multi-chapter output and adds two diagnostics: a slop figure (reliance on stock phrasing) and a degradation figure (how much quality falls over a long text). Across the models below, lower values in both columns go with higher scores. 140 entries, no printed date.
  • Lech Mazur’s Creative Story-Writing Benchmark compares stories pairwise, judged by other language models: 56 writer models, 102,592 judgments, latest changelog entry 26 September 2026.

WritingBench is left out on purpose: its page says “Last Updated: 2025-11-27”, and none of the six current models is listed on it.

Human votes: Arena Creative Writing

Model Arena variant Score Rank (spread) Votes
Gemini 3.1 Pro gemini-3.1-pro-preview 1480 ±5 11 (5–23) 22,309
GPT-5.6 Sol gpt-5.6-sol-xhigh 1468 ±8 19 (8–37) 7,537
Qwen3.8 Max qwen3.8-max 1468 ±10 20 (7–40) 4,555
DeepSeek V4 Pro deepseek-v4-pro 1446 ±7 48 (25–73) 9,721
Grok 4.6 grok-4.6-high 1445 ±9 52 (24–77) 5,198
Claude Sonnet 5 claude-sonnet-5-high 1435 ±7 65 (37–87) 9,413
Llama 4 Maverick llama-4-maverick-17b-128e-instruct 1307 ±9 224 (200–254) 5,110

Two cautions. First, the variant names are the ones Arena evaluated: the “high” and “xhigh” suffixes are reasoning-effort settings, and we did not verify which one an API or chat app uses by default. Second, the rank spreads overlap heavily. Gemini 3.1 Pro, GPT-5.6 Sol and Qwen3.8 Max sit within each other’s ranges, and so do DeepSeek V4 Pro, Grok 4.6 and Claude Sonnet 5, so “X beats Y” is not a safe reading of any close pair.

Language-model judges: EQ-Bench and Mazur

Model EQ-Bench CW v3 (Elo, position) EQ-Bench Longform (score, position) Mazur stories (score, rank of 56)
GPT-5.6 Sol 1971.6, 10th 81.7, 12th 2.488, 8th (xhigh)
Claude Sonnet 5 1794.0, 27th 78.3, 20th not listed
DeepSeek V4 Pro 1553.2, 57th 75.6, 26th 0.547, 21st (high)
Gemini 3.1 Pro 1491.3, 65th 68.2, 50th −2.180, 47th
Qwen3.8 Max not listed not listed 0.204, 27th
Grok 4.6 not listed not listed −2.837, 51st (high)
Llama 4 Maverick 860.1, 124th 31.9, 130th not listed

“Not listed” means exactly that. EQ-Bench has Grok 4.7 but not Grok 4.6, and we do not substitute one for the other. For Qwen, EQ-Bench lists an open-weights checkpoint, Qwen3.8-2.4T-A95B, which Qwen’s own model card describes as the base of Qwen3.8 Max “with more features”; it is a related model, not the same one, so it is not counted. The Mazur board has a second GPT-5.6 Sol entry (high) at 2.299, 11th. These scales are not interchangeable: Elo, rubric points and centered comparison scores measure different things.

Infographic: a rank matrix showing seven AI models across four writing leaderboards. GPT-5.6 Sol ranks 19th of 411 on Arena, 10th of 141 on EQ-Bench Creative Writing v3, 12th of 140 on EQ-Bench Longform and 8th of 56 on the Mazur story benchmark. Gemini 3.1 Pro ranks 11th on Arena but 65th, 50th and 47th on the three language-model-judged boards.

What the disagreement shows

  • GPT-5.6 Sol is the only one of the seven that places near the top of every board that lists it: 19th of 411 on Arena, 10th of 141 on EQ-Bench v3, 12th of 140 on Longform and 8th of 56 on Mazur. It also has the cleanest long-form diagnostics in the group, with slop at 12.0 and zero measured degradation.
  • Gemini 3.1 Pro splits the boards. It has the highest Arena score of the seven (1480, rank 11) but sits 65th on EQ-Bench v3, 50th on Longform and 47th on Mazur. Its Longform slop figure is 38.1, the highest of the four current models on that board. Human voters and language-model judges are rewarding different things here.
  • Claude Sonnet 5 splits the other way. It is 65th on Arena Creative Writing but 27th on EQ-Bench v3 and 20th on Longform, with slop of 13.5 and near-zero degradation (0.014). It is not on the Mazur board.
  • DeepSeek V4 Pro is mid-table everywhere: 48th on Arena, 57th on EQ-Bench v3, 26th on Longform, 21st on Mazur.
  • Llama 4 Maverick is far behind on every board that lists it: 224th on Arena, 124th on EQ-Bench v3 and 130th on Longform. Meta’s April 2025 announcement called it “great for precise image understanding and creative writing” and cited an ELO of 1417 on LMArena, but that figure was for an experimental chat variant, not the released model. Arena now scores the released model at 1327 overall.

Each judged board is only as neutral as its judge model, which is why the human-vote board and the judged boards are best read side by side rather than averaged into one number.

The practical limits: output length, context and price

Writing tasks also hit hard limits that no quality leaderboard captures. These figures come from this site’s own model data:

Model Max output (tokens) Context window Price per 1M tokens (in / out)
Grok 4.6 450,000 500,000 $2 / $6
DeepSeek V4 Pro 384,000 1,048,576 $0.66 / $1.98
Qwen3.8 Max 131,072 1,000,000 $2 / $6
GPT-5.6 Sol 128,000 1,050,000 $4 / $20
Claude Sonnet 5 128,000 1,000,000 $2 / $10
Gemini 3.1 Pro 65,536 1,048,576 $2 / $12
Llama 4 Maverick 16,384 128,000 $0.19 / $0.65
Infographic: maximum output length in tokens for seven AI models, from Grok 4.6 at 450,000 and DeepSeek V4 Pro at 384,000 down to Gemini 3.1 Pro at 65,536 and Llama 4 Maverick at 16,384

For a single long document, a manuscript chapter set or a long report written in one request, the output cap matters as much as style. Grok 4.6 and DeepSeek V4 Pro can generate several times more text in one response than the rest; Llama 4 Maverick stops at 16,384 tokens, which is roughly a long article.

Picking one, by scenario

  • Consistent quality across every judge: GPT-5.6 Sol, at the highest price in the group ($4 / $20).
  • What human voters prefer in creative writing: Gemini 3.1 Pro leads the seven on Arena, within the margin of GPT-5.6 Sol and Qwen3.8 Max, and its 65,536-token output cap is the shortest of the six current models.
  • Long-form with low stock phrasing, at mid price: Claude Sonnet 5 ($2 / $10), whose strength shows on the judged boards rather than the human-vote board.
  • Best value for drafts: DeepSeek V4 Pro, mid-table on all four boards at roughly one-tenth of GPT-5.6 Sol’s output price, with a 384,000-token output cap.
  • Very long single outputs: Grok 4.6 offers the largest cap, but it ranks 51st of 56 on Mazur and is not listed on EQ-Bench, so check its prose on your own material before committing.
  • Not for long-form: Llama 4 Maverick, on output cap and every board that lists it.

Writing quality is the most taste-dependent thing on this site’s list of comparisons. Run your own prompt, in your own voice, on two or three of these before choosing, and use the leaderboards to decide which two or three.

Frequently asked questions

Which AI model is best for creative writing?

Of the seven models this site tracks, none wins on every measure. GPT-5.6 Sol is the most consistent across all four independent leaderboards, Gemini 3.1 Pro has the highest human-vote score on Arena, and Claude Sonnet 5 does better with language-model judges than with human voters. Newer models not covered here rank above all of them.

Why do the leaderboards disagree?

They measure different things. Arena counts human preferences between anonymous responses, while EQ-Bench and the Mazur benchmark rely on language models acting as judges, with different prompts, rubrics and score scales. A model can please one kind of judge more than the other.

Do AI vendors publish writing benchmarks?

Not for these models. Six of the seven vendors’ announcement pages make no writing-specific claim. Meta’s Llama 4 announcement did mention creative writing, but its quoted ELO figure was for an experimental chat variant, not the released model.

What is style control on Arena?

A statistical adjustment that accounts for response length and formatting style, so a model does not climb the board simply by writing longer or using more markdown. The scores in this guide are from the default view, with style control on.

Is a bigger output limit better for writing?

Only for tasks that need one very long response. Quality per token is a separate question that the output cap does not answer, which is why this guide shows both.