"Best AI for writing" is a moving target. Worse, it's the wrong question for a lot of real writing work, because different stages of producing a piece of writing benefit from different strengths. Research, drafting, and editing aren't the same task, and treating them as one undifferentiated "writing" problem is why a single model recommendation rarely holds up well across an entire writing workflow.

Why writing quality is harder to benchmark than coding

Code correctness is at least partly objective: it runs or it doesn't, it passes tests or it doesn't. Writing quality is much more subjective: tone, voice, structure, and persuasiveness are judged differently by different readers and for different purposes. That's why "best AI for writing" rankings tend to be less reliable and more opinion-driven than coding benchmarks, and deserve more skepticism, not less.

What actually matters, broken down by the stage of writing

Research and fact-gathering

If a writing task depends on current information or needs to be grounded in real sources, what matters most is whether the tool can search and cite real, verifiable sources rather than just generating plausible-sounding claims. This is the stage where grounding in real retrieved information matters most. It's also where AI-generated citations most often turn out to be fabricated if they aren't checked against real sources.

Drafting

What matters here is how well a model can follow specific instructions about structure, tone, and length, and how naturally it can work from an outline or a set of points rather than needing to invent the structure itself. A model that drafts well from clear direction is more useful in practice than one that produces polished-sounding prose from vague instructions but ignores your actual intent.

Editing and refining

Editing existing text well requires preserving the original voice while improving clarity, structure, or argument. That's a different skill from generating original prose. A model that's strong at generating new content isn't automatically the strongest choice for refining something that's already mostly there.

Fact-checking

Because any model can produce a confident, wrong statement, a genuinely careful writing workflow includes a separate fact-checking step, ideally against real sources, rather than trusting a single model's output as automatically accurate. This matters more for writing than for most other AI use cases, because a factual error in published writing carries real reputational cost.

Why a multi-tool workflow often beats a single "best" model

Given how differently these stages benefit from different strengths, a workflow that uses different tools for research, drafting, and review often produces better results than forcing one model through the entire process. That's a different recommendation from "pick the best model," and a more honest one, given how the underlying capabilities differ by stage.

The one universal rule regardless of which model you use
Never publish an AI-drafted piece of writing without verifying its factual claims and citations against real sources yourself. This applies whichever model produced the draft: confident, plausible-sounding, and factually wrong is a failure mode every current model can exhibit.

How we approach this

We help clients build writing workflows around what each stage actually needs (research, drafting, editing, fact-checking) rather than picking one model and expecting it to be equally strong at all of them. We also build in a real verification step for anything that gets published.