"Best AI for writing" is a moving target. Worse, it's the wrong question for a lot of real writing work, because different stages of producing a piece of writing benefit from different strengths. Research, drafting, and editing aren't the same task, and treating them as one undifferentiated "writing" problem is why a single model recommendation rarely holds up well across an entire writing workflow.
Why writing quality is harder to benchmark than coding
Code correctness is at least partly objective: it runs or it doesn't, it passes tests or it doesn't. Writing quality is much more subjective: tone, voice, structure, and persuasiveness are judged differently by different readers and for different purposes. That's why "best AI for writing" rankings tend to be less reliable and more opinion-driven than coding benchmarks, and deserve more skepticism, not less.
What actually matters, broken down by the stage of writing
Research and fact-gathering
If a writing task depends on current information or needs to be grounded in real sources, what matters most is whether the tool can search and cite real, verifiable sources rather than just generating plausible-sounding claims. This is the stage where grounding in real retrieved information matters most. It's also where AI-generated citations most often turn out to be fabricated if they aren't checked against real sources.
Drafting
What matters here is how well a model can follow specific instructions about structure, tone, and length, and how naturally it can work from an outline or a set of points rather than needing to invent the structure itself. A model that drafts well from clear direction is more useful in practice than one that produces polished-sounding prose from vague instructions but ignores your actual intent.
Editing and refining
Editing existing text well requires preserving the original voice while improving clarity, structure, or argument. That's a different skill from generating original prose. A model that's strong at generating new content isn't automatically the strongest choice for refining something that's already mostly there.
Fact-checking
Because any model can produce a confident, wrong statement, a genuinely careful writing workflow includes a separate fact-checking step, ideally against real sources, rather than trusting a single model's output as automatically accurate. This matters more for writing than for most other AI use cases, because a factual error in published writing carries real reputational cost.
Why a multi-tool workflow often beats a single "best" model
Given how differently these stages benefit from different strengths, a workflow that uses different tools for research, drafting, and review often produces better results than forcing one model through the entire process. That's a different recommendation from "pick the best model," and a more honest one, given how the underlying capabilities differ by stage.
How we approach this
We help clients build writing workflows around what each stage actually needs (research, drafting, editing, fact-checking) rather than picking one model and expecting it to be equally strong at all of them. We also build in a real verification step for anything that gets published.