Anthropic's Claude and OpenAI's ChatGPT are both excellent, both improving fast, and both sold by companies with strong incentives to tell you theirs is best. Having built production systems on both, here's how we actually choose.
Where the differences show up
Long documents and careful reading. Claude has historically been strong with very long inputs — contracts, reports, meeting transcripts — and holds instructions well over long tasks. If your use case is "read this 80-page document and answer questions accurately", test Claude first.
Tone and drafting. Both write well. Claude tends toward a more natural, less template-shaped register out of the box, which matters for anything customer-facing. ChatGPT is highly steerable with effort. Test both on your real correspondence and read the outputs aloud — you'll know within ten samples.
Ecosystem and tooling. OpenAI's ecosystem is broader: more off-the-shelf integrations, more developers familiar with it. Anthropic has invested heavily in agentic tooling and safety controls. If a system needs to take actions rather than just answer, evaluate the current agent tooling on both — this area moves quarterly.
Price. Both offer models at several price tiers. The strategic point: route most work to cheap, fast models and reserve the expensive tier for tasks that provably need it. A well-designed system might spend under £100 a month on tasks a naive design would spend thousands on.
The per-task principle
The businesses getting the most from AI don't pick a side. A single automation might use a cheap model for classification, a mid-tier model for drafting, and a premium model for the one step needing hard reasoning. Model choice is plumbing, not identity — and it should be swappable, because both providers ship better, cheaper models every few months. We covered how to keep that flexibility in choosing a model provider.
What actually matters more than the model
- Your data. An average model with access to your quotes, history and documents beats a brilliant model without them. See RAG explained.
- The approval step. Which model drafts your customer email matters less than the fact a human approves it before it sends.
- Measurement. If you can't see what the system did each week, you can't trust it — regardless of whose logo is on the model.
Want this looked at in your business? M22 is a London-based AI consultancy. A thirty-minute call gets you an honest read on where automation would pay, and where it wouldn't.
Book a call ↗