How we work

Which AI model to use for which job, from an agency running all of them

6 min read

Somebody sends you a chart of seven logos with a column of features under each, and you are no closer to knowing which AI model to use than you were before. The charts are not wrong. They are just answering a question nobody actually has, which is what each model can do, rather than the one everybody has, which is what to open when there is a job in front of you.

We run several of these every working day - writing, research, code, images, video - across a marketing agency and the software behind it. So this is what we would tell somebody over a coffee, rather than a feature comparison. Two things up front, because they matter more than the list. The differences between the leading models are far smaller than the marketing suggests. And the thing that decides whether you get a good result is almost never which one you opened.

The short version

If the job isReach forBecause
Long documents, careful writing, anything sensitiveClaudeHolds a very large amount of text at once and stays consistent across it. Best at sounding like a person rather than like a brochure.
Everyday drafting, brainstorming, quick answersChatGPTThe broadest ecosystem and the most forgiving of a vague prompt. The safe default if you only ever learn one.
Anything already inside GoogleGeminiSits in Docs, Sheets and Gmail, and is strong on images and video in the same conversation.
Live news, what people are saying right nowGrokWired into a live feed, so it is current in a way the others are not.
Something you want to run yourself, privatelyLlama, Mistral or DeepSeekOpen weights. You can host them on your own infrastructure and nothing leaves it.

If you stop reading here you have most of the value. What follows is the part that actually changes your results.

The differences are smaller than the marketing

Every few weeks one of these models takes the lead on some benchmark and the announcement is written as though the argument is settled. A month later somebody else has it. We deliberately quote no benchmark figures in this piece, because a number like that is out of date faster than a blog post can be updated, and a stale figure is worse than no figure.

For the work an ordinary business needs - writing something, summarising something, drafting a reply, tidying a spreadsheet, explaining a contract - the top handful are close enough that you would struggle to tell them apart in a blind test. The gaps that remain are real but narrow, and they show up at the edges: very long documents, very careful reasoning, code, live information.

Choosing between the leading models is a smaller decision than almost anybody selling you a course about it would like you to believe.

Where the differences are genuinely worth knowing

Length, and holding it together

If you are working with a long document - a lease, a policy, a year of meeting notes, a full website - the useful question is how much it can hold at once and whether it stays coherent across all of it. This is the clearest remaining gap between models, and it is the one most likely to change your answer. A model that loses the thread halfway through a contract is not a small inconvenience; it is a wrong answer delivered confidently.

Whether it needs to know what happened this morning

Most models are trained up to a point in the past and then stop. Some are wired into live sources and some are not, and this is worth checking rather than assuming, because the failure mode is silent - you get a fluent, confident, out-of-date answer with nothing to indicate it is out of date.

Whether the words are the product

For anything a customer will read, tone is the whole job, and the models differ here more than they differ on capability. Some produce prose that is technically fine and unmistakably machine-written: relentlessly balanced, fond of lists, ending every paragraph by restating what it just said. That is the register that makes readers stop trusting a brand, and it is not fixed by a better model. It is fixed by giving whichever model you use a written guide to how you sound, and by a person reading what comes out.

Whether the information can leave your building

This is the one most businesses under-think. If you are working with client records, staff information or anything you are legally accountable for under POPIA, the question is not which AI model to use but where the data goes and what happens to it there. Paid business tiers generally do not train on your inputs and consumer tiers sometimes do. Open-weight models can be run on your own hardware so that nothing leaves at all, which is occasionally the only acceptable answer.

The part that actually decides your results

Here is the uncomfortable bit. In our experience the model is somewhere around a tenth of the outcome. The rest is what you gave it and who checked what came back.

  • What you gave it. A one-line prompt gets a generic answer from every model on the list. The same request with your actual brand guide, three real examples and a clear statement of what a good answer looks like gets something usable, from all of them.
  • Whether it was allowed to be wrong. A model will answer confidently whether or not it knows. If nobody checks the figure, the date or the claim, you have automated the production of plausible mistakes.
  • Who read it before it went out. Every piece of client work we produce is read by a person against a written brief before it ships. That step is not a formality, it is the reason the output is worth having.

Businesses that get very little from AI have almost always optimised the wrong end of this. They spend weeks comparing models and minutes writing the brief. The businesses getting real value usually settled on one model early, stopped thinking about it, and put the effort into the instructions and the checking.

What we would do if you were starting on Monday

  • Pick one paid business tier and use only that for a month. Paid, because the free tiers are slower, more limited and less clear about your data.
  • Write down how your business sounds and what it sells, in about a page, and paste it in at the start of anything customer-facing. This single step improves output more than switching models ever will.
  • Decide, in writing, what may never be pasted into any of them. Client records, staff information, anything under a confidentiality clause.
  • Never publish anything unread. Not because the model is unreliable, but because it is reliably confident, which is a different and more expensive problem.
  • Revisit the choice in six months. Not sooner. The field moves, but not fast enough to justify switching every time somebody announces something.

We publish how we use AI in client work in full, including what a person signs off and what never goes into a model at all. If you would rather have this conversation about your own business than read another comparison chart, a fifteen-minute call is free and there is no pitch attached to it.

All posts

Fifteen minutes, and you will know whether we are the right people

No pitch and no deck. What you are trying to grow, what is in the way, and an honest answer about whether we can help with it.