Which one should I use
The one your business already pays for — once you’ve checked that what you’re paying for does what you think it does.
That caveat is doing real work, and it’s where most of the money gets wasted.
Why the boring answer is usually right
Section titled “Why the boring answer is usually right”The assistant that comes with the software you already run is switched on, administered, and pointed at your work without anyone configuring anything per person. It can reach your mail, your files, your calendar. A separate tool can connect to a lot of that too, but somebody has to wire it up for each person and keep it working, and that somebody is you.
What a tool can see changes what you can ask for. Capability only changes how well the answer is written, and they all write well.
The rest is unglamorous and matters more than the model:
- Someone can administer it. Who has access, what your company’s terms actually are, what happens when a person leaves.
- It’s one relationship. You already have the contract, the billing, the admin console and somewhere to complain. A second vendor is a real decision, and a sticky one.
- People will actually use it. A tool already in the window they have open gets used. One that lives in another tab does not, no matter how good it is.
Check what “bundled” actually means
Section titled “Check what “bundled” actually means”The two big suites do this differently, and the difference costs money.
On Google Workspace, Gemini is included in the Business and Enterprise plans. No add-on, no second bill.
On Microsoft 365 it is split in a way that catches people out. The Copilot Chat that comes with your subscription searches the web. The Copilot that can read your mail, your files and your SharePoint is a paid per-user add-on. Microsoft’s own documentation is blunt about it: without that licence, Copilot Chat “can’t access the user’s shared enterprise data, individual data, or external data indexed via Microsoft Graph connectors.”
This is the most expensive thing to get wrong here. “We already have Copilot” usually means the free one, which cannot do the thing you are picturing — so people write the tool off having never used the version that reads their files, and some go and buy a second product they didn’t need. Find out which one you have before you judge it.
Where the default is wrong
Section titled “Where the default is wrong”Nothing came bundled. Plenty of businesses run neither suite. Then you’re choosing freely, and the answer is whichever survives the test below. Any major paid tier is a defensible start; the free tiers are fine for trying and frustrating for working.
It can’t reach the material. If the work lives in a system the assistant doesn’t connect to — your job management software, your accounting package, a shared drive it can’t see — its main advantage evaporates, and you’re comparing on capability like everyone else.
You need one specific thing. Reading very long documents, holding a real conversation out loud, writing code. These are where the assistants genuinely differ, so if your work lives in one of them, test for it directly.
You ran the test and it lost. Not “someone said the other one is better.” You gave them both the same real work and one clearly needed less fixing.
Everything else — a new release, a benchmark result, a confident post on LinkedIn — is not a reason to switch. The gap any ranking describes is almost always smaller than the gap between using one properly and using one badly, and switching costs you the background material you’ve built up in the tool plus the fluency your team built. The one thing worth noticing is a genuinely new capability rather than a better score — a tool that can now reach a system it couldn’t before, or handle a kind of document it used to refuse. A rank change is not that.
The test that settles it
Section titled “The test that settles it”An afternoon, once. Not a project, and not something to hand to whoever is most excited about the answer.
Before you start: watch what you paste. You’re about to put real work into tools you haven’t chosen yet, on consumer plans whose terms are usually worse than the business tier you’d end up on — retention, and whether your material trains the vendor’s models. Use material you’d be comfortable forwarding outside the company. If your best test case is under an NDA, use something else. The test isn’t worth the exposure.
-
Pick three pieces of work you’ve already done, and were happy with. A quote you sent, an awkward email you wrote, a long document you had to summarise for someone. Real ones — you need cases where you already know what good looks like, or you end up judging fluency, which is the one thing all of them have.
-
Mix them deliberately. One task you’re good at, so you can tell whether the output is actually right. One you find tedious, because that’s where the hours are. One in the middle.
-
Same material, same brief, each candidate. Paste the actual context — the thread, the notes, the document — not a description of it. A tool working from a summary of your problem will lose to one working from your problem.
-
Test the tier you’d actually buy. Free tiers hold something back — a smaller model, tighter caps, features switched off — and which one varies by vendor, so a free comparison tells you about free tiers and not about your decision. That usually means paying for a month. Put the cancellation date in your calendar the day you sign up: the price is quoted monthly and the commitment often isn’t.
-
Score on one question: did this need less correcting than doing it myself? Not which sounded better. Not which you enjoyed reading. Whether the finished, corrected, actually-sendable version took you less time.
Then stop. Pick the winner, or the bundled one if it was close, and go and get fluent with it. If it comes out close, that is the answer — the choice doesn’t matter much, and you’ve saved yourself the argument. Treat the result as a snapshot: re-run it when something changes in your work, not every time something changes in the news.
Why the leaderboards won’t answer this
Section titled “Why the leaderboards won’t answer this”There are public rankings of AI models, some of them good. You will be shown one. Here is what it is actually measuring.
They rank models, not products. You buy a subscription: an interface, connections to your data, an admin console, a price. They rank the engine underneath, which several products share and which changes without the product changing. No ranking can tell you whether the tool can see your files.
The most popular one measures preference, not correctness. The best-known ranking works by showing people two anonymous answers and asking which they prefer. There’s no factuality check anywhere in that loop, so a confident, well-formatted, wrong answer wins. And because the labs can see which styles win, models get tuned to produce them — longer, more agreeable, more formatted. That’s exactly the failure that costs you money elsewhere. The site publishes a separate “style control” ranking that tries to subtract the effect, which is itself an admission that it’s real.
The benchmarks are shakier than they look. A review of 445 published benchmarks by 29 expert reviewers found that barely half offered any evidence their test measures the thing it claims to measure, and one in five never defined that thing at all. Stanford’s AI Index adds the other half of the problem: tests built to stay hard for years are getting maxed out in months. The criticism comes from inside the field.
The one the open-model world relied on was shut down by the people who ran it. For a couple of years the reference ranking for freely downloadable models was Hugging Face’s Open LLM Leaderboard. In March 2025 its maintainers ended it, saying it was going obsolete and could “encourage people to hill climb irrelevant directions in the field” — that it was pushing the field to compete on things which had stopped mattering.
Who runs them
Section titled “Who runs them”- Arena — the crowd-voting one, formerly LMArena. A venture-backed company; several of its investors also hold stakes in the labs whose models it ranks, and its own funding announcement describes those labs as adopters who use its feedback to improve their models. The ranked parties are also the customers.
- Artificial Analysis — the most useful of the live ones if you’re only going to look at one, because it puts cost and speed next to capability instead of capability alone. Also a for-profit whose named backers include a sitting OpenAI board member and the chief executive of Hugging Face.
- Epoch AI — a nonprofit, transparent about its methods, and researcher-facing. Not exempt either: OpenAI commissioned and funded its hardest maths benchmark and owns the problems, which Epoch disclosed late and was criticised for.
- Stanford’s AI Index — annual, academic, free. If you read one thing a year about where this technology stands, read this rather than a ranking.
There is no Consumer Reports for this. Everyone measuring it is close to it, which is not a scandal — it’s the shape of a young field, and the reason to read any of them knowing who is paying for what. What you’ll mostly find instead, searching for “best AI model”, is affiliate content: comparison articles built to rank in search results and collect a commission.
Then go and get good at it
Section titled “Then go and get good at it”The choice matters less than what you do next, by a wide margin. Whichever one you land on, the free training from the company that makes it beats anything you’ll piece together yourself.
Ran the test and got a result that surprised you, or think we’ve called this wrong? Tell us — that’s the kind of thing that improves this page.