Writing
How to choose an AI model: a 15-minute comparison
updated 2026-09-17
Choose an AI model by testing it on your work. Give your current model and a contender the same task, check whether each result is correct, then compare time, cost, and corrections. Use benchmarks to make a shortlist. Use your results to choose.
Every model launch comes with a leaderboard victory lap. Fine. I still have to ship something by Friday.

What should you measure?
Check correctness first. Then measure the time, corrections, and cost needed to get a usable result. A fast, pleasant answer still has to meet the requirements.
| Signal | What to record | What good looks like |
|---|---|---|
| Correctness | Tests, required facts, or a written rubric | Meets the task's requirements |
| Time to useful output | Minutes through checking and repair | Less total time |
| Steering | Corrections needed to finish | Fewer interventions |
| Cost | Actual usage cost, where available | Fits the budget |
| Repeatability | Results on additional examples | Holds up beyond one lucky answer |
Tone matters too. I do not want to fight the model's personality while solving the problem. But I count that as workflow fit, not proof that the model is more capable.
How does the 15-minute comparison work?
Treat this as a quick screen. Some tasks need longer, and one result cannot establish reliability.
- Pick a small real task. Use a pull-request review, a short refactor, or a draft rewrite you can judge.
- Write the pass conditions. A code change might need to pass an existing test and preserve an API. A rewrite might need to retain five facts within 200 words.
- Run both candidates. Use fresh sessions, the same input, and comparable tools. Record model versions and effort settings.
- Check the output. Run the tests or apply the rubric. Count repairs as part of the task.
- Keep a provisional winner. Try it on several more examples before changing your default.
If you can, hide the model names while judging the outputs. It is easier to spot a weak answer when you are not rooting for its logo.
For important work, include typical cases, edge cases, and known failures. Repeat the test after material changes. OpenAI's evaluation guidance recommends task-specific tests, explicit criteria, and human review rather than judging only by whether an answer seems good.
What should stay the same?
Keep the task, source material, success criteria, and tool access comparable. Record anything you cannot control, including hidden instructions in a hosted chat product.
Decide what you are comparing. A model test holds the surrounding setup steady. A product test compares the whole experience, including its tools and defaults. Choose the test that answers your question.
What do AI benchmarks tell you?
Benchmarks measure performance on specified tasks under specified conditions. They are useful evidence, not a complete account of your working day.
For example, SWE-bench evaluates software issue resolution. Terminal-Bench evaluates agents on tasks in a terminal environment. Check the task set and agent setup before treating two scores as comparable.
Those results do not directly tell me how much editing my newsletter needs, whether my private repository is understood, or how many times I will have to restate a constraint. That is why I keep a local scorecard.
Should you always use maximum reasoning effort?
I test a moderate setting before paying for more. Where a product offers effort controls, compare the settings on the same task and record actual quality, latency, and usage.
Do not assume labels mean the same thing across providers. Do not assume a lower setting is always better either. The point is to find the least expensive setting that reliably meets your requirements.
I also check the surrounding instructions. Sometimes the problem is a confusing prompt or bloated context. Change one thing at a time, then repeat the test. Otherwise, you will not know what helped.
When should you switch models?
Switch when a contender improves the work you actually do. A new launch is a reason to test, not a reason to migrate everything.
Keep the winner for a week of ordinary work. Save failures. If it is better at code but worse at writing, use separate defaults. You do not need one champion for every job.
For a client-facing system, add privacy requirements, rate limits, uptime needs, and failure handling to the decision. A personal comparison is a starting point, not a production acceptance test. The website assistant guide covers that larger choice.
Frequently asked questions
Are AI benchmarks useless?
No. They help you narrow the field and understand measured capabilities. Check what was tested, then run a small evaluation on your own tasks.
Is one prompt enough to choose a model?
It can expose an obvious mismatch. It cannot establish reliability. Test more examples and repeat borderline results before making an important switch.
Should I compare models or chat products?
Compare the thing you plan to use. If tools, search, and the interface matter to your work, test the complete product and record that choice.
What if the cheapest model needs more corrections?
Count the repairs. Compare the total cost of a usable result, including your time, rather than the price of the first response.
Related: design a workflow with clear jobs and checks, or browse the AI directory for evaluation tools.