I’ve been a near-daily user of AI for a few years now. The landscape and available capabilities have drastically changed since ChatGPT was first unleashed. Through mostly trial and error, I’ve learned how to use AI effectively. I’ve seen a big difference in how AI models perform on different tasks.
After researching and experimenting with these models, I thought it would be helpful to put together a guide. The goal is to help you choose the best model for each task and to get the most out of the model.
Intelligence and Behavior
Like humans, there is no single model that is best at everything. Differences in models can be categorized into two buckets:
Raw capability. Think of this like IQ. Some models are better at problem solving, reasoning, and comprehending large amounts of information.
Behavioral characteristics. Some models are better at following instructions. They take initiative, speak up, and challenge assumptions. Some keep digging until they have exhausted a question, often going down a few rabbit holes along the way. Others will quickly answer and want to move on.
Models vs. Products Distinction
One of the first things to clarify is the distinction between models and products. People often use the terms interchangeably, which creates confusion.
A model is the underlying intelligence. Example: GPT 5.6 Sol, Claude Opus 5, or Grok 4.6.
The product is the application through which you interact with the model. Example: ChatGPT, Claude, Microsoft Copilot, Perplexity.
The mode is the amount of reasoning compute. Example: low, medium, high, max.
A model might work differently or less effectively when it’s embedded in a different product.
For example, Perplexity and Microsoft Copilot are product wrappers through which users can access various models. They then add their own tools, document handling, and instructions around those models. For this reason, you can have a different experience using the same models depending on what product you access it through.
Even when you use Claude or ChatGPT in your browser, your experience is shaped by their product layer too.
This evaluation of models will not cover an evaluation of products. Note that the benchmark scores provided in the guide are based on a stripped-down setup.
Different Tools Do Different Jobs
All that said, my first recommendation is to not pick just one model and try to use it for everything.
In terms of raw intelligence, some models are clearly ahead of the pack. Models with higher cognitive ability are better at reasoning through multiple steps. They excel at understanding complex documents and identifying non-obvious relationships.
The AI landscape is constantly changing as companies release new models all the time.
Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra top most intelligence benchmarks. Opus 5 is a close runner-up.
Intelligence is only one dimension to consider. A model’s greatest strength can also be its greatest weakness. Super intelligent models are known for overcomplicating things, changing the scope of the ask, and wasting time (and credits) going down irrelevant rabbit holes.
The smartest model might not always be the best one for the job.
Most people over-index on a model’s raw intelligence and ignore its behavioral characteristics. This is a mistake. Behavior often matters more.
The most important behavioral failure is sycophancy. This is AI’s tendency to conform to whatever you already believe instead of what is true. It’s a widespread feature of AI. A 2026 study by Carnegie Mellon and Emory tested 17 models. It found that sycophancy is a common failure mode in all of them.
AI loves to tell you what you want to hear.
Finance Industry Benchmark
Before you assign any work to a model, you need to understand what it can actually do in this domain.
Researchers test AI models against a standardized benchmark: Vals AI’s Finance Agent v2.
This tests models using 927 expert-reviewed questions. These questions match the analytical depth expected of a second or third-year investment banking analyst.
Every model gets the same toolkit: SEC EDGAR search, web search, a page parser, a retrieval tool, a calculator, and price history. Two hours per task.
It is graded by weighted checks. Certain critical facts are marked as dealbreakers. If the model misses one, the answer scores zero regardless of what else is right.
Here is the best score achieved by any model in the entire field, broken out by type of task:
The best model on the market fails roughly two out of three valuation modeling tasks. And that’s the leader, not the average. Keep that in mind.
Let this ranking give you an operating rule. Delegate the tasks AI models are strong at (retrieval and summarization). Draft-and-verify the mediocre ones (reconciliation). Leave the modeling to yourself.
The Most Expensive Models Are Not Better
The most expensive models aren’t winning.
On Finance Agent v2, Claude Sonnet 5 (53.9%) beats GPT-5.6 Sol (53.8%), Grok 4.6 (53.7%), and GPT-6 Astra (53.5%), one of the priciest models on the market. Within OpenAI’s own lineup, the budget model GPT-5.6 Luna outscores flagship Sol. Vals: mid-tier models trail the leaders by a few points, at a fraction of the price.
Save yourself the money and don’t always reach for the most expensive model thinking it will perform better.
Models are Like Employees
Think of models as employees rather than as machines or calculators. Each has different strengths, weaknesses, and quirks. The magic happens when they are in the right seat where they can play to their unique strength.
That said, here’s how I would use different models to assemble your AI dream team.
The Lean Operating Model
To avoid stacking up AI subscriptions, here’s what to buy to cover most of this. (Figures as of September 2026, subject to change.)
ChatGPT Plus + Claude Pro: $40/month
ChatGPT Plus ($20) is the first individual tier where GPT-5.6 Sol appears in standard chat, at Medium and High reasoning. Sol Pro and Extra High require the $100 or $200 Pro tiers. GPT-6 Astra appears in regular chat only on Pro, Business, and Enterprise.
Claude Pro ($20) covers three seats: Opus 5 for filings, Sonnet 5 as your associate, and Opus 5 again as the red team.
Claude Fable 5.1 is not included in Pro. Don’t upgrade to Max ($100) just for the red team. Using Opus 5 with a depersonalized prompt is an adequate replacement.
Here’s what you lose with this leaner setup:
The Market & Earnings seat. Sol is a decent substitution. Although, you’re giving up native audio and video input.
Gemini’s financial expertise. The free tier of Gemini uses Gemini 3.6 Flash. It scores 56.3% on the finance benchmark, while 3.8 Flash scores 61.4%.
The real problem with Gemini 3.6 Flash is that is allows such a small context window, you can’t even fit a 10-K in it.
If I Had to Pick
If you can only have one model, the one I would recommend depends on your workflow. Use ChatGPT Plus if your work typically starts with a question and goes out to the web. Use Claude Pro if your work usually starts with a document.







