Using AI to research stocks can save you so much time. It reads a 10-K in less time than I spend opening a 10-K. It also makes mistakes constantly. Costly mistakes. Those can be hard to catch because AI presents them with such a confident tone.
Imagine your AI is like an intern. A really smart intern: fast, well-read, and eager to help. But also new to the job and have a “fake it ‘til you make it” kind of attitude. It doesn’t tell you when it doesn’t know an answer. It just makes up an answer that sounds plausible.
Research shows that many of AIs failure points are predictable. More on that below.
It’s Predicting, Not Remembering
Every AI model has a training cutoff. This is the date when it last read anything. After that, its training data stops.
Researchers Lopez-Lira, Tang, and Zhu found that when you ask a model about anything before its cutoff date, the model remembers an answer based on its training data. It is not forecasting or predicting anything. This applies to stock prices, economic data, or earnings calls.
The researchers tried telling the model to ignore what it knows. They also tried hiding the company’s name. Neither of these approaches worked. When you ask the model about anything after the cutoff, the answers are rarely accurate.
Some AI tools connect to the live internet. ChatGPT, Perplexity, and Claude can search the web when you ask something current. That’s a separate tool layered on top of the model. It does not update the model’s underlying training data.
It Tells You What You Want to Hear
In a 2025 study, researchers had the user ask AI a question with a wrong opinion. The model agreed with the user, on average, 63.7% of the time. Various individual models agreed with the wrong opinion 46.6-95.1% of the time.
Research has also found that telling the model to act like an expert in the subject does not improve the result.
For tasks that involve math or coding, telling the AI to act as an expert can actually make its answers worse. No facts have been added to its training data that would impart expertise.
Ask It Three Times
Researchers tested the same finance questions with AI models 50 times for each of five tasks. The testing produced over 3.4 million answers in total. Simple tasks, like sorting a sentence into positive or negative, often returned the same result. Complicated tasks, like writing a summary or making a prediction, varied a lot from run to run. A more advanced model wasn’t necessarily a more consistent one.
Asking the same question three to five times and comparing the answers made the results more reliable.
The Portfolio a Chatbot Hands You
Researchers studied how an AI Chatbot would manage a stock portfolio. They tracked daily stock picks from ChatGPT 5.0 and 5.2, Claude Sonnet 4.5, Gemini 2.5 Flash, and Grok 4.1.
The AI portfolios put about 41% of their money into semiconductor stocks, compared to about 20% for the S&P 500. Their average stock carried a beta near 1.6, meaning it swings in either direction roughly 60% more than the overall market in either direction. When researchers told the models to keep the beta between 0.9 and 1.1, the models steered toward the top of the range and sometimes disregarded the instructions and exceeded it.
The models tended to pick whatever stocks were being discussed in headlines, and build concentrated positions in those names.
Keep in mind that this study covers eight months of an unusually strong stock market. Nobody has tested what these AI portfolios do once the market turns down.
The Checklist
AI is good at quickly ingesting a pile of numbers and telling you what they mean. If you give it a company's raw financial statements with the name stripped out, GPT-4 called the direction of future earnings right 60% of the time, beating the average human analyst's 53%.
When it comes to multi-step judgment calls, the AI answer often falls apart.
Ask these questions before you trust an answer from AI.
If you ask AI to pick stocks or construct a portfolio, it will likely suggest a concentrated bet on whatever’s trending. AI can be highly useful for stock research, if it’s used effectively.



