AI models are now part of the daily toolkit of portfolio managers, analysts and traders. A recent study[1] shows that, like human analysts, these models carry biases of their own. This article summarizes two of them that matter most for investment decisions: what the model already knows about a company, and how the wording of a request can change the answer.
Note on the models. The study tested four models (Opus 5, GPT-5.6 Sol, Gemma 3 27B and Qwen3 32B), which are not necessarily the latest versions available. The results are not a benchmark or a ranking of current models. They are illustrative: they show the kind of bias investment professionals should check for in whatever tools they use.
1. AI Use in Finance
AI has moved quickly from experiment to everyday tool across the investment industry. 70% of investors now use AI to support their investment decisions[3]. Uses range from asking a chatbot to summarize an earnings call to automated systems that screen companies and feed portfolio decisions.
Adoption has moved faster than scrutiny. A February 2026 review of 164 academic papers on AI in finance found five recurring biases that can distort results, and none of them was addressed in more than 28% of the studies[2]. Some of the impressive results circulating in the industry may therefore be overstated.
The study tested these risks on real data from large US-listed companies, using two leading models of the kind used through chat interfaces and two smaller, lower-cost models of the kind often built into internal tools.
2. The Model May Already Know the Answer
AI models are trained on vast amounts of public information, which includes years of company reports, news and market commentary. As a result, a model may recognize a company from its numbers alone, even when its name has been removed.
This matters most for backtesting. A strategy that appears to pick winners from fundamentals may in fact be recognizing the companies and drawing on what it already knows about how their share prices performed. The backtest would look strong, but it would not be repeatable with live money.
To measure this, each model received real reported figures (revenue, net income, financial ratios and the reporting period) with no company name and no sector, and was asked to identify the company.
| Model | Raw figures | Compared to peers | Compared to peers, unlabeled |
|---|---|---|---|
| Gemma 3 27B | 14.7% | 5.3% | 0.0% |
| Opus 5 | 12.6% | 4.2% | 1.1% |
| GPT-5.6 Sol | 11.6% | 3.2% | 0.0% |
| Qwen3 32B | 6.3% | 4.2% | 3.2% |
Figure 1: Share of companies identified from their financial figures alone
Every model identified some companies. Recognition was uneven: Apple was identified every time and NVIDIA 75% of the time, while Microsoft was identified only 5% of the time and Tesla never. The models are not simply guessing the largest companies; they are matching specific figures they have seen before. In a backtest, a single recognized position can be enough to distort results, especially over a long period.
Removing company names is therefore not enough. Expressing the figures relative to peers reduced identification sharply, and further still when the model was not told the figures had been transformed. Even then, two of the four models still recognized 1–3% of companies.
3. Rank, Allocate or Score: the Wording Changes the Answer
With the same data and the same companies, the answer can also change with the way the request is worded. Each model received a set of six companies and was asked for the same judgment in three different ways, always with the goal of maximizing expected return:
Rank
Order the companies from best to worst.
Allocate
Split $100 across them.
Score
Rate each one from 0 to 100.
All three requests ask the same underlying question, so the ordering of the companies should be the same. The example below, from one of the leading models, shows how different the results can be.
Figure 2: Same companies, three wordings, three different orderings
Across 10 sets of six companies, the leading models were fairly but not fully consistent: all three wordings agreed on the top pick in 90% of sets for GPT and 70% for Opus. The smaller models were much less stable: ranking and scoring produced the same ordering in only 25% of sets for Gemma and 15% for Qwen. For Qwen, simply reversing the order in which the companies were listed changed the top pick in 6 of 10 sets.
Figure 3: How often the three wordings agree, by model
In practice, two analysts on the same desk, with the same data and the same tool, could ask to "rank these companies" and to "score each from 0 to 100" and reach different conclusions for no reason other than wording. The risk is highest with smaller models, which are often the ones embedded in internal tools.
Key Takeaways for Investment Professionals
-
•Treat AI-driven backtests with caution
A model may recognize companies even without their names, and draw on what it knows about their past performance. Strong backtest results should be checked for this before capital is committed.
-
•Standardize how questions are asked
Choosing a model is not enough. Teams should agree on a fixed way of asking each question, and pick the wording that best matches the actual decision (ranking, sizing or scoring).
-
•Test the tools actually in use
Behavior varies widely from one model to another, and between versions. Results from one setup should not be assumed to hold in another; each tool should be tested on its own.
Sources
- LLMs for Financial Analysis: More Capabilities, More Biases (Sabr Research, 2026)
- Evaluating LLMs in Finance Requires Explicit Bias Consideration (arXiv 2602.14233, Feb 2026)
- Amundi, 2026 annual report
Disclaimer
This content is provided for educational and informational purposes only and does not constitute investment advice, an offer to sell, or a solicitation of an offer to buy any securities or investment products. The experimental results summarized here are drawn from the study cited in source [1], use specific model versions that may not be the most recent available, and are illustrative only; they are not intended as a benchmark or comparative evaluation of any AI model or provider. The information contained herein is not intended to be a comprehensive statement of approach and is subject to change without notice. Any investment involves risk, including the possible loss of principal. Readers should consult with their financial advisors before making any investment decisions. Nothing in this article should be construed as legal, tax, or financial advice. All information is provided on an "as-is" basis without any warranty of completeness, accuracy, or timeliness.