AI models are giving people who ask for financial advice the wrong answers most of the time, according to new research.
It showed that the most popular AI models gave wrong answers to financial inquiries on average 57% of the time.
The likes of ChatGPT, Claude, CoPilot, Grok and Gemini only give accurate answers 43% of the time, according to the study, while on harder questions the AI models made mistakes, on average, in 88% of cases.
Some AI models gave wrong answers to 99% of the more complex questions.
The new report, ‘Artificial Authority: Should you trust AI to deliver financial advice?’, has been published by financial technology firm Saturn.
It tested 18 popular AI models against 121 different financial questions, with each question repeated five times to check consistency.
In total, more than 10,000 questions were put through the AI models with answers containing errors in calculations, missed risk warnings, ignored upcoming tax changes or hallucinated rules that did not exist. In the worst cases, the cost of following the incorrect advice could run to tens of thousands of pounds, Saturn warned.
Free-to-use models gave more inaccurate advice than paid-for models, making mistakes in 63% of answers. Paid-for models made mistakes in 49% of answers. On the hardest questions, the free models made mistakes in 93% of answers.
Amal Jolly, Saturn chief executive, said: “The low quality of financial advice from mainstream AI models risks leading to widespread consumer harm. Millions of people are trusting the AI models for money advice, but they are getting wrong answers that can lose them money.”
The worst-performing model in the survey was Claude Haiku 4.5, which made mistakes in 82% of answers. The second-worst was Google’s Gemini 3.1 Pro, failing 73% of tests. xAI’s Grok 4.5 made mistakes 59% of the time while ChatGPT 5.6 Luna made mistakes 58% of the time.
When asked more complex financial questions, Google’s Gemini 3.5 Flash and Claude Haiku 4.5 and gave wrong answers 99% of the time.
The best performing model overall was Claude Opus 5 (reasoning), which made mistakes in 39% of answers.
The research comes after the FCA published The Mills Review and found 26% of consumers trust general-purpose AI tools like ChatGPT and Claude for financial advice. It warned of the potential dangers for consumers who are left without any protection when taking financial advice from AI.
The FCA is considering whether to regulate the financial advice that AI models provide.
Mr Jolly said: “AI financial advice is currently unregulated, leaving consumers with none of the protections, including compensation, that they would get if they went to a human adviser. The FCA has started to think about this, but it needs to act fast to protect people. The FCA should regulate AI to ensure consumers are protected.”