One in ten answers (11%) provided by AI models to common pension questions could harm savers or lead to mistakes they cannot undo, according to a new study.
The research by pension consolidator PensionBee echoes a similar survey published last month which found that the most popular AI models gave wrong answers to financial inquiries on average 57% of the time.
The likes of ChatGPT, Claude, CoPilot, Grok and Gemini only gave accurate answers 43% of the time, according to the study by financial technology firm Saturn, while on harder questions the AI models made mistakes, on average, in 88% of cases. Some AI models gave wrong answers to 99% of the more complex questions.
In the latest study by PensionBee, while almost nine in 10 (89%) answers were accurate, accuracy did not always mean an answer was safe, PensionBee warned.
The firm said that while the answers were not factually wrong, they left out crucial details that changed the answer. For example, excluding that transferring a defined benefit pension worth more than £30,000 legally requires regulated advice.
More than half, 58%, of potentially harmful answers were judged broadly accurate or better, with the accuracy score judged on whether the answer was correct, not whether the answer was complete.
Summary of harm and accuracy for major AI services
Measure | Copilot | ChatGPT | Gemini | Claude |
Harm rate | 8.1% | 9.6% | 10.4% | 14.1% |
Answers scoring 2 or more for accuracy | 94.0% | 91.0% | 87.3% | 85.3% |
Answers scoring 3 for accuracy (full marks) | 80.7% | 69.6% | 70.9% | 66.7% |
Answers scoring under 2 | 2.2% | 3.7% | 9.7% | 10.4% |
Answers scoring zero | 0% | 1% | 4% | 6% |
Source: PensionBee AI Pensions Stress Test, 2026
The findings come after the FCA’s Mills Review of AI found that just 9% of adults received regulated financial advice about their pensions or investments, leaving millions to make major pension decisions on their own.
Meanwhile nearly a third of people who engaged with their pension in the past year used AI to help them do so.
Becky O'Connor, head of pensions at PensionBee, said: "Using AI chatbots for pension advice can be a bit like playing Russian roulette with your retirement planning. While for the most part it gets things technically right; the confident, helpful tone of answers occasionally masks some worrying omissions, it may fail to detect vulnerability, or just straight up get things wrong.”
She said AI can help bridge the advice gap by giving people useful pension information when they might otherwise struggle. It can make complicated subjects more accessible and help people get started.
Ms O’Connor added: “But these results also show why consumers need to understand the limits of what an AI chatbot can safely tell them, because according to our research, one in ten times, it could turn out to be a false friend."
• The AI Pensions Stress Test was a human-run test of Copilot, ChatGPT, Gemini and Claude, to see how accurately they answer common pension questions. Human testers manually asked each AI chatbot 45 questions, three times each, using fresh accounts and clean conversations. The questions were based on real saver queries and had correct answers agreed and locked before testing, each one sourced to an official UK body.