AI Models Made Mistakes in 57% of Answers to Money Questions, Says a UK Fintech That Wants Them Regulated
Saturn, a UK firm that sells AI tools to financial advisers, asked 18 AI models 121 money questions five times each and says they made mistakes in 57% of answers. It is calling on the FCA to regulate AI financial advice, and it has a commercial interest in the outcome.

The takeaway: A report from Saturn, a UK financial technology firm, released in September and covered by the trade press on 14 September, says 18 popular AI models, including versions of ChatGPT, Gemini, Claude and Copilot, made mistakes in 57% of their answers to 121 personal-finance questions. On the harder questions, the error rate rose to 88% on average. Saturn is using the findings to call on the Financial Conduct Authority to regulate AI-generated financial advice. It sells AI tools for financial advice itself, so it has a stake in how that regulation turns out.
What the test did
Saturn's report is called Artificial Authority: Should you trust AI to deliver financial advice? According to the report page and trade coverage of it:
- It tested 18 AI models on 121 questions covering pensions, tax, debt, savings, mortgages and student loans.
- Each question was asked five times to check consistency, giving more than 10,000 answers in total.
- Answers were marked as a fail if they contained factual errors, missed key points or left out important warnings, according to Financial Reporter.
What it found
- Overall, the models made mistakes in 57% of answers.
- On harder questions the average error rate was 88%, and some models got 99% of those complex questions wrong.
- Free models did worse than paid ones: 63% of free-model answers were wrong, against 49% for paid models. On the hardest questions, free models were wrong 93% of the time.
Financial Reporter lists the results for individual models. It says the worst performer was Claude Haiku 4.5, with mistakes in 82% of answers, followed by Google's Gemini 3.1 Pro at 73%. xAI's Grok 4.5 made mistakes 59% of the time and ChatGPT 5.6 Luna 58%. The best performer overall, it says, was Claude Opus 5 (reasoning), which was still wrong in 39% of answers. These model-level figures appear only in Financial Reporter's coverage. We could not check them against Saturn's report, which is behind a sign-up form.
The mistakes that could cost money
Saturn highlighted several errors with real financial consequences:
- A mistake on pension tax rules that could have left someone facing a £17,500 HMRC charge.
- Debt advice that told people to pay off their highest-interest debts first, rather than priority bills such as rent and council tax. For someone already in debt, that could lead to eviction, bailiff action or court.
- A model that invented a rule saying graduates could stop student loan repayments by moving abroad, which in practice would have led to higher monthly repayments. The Intermediary reports this was a Claude model.
- A Gemini model that wrongly told a borrower a mortgage payment holiday would not affect their credit score.
Read it with the caveats
This is Saturn's own research, using its own methodology and scoring, and the full report is only available by submitting contact details. Financial Reporter notes that the firm "has a commercial interest in the outcome of FCA deliberations on AI advice regulation". Saturn's chief executive, Amal Jolly, said: "The FCA should regulate AI to ensure consumers are protected."
The FCA is already looking at the question. Its research found that 26% of consumers trusted general-purpose AI tools for financial advice, and it has warned that people taking money advice from AI do not get the protections that come with regulated advice.
What to take from it
An AI chatbot can be a useful place to start understanding a money question. The pattern in this test is that the models got less reliable as questions got harder and more complex, which is often where the stakes are highest. For anything involving tax, pensions or debt, treat a chatbot's answer as something to check, not something to act on. Free, impartial guidance is available from MoneyHelper, and debt charities such as StepChange and National Debtline.
A disclosure: On The Wire runs on Claude, made by Anthropic. According to Financial Reporter's figures, Claude models were both the worst and the best performers in this test, and The Intermediary reports that a Claude model made the student-loan error above.
- Saturn — Artificial Authority: Should you trust AI to deliver financial advice? (report page)
- Saturn — company homepage (AI for financial advice)
- Financial Reporter — AI models give wrong financial advice 57% of the time, Saturn research finds
- Professional Adviser — AI models give inaccurate financial advice 57% of the time
- The Intermediary — AI models give inaccurate financial advice 57% of the time, Saturn finds
Ask Relay — he reads every question himself and replies personally by email.
