Observed Signal · Sep 21, 2026 · Research Study · Source: t3n · Impact: 2/5 · Sentiment: Negative

Study: AI Chatbots Give Wrong Financial Answers in 57% of Cases

Executive Signal Summary

A study by technology company Saturn found that popular AI chatbots including ChatGPT, Claude, Copilot, Grok, and Gemini provide incorrect financial answers in 57% of cases on average. The error rate spikes to 88% for complex questions requiring multiple calculation steps, with some models failing 99% of the time. The best performer, Claude Opus 5 in reasoning mode, still erred in 39% of responses. The study, involving over 100 financial questions, revealed calculation errors, outdated tax information, and hallucinated rules. For example, Claude Haiku 4.5 incorrectly stated a UK saver could owe £17,500 in taxes, and another model invented a rule about student loan repayment suspension when moving abroad. Saturn's CEO Amal Jolly warned that millions of people risk significant financial losses by trusting these AI models.

Polaris7 AgentPolaris7 Strategic Assessment
High Confidence

The study highlights significant reliability issues in AI chatbots for financial advice, relevant to AI trust but not a groundbreaking industry event.

SIGNAL RADAR

Track MediaMarktSaturn Signals & Market Shifts in Real-Time

Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.

Start Free in Explorer
Free Explorer tierNo credit card requiredInstant watchlist setup

Key Takeaways & Evidence Grounding

  • Study by Saturn found AI chatbots give wrong financial answers in 57% of cases.
  • Error rate rises to 88% for complex financial questions.
  • Claude Opus 5 was the best performer, with 39% error rate.
  • One model hallucinated a non-existent student loan repayment rule.
  • Incorrect pension tax advice could cost a UK saver £17,500.

Connected Companies & Entities

6 Entities mapped

“Populäre KI-Modelle wie ChatGPT, Claude, Copilot, Grok und Gemini lieferten demnach bei Finanzfragen im Durchschnitt in 57 Prozent der Fälle...”

“Populäre KI-Modelle wie ChatGPT, Claude, Copilot, Grok und Gemini lieferten demnach bei Finanzfragen im Durchschnitt in 57 Prozent der Fälle...”

“Populäre KI-Modelle wie ChatGPT, Claude, Copilot, Grok und Gemini lieferten demnach bei Finanzfragen im Durchschnitt in 57 Prozent der Fälle...”

“Populäre KI-Modelle wie ChatGPT, Claude, Copilot, Grok und Gemini lieferten demnach bei Finanzfragen im Durchschnitt in 57 Prozent der Fälle...”

Primary Source Grounding & Direct Attribution
Direct Origin Attribution
Primary Reporting: t3n•Published: Sep 21, 2026
Original Coverage Title: “Studie zeigt: KI-Chatbots liefern bei Finanzfragen in 57 Prozent der Fälle falsche Antworten”

Related Market Signals & Shifts

Recent verified developments and strategic activity across this market segment.

Conversational AI & ChatbotsApr 20, 2026

MIT Professor: How to Prompt ChatGPT for Finance

MIT professor Andrew Lo says AI chatbots (ChatGPT, Gemini, Claude) can provide valid personal finance guidance but the usefulness of their answers depends heavily on prompt quality. Citing an Intuit CreditKarma survey (Fall 2025) the article notes broad consumer use of generative AI for financial advice—two thirds of U.S. users overall and over 80% for Millennials/Gen Z—and that ~85% of users acted on chatbot recommendations. Lo recommends detailed prompts that include goals, constraints, tax bracket, residence, assets, timelines and risk tolerance, and that users ask the model to list missing information and express uncertainty. The piece warns about hallucinations and limits of model calculations and urges verifying sources and re-running prompts.

Read assessment
Large Language Models (LLM) & AIAug 5, 2026

AI Widely Used for Financial Advice, Often Wrong

A 2026 analysis finds AI has become a primary financial adviser for many Americans — 55% now use AI to help manage money, up from 10% a year earlier — but personalized AI advice can be biased and inaccurate. MIT Sloan researchers ran simulations feeding prompts into chat models and found that users with low financial literacy received guidance that left them with nearly $50,000 less wealth by age 60; prompts written by women produced roughly $60,000 less in simulated lifetime wealth than prompts written by men. Other evaluations showed AI answered correctly about 56% of the time on personal finance questions, was deceptive or misleading 27% of the time, and outright wrong 17% of the time. The report also highlights data-privacy risks (some users shared SSNs and bank numbers) and notes AI lacks a fiduciary duty.

Read assessment
Conversational AI & LLM guidanceApr 20, 2026

MIT Professor: Write Specific Prompts for Financial AI Advice

Millions use AI chatbots such as ChatGPT, Gemini and Claude for financial advice, but the quality of recommendations depends heavily on prompt design, MIT professor Andrew Lo says. Citing an Intuit Creditkarma survey (Fall 2025), the article reports that two-thirds of U.S. generative-AI users have asked such tools for finance help (over 80% among Millennials and Gen Z) and that about 85% of users implemented the chatbots' advice. Lo recommends prompts include concrete personal details — goals, constraints, tax bracket, residence, assets, timelines, and risk tolerance — and to request identified missing information and uncertainty estimates. He also advises assigning the AI the role of a financial/advisory professional and warns users to verify outputs because hallucinations and calculation limits can distort recommendations. t3n conducted an experiment on using ChatGPT for portfolio creation.

Read assessment

Track Real-Time Market Signals & Shifts

Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.