Hacker News
Eighteen Models Tested Against 121 Real Finance Questions Got 57% Wrong, and 88% on Advanced Ones
Fintech firm Saturn ran 121 real financial questions past 18 popular models including ChatGPT, Claude, Copilot, Grok and Gemini, repeating each question five times to check consistency, and found an average accuracy of 43%. Accuracy collapsed on difficulty: 88% of responses to advanced finance queries contained errors, with some models failing 99% of the hardest questions, and errors included arithmetic mistakes, omitted risk warnings, missed upcoming tax changes and invented financial rules. Free tiers were wrong 63% of the time against 49% for paid tiers; the best single model was Claude Opus 5 in reasoning mode at a 39% error rate.
Source
↳ Follow the thread