I was sitting in a client meeting last month when someone casually suggested we use Claude to validate transaction reconciliations. I nodded politely, but internally I cringed. A few weeks earlier, I'd watched a model confidently flip a positive number to negative while claiming the logic was sound. That moment stuck with me. Money is one of the few domains where being "mostly right" doesn't exist. There's no partial credit when a model misses by a single cent across a thousand transactions.
This is what bothered me about the work I just read. Someone actually benchmarked whether current LLMs can handle precise financial reasoning, the kind of thing we assume they'd be good at. The results aren't just disappointing; they're revealing in ways that should matter to anyone building fintech products.
The Test: Making Models Sweat on Money
The benchmark is elegantly designed. Forty test cases, split into twenty minimal pairs. Each pair changes exactly one material fact, a currency locale, a fee direction, whether two transactions are distinct or duplicate, and the correct answer must change accordingly. The model needs to return three fields: an error status, the expected amount in cents, and the discrepancy. All three must be correct, or the case fails.
What I appreciate here is the rigor. This isn't testing whether a model can recite tax law. It's testing whether it can follow explicit rules when the stakes are numerical. Brazilian Portuguese, different rounding rules, refunds that flow in different directions, the scenarios force precise reasoning, not pattern matching.
The results? Gemini 3.7 Flash got 39 out of 40 right. GPT OSS got 38. Claude Haiku got 19. But here's what matters: when you look at paired accuracy, both variants of the same scenario answered correctly, the picture fractures. Haiku managed only two complete pairs despite passing nineteen individual cases. That's not just a weakness; that's a red flag about consistency.
Why the Correct Answer Isn't Actually Correct
The most interesting failure involves all three models. In one scenario with a completed receipt, a refund, and a fee that either leaves or enters the account, Gemini recognized the record was wrong but proposed the wrong correction. It calculated receipt minus refund, ignoring the fee entirely. The other two just returned the recorded amount unchanged.
This is the kind of error that kills me about LLM confidence. Gemini was partially correct, it knew something was off, but wrong in a way that's more dangerous than complete silence. And the others? They failed the consistency test entirely. Same scenario, different facts, same wrong answer.
I've seen this pattern in production. A model catches that something looks suspicious but proposes a correction based on incomplete logic. A human reviewer flags it, and suddenly you're dealing with liability.
My Take: This Changes How I'd Deploy LLMs Financially
Here's my honest take: if you're building anything that touches real money, you cannot use these models as decision-makers. You can use them to flag anomalies or suggest where to look, but not to calculate, validate, or reconcile.
What bothers me most isn't the accuracy percentage. It's the pair failures. Claude Haiku passed individual cases but couldn't maintain consistency across minimal variations. That tells me the model isn't reasoning through the logic, it's pattern-matching. Change the surface details slightly, and it breaks.
If I were designing a financial system today, I'd use LLMs only for triage work: parsing invoice text, categorizing transactions, suggesting which transactions need manual review. The actual calculations? I'm writing those in Python with the Decimal module, unit-tested until my eyes bleed. No negotiation.
The benchmark also reveals something about how we test AI. The researcher didn't give models calculators or search tools. They ran it once, no retries. That's realistic for production, but it also means we're testing raw reasoning under constraints. Remove those constraints and results might shift. But that's exactly the point, in production, you don't get retries for financial calculations.
What This Means for Your Architecture
# What I'm doing: hard logic, proven math
from decimal import Decimal, ROUND_HALF_UP
def calculate_balance(receipt_cents, refund_cents, fee_cents):
# No ambiguity, no magic
balance = Decimal(receipt_cents) - Decimal(refund_cents) + Decimal(fee_cents)
return int(balance)
# What I'm NOT doing:
# - Running this through an LLM without verification
# - Trusting model output without a separate calculation
# - Assuming "mostly correct" is acceptable
The Question for You
If you're working on fintech or anything touching transactions, what's your current workflow for validation? Are you testing your models on exact variations like this? Because if you're not, you're betting on luck.
Source: This post was inspired by "Cents Matter: Does a Model Notice When One Financial Fact Changes?" by Dev.to. Read the original article