AI & Machine Learning

LLMs Can't Do Your Accounting: Why I'm Not Trusting ChatGPT With Financial Logic

A

Adil Sher

Author

Oct 8, 2026
4 min read
4 views
LLMs Can't Do Your Accounting: Why I'm Not Trusting ChatGPT With Financial Logic

I was sitting in a client meeting last month when someone casually suggested we use Claude to validate transaction reconciliations. I nodded politely, but internally I cringed. A few weeks earlier, I'd watched a model confidently flip a positive number to negative while claiming the logic was sound. That moment stuck with me. Money is one of the few domains where being "mostly right" doesn't exist. There's no partial credit when a model misses by a single cent across a thousand transactions.

This is what bothered me about the work I just read. Someone actually benchmarked whether current LLMs can handle precise financial reasoning, the kind of thing we assume they'd be good at. The results aren't just disappointing; they're revealing in ways that should matter to anyone building fintech products.

The Test: Making Models Sweat on Money

The benchmark is elegantly designed. Forty test cases, split into twenty minimal pairs. Each pair changes exactly one material fact, a currency locale, a fee direction, whether two transactions are distinct or duplicate, and the correct answer must change accordingly. The model needs to return three fields: an error status, the expected amount in cents, and the discrepancy. All three must be correct, or the case fails.

What I appreciate here is the rigor. This isn't testing whether a model can recite tax law. It's testing whether it can follow explicit rules when the stakes are numerical. Brazilian Portuguese, different rounding rules, refunds that flow in different directions, the scenarios force precise reasoning, not pattern matching.

The results? Gemini 3.7 Flash got 39 out of 40 right. GPT OSS got 38. Claude Haiku got 19. But here's what matters: when you look at paired accuracy, both variants of the same scenario answered correctly, the picture fractures. Haiku managed only two complete pairs despite passing nineteen individual cases. That's not just a weakness; that's a red flag about consistency.

Why the Correct Answer Isn't Actually Correct

The most interesting failure involves all three models. In one scenario with a completed receipt, a refund, and a fee that either leaves or enters the account, Gemini recognized the record was wrong but proposed the wrong correction. It calculated receipt minus refund, ignoring the fee entirely. The other two just returned the recorded amount unchanged.

This is the kind of error that kills me about LLM confidence. Gemini was partially correct, it knew something was off, but wrong in a way that's more dangerous than complete silence. And the others? They failed the consistency test entirely. Same scenario, different facts, same wrong answer.

I've seen this pattern in production. A model catches that something looks suspicious but proposes a correction based on incomplete logic. A human reviewer flags it, and suddenly you're dealing with liability.

My Take: This Changes How I'd Deploy LLMs Financially

Here's my honest take: if you're building anything that touches real money, you cannot use these models as decision-makers. You can use them to flag anomalies or suggest where to look, but not to calculate, validate, or reconcile.

What bothers me most isn't the accuracy percentage. It's the pair failures. Claude Haiku passed individual cases but couldn't maintain consistency across minimal variations. That tells me the model isn't reasoning through the logic, it's pattern-matching. Change the surface details slightly, and it breaks.

If I were designing a financial system today, I'd use LLMs only for triage work: parsing invoice text, categorizing transactions, suggesting which transactions need manual review. The actual calculations? I'm writing those in Python with the Decimal module, unit-tested until my eyes bleed. No negotiation.

The benchmark also reveals something about how we test AI. The researcher didn't give models calculators or search tools. They ran it once, no retries. That's realistic for production, but it also means we're testing raw reasoning under constraints. Remove those constraints and results might shift. But that's exactly the point, in production, you don't get retries for financial calculations.

What This Means for Your Architecture

# What I'm doing: hard logic, proven math
from decimal import Decimal, ROUND_HALF_UP

def calculate_balance(receipt_cents, refund_cents, fee_cents):
 # No ambiguity, no magic
 balance = Decimal(receipt_cents) - Decimal(refund_cents) + Decimal(fee_cents)
 return int(balance)

# What I'm NOT doing:
# - Running this through an LLM without verification
# - Trusting model output without a separate calculation
# - Assuming "mostly correct" is acceptable

The Question for You

If you're working on fintech or anything touching transactions, what's your current workflow for validation? Are you testing your models on exact variations like this? Because if you're not, you're betting on luck.

Source: This post was inspired by "Cents Matter: Does a Model Notice When One Financial Fact Changes?" by Dev.to. Read the original article

Written by Adil Sher

Full stack developer building high-traffic platforms, AI services, and custom web applications. Explore my portfolio, learn about my background, or get in touch.

Related Articles

Why I'm Suddenly Worried About the LLMs I'm Shipping to Clients
AI & Machine Learning Oct 7

Why I'm Suddenly Worried About the LLMs I'm Shipping to Clients

I was building a code review assistant last month when something unsettling happened. Claude gave me a perfect explanation of why you shouldn't mutate state in React hooks. Then I pushed back, confidently, casually, with "actually, I'm pretty sure React handles that automatically n...

I Wasted $200 on an AI Tool Before Learning This Lesson
AI & Machine Learning Oct 6

I Wasted $200 on an AI Tool Before Learning This Lesson

Last month, I signed up for one of those shiny all-in-one AI platforms. The landing page had everything: ChatGPT, Claude, Gemini, image generation, document summarization. Twelve different features arranged beautifully. The price seemed reasonable for a month of "productivity." I...

I've Been Sleeping on Network Protocols While My GPUs Burn Money
AI & Machine Learning Oct 5

I've Been Sleeping on Network Protocols While My GPUs Burn Money

Last month, I got pulled into a late-night Slack thread about our training jobs taking longer than expected. My colleague casually mentioned switching from TCP to something called Homa, and I had the embarrassing realization that I'd been treating network transport as a black box...