How Accurate Is Auto-Categorization? Honest Numbers

Bobby Huang

Partner, SDO CPA LLC / CEO, Growthy

July 14, 2026
8 min read
AI Bookkeeping
How Accurate Is Auto-Categorization? Honest Numbers

Julia Eskander is a CPA. She reviews client books for a living. She says she's constantly finding gross errors made by AI. She's skeptical, but open. That's the right posture for anyone deciding whether to trust software that categorizes transactions on its own.

Most vendors answer the auto categorization accuracy question with one round number and a smile. This piece answers it with the real ladder. It covers what different tools actually hit, how Growthy's own number is measured, and where the software still needs you. No inflated claims. Just the numbers.

How accurate is auto-categorization?

It depends on the tool and the books. QuickBooks Online's native rule-matching runs around 50%. That's close to a coin flip. A generic large language model like GPT-5 gets to roughly 70-71% with no training on your specific books. Growthy hits 85% on first import, before it's learned anything about your business. On returning books, Growthy's accuracy holds at that same 85%. That's the number we publish. Nobody serious should claim 95% or 100%.

Key Takeaways

  • 85% first-import accuracy is Growthy's baseline the day a new client's books connect, before any pattern learning on that client's books has happened.
  • One number for new and returning books. Growthy publishes 85% for returning books too, so there's no qualifier to lose.
  • A confidence score isn't an accuracy score. The system can be very sure and still wrong.
  • Roughly 15% of transactions stay difficult. How many need a human eye depends on the confidence threshold you set.
  • QBO's native categorization runs about 50%. A generic LLM like GPT-5 lands around 70-71% with no context on your books.
  • Review never fully goes away. At 85%, a bookkeeper still checks the queue.

What 85% first-import accuracy actually means

The 85% figure comes from Bobby Huang's own dogfood data. Bobby is a partner at SDO CPA LLC. He has 18 years of bookkeeping experience. He's not a CPA himself. He ran the comparison on his firm's own books and Growthy's books during a March 2026 window. At that point the system had learned nothing about either business yet.

First import means exactly that. Day one. Cold start. No history. Growthy has never seen this client's vendors, categories, or quirks before. It's working from patterns learned across many books, not this specific one. That's a different measurement than a steady-state number collected after months of use. Mixing the two is how marketing claims get inflated.

For contrast, Digits advertises 96% accuracy. That's their claim. We're not going to compete on that number. We can't verify how they measured it, and an unqualified 95%+ figure doesn't match what real categorization work looks like on messy books. A first-import number like 85% is lower. It's the number we can actually stand behind.

Take a transaction like a $3,847.92 Stripe deposit that doesn't match any invoice in the system. This is an illustrative example, not a real client record. On first import, that's exactly the kind of line item that can trip up any categorization engine. Human or software, there's no history to draw on yet.

The 85% number also tells you something about scope. It's measured across a full month of real transactions, not a cherry-picked sample of easy ones. Recurring subscriptions and simple vendor payments are usually easy. Odd deposits, split transactions, and one-off vendors are the ones that pull the average down. An honest first-import number has to include all of it, not just the easy ones.

What happens on returning books

Categorizes the routine. Flags what needs you.

See Growthy on a sample book. Read-only bank access.

Get started

Here's where pattern learning matters. Growthy isn't training a brand-new model on your data. It's building a working memory of how a specific set of books behaves. That means which vendor names map to which categories, which recurring charges are subscriptions versus one-time purchases, and how this business tends to code its own transactions.

On returning books, accuracy holds at about 85%. What changes is the repeat work: once a client's patterns are in place, you correct the same vendor or recurring charge far less often. We publish one number for first import and returning books, so it doesn't need a qualifier to stay honest. It doesn't reach 100% because some transactions will always need judgment that only a person can supply.

Compare that against the other end of the spectrum. QBO's native rule-based categorization runs around 50%. One bookkeeper, Natalia P., called it "optimistically random." A generic LLM like GPT-5, with no training on your books at all, lands around 70-71%. Patterns learned across many other books are the difference between those numbers and 85%, since the first-import figure comes before Growthy has seen your specific books.

Outsourced human bookkeeping lands around 80%, used cautiously as a comparison point. That number moves with the person doing the work and how many clients they juggle at once. It's a useful middle marker. Growthy's 85% sits above it, and review still applies.

Confidence is not accuracy

Skeptics like Julia are right to push back here. A system can look certain and still be wrong. A confidence score tells you how sure the system is. It doesn't tell you whether the system is right.

A raw bank string like "ACH PAYMENT 847293847 WEB" can pattern-match cleanly to a category the system has seen before. The system can report high confidence on that match. But a clean pattern match isn't the same as a correct one. The string might belong to a vendor the system has never actually verified. It might just resemble a familiar pattern closely enough to trigger a confident guess.

That gap, between "the system is sure" and "the system is correct," is exactly why review stays part of the workflow. A confident wrong answer is more dangerous than an uncertain one. It doesn't ask for a second look.

Think about how this plays out over a full month of books. A low-confidence guess gets flagged and lands in front of a bookkeeper. A high-confidence wrong guess slides straight through, unless someone happens to spot it later during reconciliation. That's why a bookkeeper's spot-checks matter even on transactions the system marked as certain, not just the ones it flagged as uncertain.

The difficult 15%: which transactions stay human

Not every transaction is equally hard. The hard ones tend to repeat. A few patterns show up again and again in the categories that need a person:

  • The $3,847.92 Stripe deposit that doesn't match any invoice on file (illustrative example).
  • Raw bank strings with no vendor context, like "ACH PAYMENT 847293847 WEB," where the description alone isn't enough to work from.
  • Clearing accounts that never return to zero. This usually signals a miscoded transaction sitting somewhere in the chain.

A number like "13 out of 247 need you" gets used sometimes to describe this slice. It's worth saying plainly: that ratio is threshold-dependent. Set the confidence bar higher and more transactions get flagged for review. Set it lower and fewer do, at the cost of more silent errors. There's no single fixed rate. The rate moves with where you set the line.

Why review never fully goes away

Growthy doesn't replace a bookkeeper's judgment. It doesn't send invoices, run collections, or remove the need for someone to check the work. Even at 85%, a share of transactions still needs a person to make the call. That's by design, not a shortfall to apologize for.

The honest version of this story is simple. Growthy's published accuracy is 85%, not 100%. The bookkeeper stays the judgment-holder. Software handles the repeatable pattern matching so a person can spend their attention on the transactions that actually need it.

Frequently asked questions

Is auto-categorization accurate enough to trust?

On first import, plan on roughly 85%, and plan on the same 85% for returning books. Either way, budget time to review the flagged transactions. Don't assume a clean pass.

What does 85% first-import accuracy mean in practice?

Out of every 100 transactions on a brand-new set of books, about 85 get categorized correctly on the first pass. That's before the system has learned anything specific about that business.

Why isn't the number higher, like 95% or 100%?

Some transactions genuinely need judgment a system can't supply on its own. Think of an unmatched deposit or an ambiguous bank description. Growthy publishes one number, 85%, and doesn't claim more.

For the buying-decision side of these numbers, read the companion checklist for evaluating AI categorization tools.

If you run client books and want to see where Growthy's numbers land on your own data, get started and bring a messy month. You'll also want to read how automated expense categorization handles the mechanics behind these numbers, and how Growthy fits into the broader picture of AI for accountants. For a wider look at how AI accounting software claims stack up against each other, see AI accounting software.


Growthy is bookkeeping software, not a CPA firm. This content is educational, not professional advice.

See It Work on Your Data

See Growthy on a sample book. Read-only bank access.

✓ 14-day free trial✓ Works with QuickBooks Online✓ 85% accuracy
Get started

Bobby Huang • Partner, SDO CPA LLC / CEO, Growthy

Partner at SDO CPA. 18 years of hands-on bookkeeping. Bobby still reconciles real client books and builds Growthy from that operating work.

View author profile

Growthy content is written and reviewed by people who keep real books. Worked examples come from real bookkeeping scenarios, and product claims are checked against what the product does today. Our editorial guidelines cover how we source, verify, and update every article.

Keep reading

ChatGPT for Bookkeeping: What Works and What Breaks
A person working on a laptop at a desk
AI Bookkeeping

ChatGPT for Bookkeeping: What Works and What Breaks

ChatGPT's official QuickBooks app is no longer read-only. See what it handles well, where it breaks on real books, and what to do about each break.

How to Use Claude Cowork for Accounting and Bookkeeping
Featured image for How to Use Claude Cowork for Accounting and Bookkeeping
AI Bookkeeping

How to Use Claude Cowork for Accounting and Bookkeeping

Set up Claude Cowork for accounting and bookkeeping, install the Small Business plugin, connect QuickBooks, and run 10+ workflows with real security detail.

Do Your Books From Bank Statements With Claude
Featured image for Do Your Books From Bank Statements With Claude
AI Bookkeeping

Do Your Books From Bank Statements With Claude

You've got a folder of bank statements and no bookkeeping software. Claude can turn that into a first draft of your books. Not finished books. A first pass you read and fix, the same way you'd check a new hire's work.