
It's a few business days into the new month, the board meeting is coming, and finance is still waiting on stragglers — expense reports that came in late, transactions that employees coded wrong, hundreds of line items that need review and recategorization before the numbers can be trusted. The CEO wants preliminary financials tomorrow; you know you're a couple of days away from anything you'd actually stand behind.
This plays out in mid-market finance departments every month. Expense categorization isn't the only thing slowing the close — reconciliations and data-gathering carry plenty of the weight too — but it's one of the genuinely fixable bottlenecks, and it's worth understanding why it drags and how machine learning changes it.
For context on what "good" looks like: APQC's benchmarking of roughly 2,300 organizations puts the median monthly close at around 6 calendar days, with top-quartile performers closing in under 5. The gap between median and top-quartile is largely about how much manual processing sits in the critical path — and categorization is squarely in that category.
The people submitting expenses don't think in GL codes, and they shouldn't have to. The distinction between "Meals & Entertainment," "Travel – Meals," and "Employee Morale" is invisible to them, so they guess — a conference fee coded as travel instead of professional development, a client dinner coded as meals instead of entertainment, a software subscription filed under office supplies. None of it is malicious; it's just that categorization requires context the submitter doesn't have. And every wrong code becomes finance's problem to catch and fix.
That fixing isn't trivial. During close week, finance staff spend real hours reviewing and re-coding expenses — and doing it under time pressure, which introduces its own errors and second review cycles. It's tedious, low-value work landing at exactly the moment the team is most stretched.
Different reviewers code the same ambiguous expense differently, based on their own read of policy. That inconsistency quietly corrupts period-over-period comparisons and trend analysis, makes variances harder to explain to leadership, and weakens audit defense when the categorization logic was never documented in the first place.
The close can't really finish until expenses are properly categorized, so any delay early — late submissions, a backlog of miscoded items — pushes everything downstream and compresses the time left for the analysis that actually matters.
Machine learning categorization works differently from the static rule tables most systems use — it learns your organization's actual patterns and improves as it goes.
A rule that says "Starbucks = Meals" is brittle. An ML model weighs multiple signals at once — the merchant, the amount, the time of day, the employee's role, surrounding transactions — so it can distinguish a client coffee from an office supply run, or recognize that charges clustered around a conference date belong to professional development rather than generic travel. It's replicating the contextual judgment a good accountant applies, at a scale a person can't sustain.
Unlike static rules, the model improves with every approved expense. It trains on your historical data to establish baseline patterns, then each time finance overrides a categorization, it learns from that correction. Critically, it assigns a confidence score to each decision — so low-confidence items get routed to a human while the high-confidence majority flow through automatically. Accuracy climbs as it sees more of your real data.
Beyond the top-level category, a capable system can apply department allocation based on the employee, the right GL and sub-account codes, project or client attribution, deductible-vs-non-deductible tax treatment, and — for multi-entity companies — allocation to the correct legal entity. That's the layered coding that eats the most manual time.
The real value isn't just auto-coding the obvious; it's isolating the genuinely ambiguous. The system surfaces expenses that don't match an employee's or department's normal pattern, first-time merchants, transactions with conflicting signals, and items that straddle categories — with an explanation of why each needs attention — so finance applies its expertise where it actually matters instead of reviewing everything.
Rather than quote invented precision, here's the honest shape of the improvement:
The magnitude depends on your expense volume, how clean your historical data is, and how much of your close time categorization actually represents. Worth measuring your own baseline before assuming a number — and worth being honest that categorization is one lever, not the whole close.
Training phase (weeks 1–3): Export 6–12 months of expense history with final approved categories — the more data, the better the starting model — and map your categories to GL codes with any special rules. The system builds its initial models from that history in the background; no data-science expertise required on your end.
Pilot phase (weeks 4–6): Run ML categorization in parallel with the manual process for a full expense cycle and compare, feed corrections back so the model learns, and set the confidence threshold that triggers human review (start conservative and relax it as trust builds).
Full deployment (weeks 7–8): Enable it for all submissions, with finance reviewing only low-confidence items and exceptions, and keep monitoring accuracy as it continues to improve with use.
Expect the automation rate to start solid and climb over the first few months as the model sees more of your real data.
"What if the AI miscodes something?"Confidence scoring plus human-in-the-loop review means low-confidence items get flagged rather than posted silently, and every correction improves the model. The honest comparison is to manual coding, which also errs — just without the confidence flag or the learning.
"Our categorization rules are too complex for AI."Complex, nuanced rules are exactly where ML does well — if a human can make the call from available data, the model can learn to replicate it, and more consistently.
"We have industry-specific categories."Because the model trains on your own historical data, it learns your specific patterns rather than a generic template.
"What about month-end adjustments?"It integrates with normal close processes — finance still makes journal entries and reclassifications, and the system learns from those too.
Set targets after establishing your baseline — the right numbers depend on where you're starting.
The point of automating categorization isn't the time savings on their own — it's what finance does with the reclaimed time. When the team isn't spending close week re-coding expenses, it can spend that time on variance analysis, forecasting, and helping the business make better decisions. The move is from processing transactions to advising on them. Categorization is one piece of getting there — but it's a piece worth fixing.
Want to see where categorization is costing you in the close? It's part of Convor's AI for Finance work — get in touch for an assessment of your close process.
