NYC Legal Innovation Exchange at IBM, Oct. 26. Request your invitation.

New Research: AI Outperforms Experienced Lawyers in Legal Invoice Review

4 min read

A legal operations or finance professional reviewing an invoice on screen next to a printed billing guidelines document, in a polished modern office setting. Avoid robots, glowing AI brains, gavels, and generic courthouse imagery.

Legal Spend Management

Every legal department reviews outside counsel invoices the same way it always has: line by line, guideline by guideline, one tired reviewer at a time. It is slow, inconsistent, and expensive, and most legal teams have accepted that as the cost of doing business. 

A 2025 study by Onit’s AI Center of Excellence benchmarked large language models against early-career lawyers, experienced lawyers, and legal operations professionals on the same task: reviewing line items from anonymized client and synthetic invoices for compliance with billing guidelines. The result was not a close call. Two models achieved a 0.92 invoice-decision F-score, compared with 0.72 for experienced lawyers. One of those models completed reviews about 22x faster than the fastest human group. 

The results come from a published research paper. They give legal teams a reason to examine where AI-assisted review could fit in their legal spend process.  

What the research actually found 

The study, “Better Bill GPT: Comparing Large Language Models against Legal Invoice Reviewers” (arXiv:2504.02881), set out to answer a question legal operations teams have quietly wondered for years: is manual invoice review actually as accurate as it feels? Researchers built a ground-truth set of invoice decisions validated by expert legal professionals, then tested how closely different reviewer groups matched it, human and machine alike. 

Experienced lawyers, the highest-scoring human group in the study, achieved a 0.72 invoice-decision F-score. GPT-4o and Gemini 2.0 Flash Thinking each achieved 0.92. An F-score combines precision and recall; it is not a simple accuracy percentage. The speed difference was just as stark: Gemini averaged 8.68 seconds per invoice, about 22x faster than experienced lawyers, who averaged 194.75 seconds. Those times cover the review task measured in the study, not the full process of resolving and approving an invoice. 

For a function that has run on the same review model for decades, that is a meaningful signal, not a marginal improvement. 

Why manual review misses what it misses 

The accuracy gap is not really about human effort or diligence. It is about what a reviewer can reliably catch line by line, invoice after invoice, at volume. Traditional invoice validation tools, and human reviewers under time pressure, tend to check for the presence of required fields rather than the substance of what was billed. 

Consider a task description that reads: “General status and strategy work on the file, 6.2 hours.” It has a valid task code, a plausible hour count, and nothing to verify. No document, call, or decision is named, yet it passes format checks because format checks only confirm a field exists, not that the content behind it is defensible. Or take two entries describing what amounts to the same revision in slightly different language. Each line looks fine in isolation. Read in context against your guidelines, they describe duplicate work. 

These are illustrative examples, not drawn from a real matter, but they represent exactly the kind of pattern that AI-based review, working in natural language context rather than keyword rules, is built to catch. 

What this means for legal and finance teams 

None of this means AI should run unsupervised across every invoice a legal department receives. The research points to something more useful: a published, independent benchmark that legal and finance leaders can point to when deciding how much of the review process to automate, and how much to keep in human hands. 

That distinction matters in practice. Solutions like Onit’s Spend Agent are built around it, letting legal operations teams choose, vendor by vendor, whether findings route to a human for a final decision or whether compliant invoices move forward automatically. Every flagged line comes with a clear explanation tied to the specific guideline it violates, so legal and finance teams can make faster, defensible decisions instead of taking an AI’s word for it. Human oversight stays in the loop by design, not as an afterthought. 

For a General Counsel or CFO trying to get ahead of unpredictable legal spend, or a Legal Ops leader tired of chasing billing-guideline compliance manually, the research is a reason to look at what AI-assisted review actually does today, backed by data instead of a sales deck. 

See the research applied to real invoice patterns 

The full study is worth reading if you own legal spend. If you want to see how these findings translate into day-to-day invoice review, our ROI datasheet walks through the block-billing, vague-description, and duplicate-billing patterns above in more detail, and our upcoming whitepaper digs deeper into the methodology and what it means for legal operations teams building the case for AI-assisted review. 

Ready to see how AI-assisted invoice review holds up against your own billing guidelines? Book a demo with Onit and find out.