RAG or fine-tuning: how to choose without burning a quarter
A decision framework for RAG vs fine-tuning — what each actually fixes, when the combination is right, and the volume threshold where fine-tuning starts paying for itself.
The ingest-classify-extract pipeline is not the hard part of document AI. What breaks in production is messy scans, tables, and knowing when to trust an extraction — here's the review-queue pattern that actually works.
Ask most vendors to explain document intelligence and you get the same five-box diagram: ingest, classify, extract, validate, load. It's accurate, and it's also not the hard part. Any team can wire that pipeline together in a sprint with an LLM behind the extract step. What determines whether the system survives contact with real documents is everything the diagram doesn't show.
We build these systems for a living, and the failure modes are consistent enough to write down.
Vendor OCR and extraction benchmarks are run on clean scans of a single document type, usually one per file, upright, at good resolution. Report "98% accuracy" against that and it's true — for that input.
Production input is a batch of forty pages that mixes three document types, scanned on a shared office scanner at an angle, half of them phone photos from the field, with a stapled cover sheet the classifier has never seen. The same model that scored 98% on the benchmark set routinely lands at 80-85% field-level accuracy on that mix. This isn't the vendor lying — it's the benchmark answering a different question than the one you're asking. The number that matters is accuracy on your document distribution, and the only way to get it is to run the pipeline on a real sample of your own documents before you commit to an approach.
The pipeline diagram assumes one document per file. Real intake — an email with six attachments, a scanned folder of loosely related pages, a fax that merged three separate forms into one PDF — does not arrive that way.
The classify step has to first decide where one document ends and the next begins, before it can decide what type each one is. Get this wrong and every downstream step operates on the wrong unit: an extraction pulls the invoice total from the packing slip stapled behind it, and the error looks like a model mistake when it's actually a segmentation mistake three steps upstream. Building a dedicated document-boundary detection pass — separate from and prior to classification — is the difference between a demo that handles single uploads and a pipeline that survives a real inbox.
Free text extracts reasonably well even from a model with no special document training, because language models are good at language. Tables are a different problem: a line-item table with merged cells, a total row that isn't visually distinguished from the last item, columns that shift position between two invoices from the same vendor a year apart. An LLM reading a table as flattened text loses the row/column relationships that give the numbers meaning, and it will confidently return a wrong total rather than flag that it's unsure.
The fix isn't a smarter prompt. It's treating table extraction as a structured problem — using layout-aware extraction that preserves cell geometry before handing rows to the model, and scoring table extractions separately from single-field extractions in your evaluation set, because they fail differently and at a different rate.
The instinct on every document AI project is to chase full automation — 100% straight-through processing, no human touches a document. That instinct is expensive and it never fully arrives, because some fraction of real documents are genuinely ambiguous: a smudged number, a field the sender left blank, a template the system has never seen. Chasing the last few percent of automation on those cases costs more engineering time than the entire rest of the pipeline.
The pattern that holds up in production instead: every extracted field carries a confidence score, a threshold routes low-confidence fields to a human review queue, and everything above threshold goes straight through. In one of our own console builds — the kind of system where an ops team watches processing volume, auto-clear rate and a review queue for anything low-confidence — the review queue isn't a fallback bolted on for edge cases. It's the mechanism that makes the system trustworthy from the first week, because a human catches what the model is unsure about before it reaches a downstream system, instead of a wrong number silently propagating into your ERP.
Most pipelines that claim to "validate" extractions are running a formatting check — is this a valid date, does this look like a PO number. That catches typos. It doesn't catch a correctly-formatted, confidently-extracted, wrong value: a vendor name that's real but isn't in your supplier master, a PO number that's well-formed but doesn't exist in your ERP.
Real validation is a lookup against the system of record at extraction time, with a defined behaviour for a miss — route to review, flag as a new vendor, whatever your process requires — rather than loading the value and finding out three weeks later during reconciliation. This is usually the step that turns out to need the most integration work on a project, and it's the one most demos skip entirely because it requires access to systems the vendor doesn't have in a sales cycle.
If you're evaluating document AI, spend less time on the accuracy number in the pitch and more time asking: what happens to the documents it gets wrong? Is there a review queue, does it route on a real confidence score, and does validation check against your actual records or just check that a number looks like a number? A system honest about where it needs a human is more production-ready than one promising it doesn't.
That's the shape of work our document intelligence projects take — confidence-scored extraction with a review queue and validation against your systems, built for your actual document mix rather than a benchmark.
A decision framework for RAG vs fine-tuning — what each actually fixes, when the combination is right, and the volume threshold where fine-tuning starts paying for itself.
Quality is the top-cited reason AI agent pilots stall, yet barely half of teams run evaluations at all. What an AI agent evaluation framework actually has to check — trajectory, not just output — and how to build one before you scale.
AI agent memory cost is the budget line most teams discover only after their context window balloons. What naive context injection actually costs, what tiered retrieval saves, and real 2026 pricing to plan against.