Every retrieval demo you have ever seen used clean input. Markdown files, a folder of blog posts, maybe a tidy CSV. You chunk it, embed it, query it, and the answers come back right.
Then somebody points it at the documents the business actually runs on. A supplier contract in PDF. A quarterly report where the numbers live in a table spanning two pages. A scanned form somebody signed, photographed, and emailed.
The pipeline does not error. That is the problem.
It returns an answer, confidently, built from text that arrived in the wrong order, from a table that got flattened into a single line of digits with no column headers attached. Nothing in your stack knows this happened. Your evals pass, because your evals were written against the clean corpus too.
This is the least glamorous version of the demo-to-production gap. Your retrieval layer sits downstream of a parser, and the parser is the step that threw the structure away. (It is also the one component nobody puts on the architecture diagram.) Rebuilding that structure before anything downstream touches it is the job of a document-understanding parser like LlamaParse.
A PDF Does Not Contain a Document
This is the part that surprises people, so it is worth being precise about.
A PDF is a set of drawing instructions. Put this glyph at these coordinates, in this font, at this size. That is close to the whole format. It was designed so a page would print identically everywhere, and it succeeded completely at that job. Describing what the page means was never part of the brief.
What it does not store is any of the structure you care about.
The page content has no table object. What looks like a table is a set of text runs sitting at aligned coordinates, plus possibly some lines drawn near them.
PDF can carry table structure, in optional accessibility tags that almost nobody writes. In a study of 20,000 scholarly PDFs, 74.9% had none of it. Without those tags, the spec says matching headers to data is guesswork that “may fail for complex tables.”
So PDF guarantees you no reading order either. The drawing instructions appear in whatever sequence the generating application emitted them, which for a two-column layout frequently means the parser reads across the columns when a human reads down them.
A text extractor pulls the glyphs and gives you a string. Everything about how those glyphs were arranged, which is where a human reads the meaning, is thrown away in that step.
Your chunker then splits that string. Your embedder turns the pieces into vectors. I took the whole pipeline apart in What is RAG? If you want the mechanics of what happens after the parse. None of those later stages can recover what the parser dropped, because the information is no longer in the input.
Four Ways a Parser Loses Your Data Without Throwing an Error
Tables become sentences. A financial table has row labels, column headers, and cells, and the meaning lives entirely in which cell sits under which header. Flatten it, and you get a run of numbers. Ask “what was Q3 revenue in the EMEA segment” and the retriever finds the chunk with the right words in it, hands the model a row of digits with no headers, and the model produces a number with nothing tying it to a column.
Reading order scrambles. It happens on two-column academic papers, on invoices with a sidebar, on anything with a header block. The extractor emits text in drawing order, which can interleave the columns. The resulting chunks read as fluent nonsense: real sentences, sliced across the fold.
Scans have no text layer at all. A photographed or scanned page contains one image. A plain text extractor returns an empty string, so the document silently contributes nothing to your index. It does not fail loudly. It just is not there.
Charts and annotations carry meaning; nothing extracts. A trend line with no data table beside it, a handwritten note in a margin, a stamp. All of it is visual, and all of it disappears.
The common thread is that none of these throw an exception: the system runs, the output looks plausible, and the defect only surfaces when someone checks a number by hand.
Why Throwing a Frontier Model at It Does Not Close the Gap
The obvious move in 2026 is to skip the parser. Send page images straight to a vision-capable model and ask it for structured output.
It works, and it works well enough on a good page that a lot of teams stop there. Three things push back at production scale.
None of them show up on one page.
Cost is the first, and it is not where most people expect. A single cheap model call on a page image is competitive with a dedicated parser, sometimes cheaper. The bill grows when you drive a frontier model agentically until it reaches parser-grade accuracy, which is what a hard document actually requires.
LlamaIndex’s own ExtractBench measured that at 27.8 cents per page for a frontier coding agent against 8.1 cents for their extraction tier, with the cheaper one scoring higher. Their benchmark, so weigh it accordingly. It does match how the pricing works: every provider charges several times more for what a model writes than for what it reads, and an agentic run writes a lot.
Consistency is the second, and it breaks in the half nobody watches. Ask a general model for a table, and the shape comes back right every time, because the output format is enforced. What moves is the contents. Researchers ran identical prompts at temperature zero, the setting meant to make output repeatable, and still measured accuracy swings of up to 15 points.
On a dense table that shows up as a cell dropped here, a digit misread there. Your code gets a table that validates, with a wrong number in it, and nothing raises an error.
Page-level context is the third. A table continued across a page break needs the model to know what the previous page’s headers were, and a single-page call cannot know that.
None of this makes the approach wrong. It makes it a different tool from an ingestion pipeline that has to run on a hundred thousand pages and hand the same schema to the same downstream job every time. (The same split shows up with inference generally: what is fine for one call gets priced differently at corpus scale.)
Reach for it on the hard page. Reach for a pipeline on the corpus.
What Document-Understanding Agents Do Differently
This is the category LlamaIndex built LlamaParse for, and it is worth understanding the approach whether or not you use theirs.
The system reconstructs the document’s structure first, ahead of any text extraction. Their word for it is layout-aware parsing, OCR and language models working together: the vision layer finds the regions and the table boundaries, OCR recovers text from anything without a text layer, and the language model resolves what the reconstructed regions mean. The output is structure, tables as tables, with reading order restored.
Underneath the label, any system in this category has to do the same four jobs in order, and each one is a research problem with its own benchmarks.
Layout detection comes first. The page gets rasterized and a vision model labels regions: this block is body text, this is a table, this is a figure caption, this is a header. Nothing can be read correctly until the page has been carved up.
Reading order resolution comes second, because the regions arrive as a set and prose needs a sequence. This is where two-column layouts are won or lost, and it is a separate model decision from finding the regions in the first place.
Table structure recognition is the hard one. Finding a table’s bounding box is easy; recovering which cell belongs to which row, which column, and which header, including merged cells that span three columns, is its own field with its own metric. It is called TEDS, tree edit distance similarity, and it scores a predicted table tree against the true one.
OCR fills in anything with no text layer, and the language model assembles the labelled, ordered, structured regions into output your chunker can use.
That TEDS metric is why the specialist-versus-generalist question has an answer rather than an opinion. On OmniDocBench, a 1,355-page benchmark from CVPR 2025, GPT-4o scores 67.1 on table TEDS while a dedicated document pipeline scores 81.7 and a current fast model reaches 88.0. A fourteen-point gap on table reconstruction is not a preference. It is cells landing in the wrong column.
The performance claims are theirs, and I am attributing them: LlamaIndex says LlamaParse is several times cheaper than frontier-lab parsing and several times more accurate than other document parsing tools. Their adoption is the more checkable signal: tens of millions of monthly downloads and a GitHub repository with tens of thousands of stars, with KPMG and Pepsi among the companies they list as customers.
I have not run my own corpus through it, and you should not take a parsing accuracy claim from anybody, including me, without testing it on your documents. LlamaIndex is giving away enough credits to test exactly that.
What This Still Does Not Solve
Better parsing moves the failure further down the pipeline.
Your chunker is next in line. A perfectly reconstructed forty-row table still has to survive chunking, and a naive splitter will cut it in half. Structure-aware parsing gives you the option to chunk on real boundaries. It does not take that decision for you.
Retrieval quality is a separate problem. A correctly parsed table that your embedder never surfaces for the right query has not helped anybody.
Handwriting and bad scans remain hard. Better than a plain text extractor, which returns nothing at all, and still not solved.
You are adding a dependency on the ingest path. A parsing API sits between your documents and your index, which is a service to monitor, a bill to watch, and a vendor to evaluate on the questions you would ask any vendor.
Where to Start
Take the worst document you have. The scanned one with the merged-cell table that somebody has been fixing by hand for a year.
Run it through whatever you use now and read the raw output, ahead of any answers. Most teams discover the problem right there, in the extracted string, before any model is involved.
Then run the same page through a structure-aware parser and compare the two outputs directly. That comparison takes an afternoon, and it settles the question for your corpus, the only corpus that counts.
LlamaIndex is running an offer that covers exactly that test. Signing up is free and includes $250 in credits, no purchase required, which is enough to put a real corpus through it. Upgrade to Pro within 30 days of signing up and that is 50% off for your first three months, which puts Pro at $250 a month against a list price of $500.
Two details worth having straight. The 30-day clock starts at signup, so signing up now and deciding later is the version that keeps the discount available. And the campaign closes on 30 October.
The parser is the least interesting part of your stack and the first place your accuracy goes. Worth an afternoon before you tune another retriever.
I write The AI Engineer twice a week on production AI systems, what breaks in them, and why.