cd /news/artificial-intelligence/from-pdf-to-evidence-building-a-prod… Β· home β€Ί topics β€Ί artificial-intelligence β€Ί article
[ARTICLE Β· art-141408] src=dev.to β†— pub= topic=artificial-intelligence verified=true sentiment=Β· neutral

From PDF to Evidence: Building a Production-Ready RAG Pipeline

A developer outlined a production-oriented RAG pipeline for processing complex PDFs such as contracts, invoices, and reports, arguing that retrieval quality depends heavily on document-processing quality. The architecture chains OCR/document intelligence, layout analysis, page-aware structure, metadata-rich chunking, hybrid keyword-plus-vector retrieval, and LLM output that returns structured findings with page-level citations as evidence. The developer recommends evaluating retrieval separately from answer generation, using precision and other retrieval metrics rather than judging only the final LLM response.

by read6 min views2 publishedSep 29, 2026

RAG is often described as a simple architecture:

Documents β†’ Embeddings β†’ Vector Database β†’ LLM

That diagram is useful for understanding the concept, but it is far from enough for a production document-processing system.

When the source documents are PDFs containing tables, scanned pages, headings, footnotes, forms, and complex layouts, the real challenge is not simply retrieving similar text.

The real challenge is:

Can the system generate an answer that can be traced back to the correct evidence in the original document?

This article describes the architecture and engineering considerations behind a production-oriented RAG pipeline for document processing.

Consider a system where users upload contracts, invoices, reports, or other business documents.

A typical workflow looks like this:

PDF
 ↓
Document Intelligence / OCR
 ↓
Layout Analysis
 ↓
Page-aware Document Structure
 ↓
Chunking + Metadata
 ↓
Search Index
 ↓
Hybrid Retrieval
 ↓
LLM
 ↓
Structured Finding
 ↓
Evidence + Citation

The important observation is that retrieval quality depends heavily on document processing quality.

If the original PDF is poorly converted into text, even the best embedding model cannot recover information that was lost during extraction.

A PDF is not necessarily a collection of plain text.

It can contain:

Simply extracting all text and splitting it every 500 tokens can destroy important relationships.

For example:

Invoice Amount: Β₯12,500,000

Tax: Β₯1,250,000

Total: Β₯13,750,000

If these values are separated incorrectly during chunking, the LLM may retrieve only part of the information.

Therefore, the ingestion pipeline should preserve document structure whenever possible.

One of the most important pieces of metadata in a document RAG system is the source page.

Instead of storing only:

{
  "text": "The contract amount is..."
}

I prefer a structure closer to:

{
  "text": "The contract amount is...",
  "document_id": "contract-001",
  "page_number": 14,
  "section": "Contract Amount",
  "chunk_id": "contract-001-page14-chunk03"
}

This gives the retrieval system additional context and, more importantly, allows the generated result to point back to the original source.

For reviewed business documents, this distinction is extremely important.

A generated statement without evidence is difficult to trust.

A generated statement with:

Finding:
The contract amount exceeds the configured threshold.

Evidence:
Contract.pdf β€” Page 14

is much easier for a human reviewer to validate.

A common approach is to split documents by a fixed number of tokens.

Every 500 tokens β†’ one chunk

This is simple, but document structure can be more important than chunk size.

A better strategy can combine:

Page 14
 β”œβ”€β”€ Contract Overview
 β”‚    β”œβ”€β”€ Contractor
 β”‚    β”œβ”€β”€ Contract Number
 β”‚    └── Contract Amount
 β”‚
 └── Payment Terms
      β”œβ”€β”€ Payment Schedule
      └── Conditions

The resulting chunks carry both the content and its context.

Pure vector search is powerful for semantic similarity.

However, business documents often contain exact identifiers:

ABC-2026-00125
Β₯13,750,000
Project ID: PJ-10293

These values may be better handled by keyword or exact matching.

This is where hybrid search becomes useful.

Conceptually:

User Query
     β”‚
     β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
     ↓               ↓
Keyword Search   Vector Search
     β”‚               β”‚
     β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
             ↓
       Result Ranking
             ↓
        Top Evidence

Keyword search can capture exact terms.

Vector search can capture semantic meaning.

Combining them gives the system more flexibility across different document types and query patterns.

A common mistake is to evaluate a RAG system only by asking:

β€œDoes the LLM produce a good answer?”

That is too late in the pipeline.

We should separately evaluate retrieval.

Precision

How many retrieved documents are actually relevant?

Recall

How many of the relevant documents did we successfully retrieve?

MRR (Mean Reciprocal Rank)

How high does the first relevant result appear?

These metrics help identify whether a problem comes from:

Retrieval
    ↓
Context construction
    ↓
Prompt
    ↓
LLM generation

Without separating these stages, it becomes difficult to determine why the system is failing.

For business applications, free-form LLM responses can be difficult to validate.

Instead of asking the model to return:

I found that the amount appears to exceed...

we can define a structured schema:

{
  "finding": "Contract amount exceeds threshold",
  "severity": "warning",
  "value": 13750000,
  "threshold": 10000000,
  "evidence": [
    {
      "document": "contract.pdf",
      "page": 14
    }
  ]
}

Now the application can validate:

before displaying the result to the user.

This creates a much stronger boundary between the probabilistic LLM layer and the deterministic application layer.

One of the most important design decisions is to treat evidence as part of the generated resultβ€”not as an optional UI feature.

Finding
   β”‚
   β”œβ”€β”€ Generated statement
   β”‚
   β”œβ”€β”€ Supporting evidence
   β”‚
   β”œβ”€β”€ Source document
   β”‚
   └── Source page

The application can then allow a reviewer to move directly from the finding to the relevant page.

This creates a human verification loop:

AI Finding
    ↓
Evidence
    ↓
Human Review
    ↓
Approve / Reject / Correct

For systems used in document review, this workflow can be more valuable than simply trying to maximize the amount of text generated by the model.

Another useful architectural principle is to distinguish between:

What the model thinks is happening

and

What evidence supports that conclusion

Finding Generation
        ↓
"Amount exceeds threshold"
        ↓
Evidence Retrieval
        ↓
Page 14
        ↓
Source text / table
        ↓
Validation

This allows the application to verify that an AI-generated finding actually has supporting evidence.

It also makes debugging easier.

If a finding is incorrect, we can ask:

That is much more actionable than simply saying:

β€œThe AI hallucinated.”

Putting these ideas together:

                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                 β”‚   PDF Upload    β”‚
                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ OCR + Layout      β”‚
                β”‚ Understanding     β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ Page-aware        β”‚
                β”‚ Document Model    β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ Chunking +        β”‚
                β”‚ Metadata          β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ Search Index      β”‚
                β”‚ Keyword + Vector  β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          ↓
                    User Query
                          ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ Hybrid Retrieval  β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ LLM Generation    β”‚
                β”‚ Structured Output  β”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          ↓
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β”‚ Evidence +        β”‚
                β”‚ Citation Validationβ”‚
                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          ↓
                    Human Review

The LLM is only one component of the system.

The surrounding engineering determines whether the system is reliable enough for production.

Building AI systems around real-world documents has changed how I think about RAG.

The difficult part is rarely:

β€œHow do I call an LLM?”

The difficult part is designing the entire pipeline around it.

A production system needs to consider:

Most importantly, the system should make it easy to verify what the AI is saying.

For many enterprise AI applications, trust does not come from the model alone.

It comes from the combination of:

Good retrieval + structured generation + verifiable evidence + human review.

That is the foundation I would use when designing a production RAG system for document-heavy applications.

RAG should not be viewed simply as:

β€œGive documents to an LLM.”

A better mental model is:

β€œBuild a reliable evidence pipeline around an LLM.”

Once you think about RAG this way, many engineering decisionsβ€”from page-aware ingestion to hybrid search and citation validationβ€”become much clearer.

── more in #artificial-intelligence 4 stories Β· sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/from-pdf-to-evidence…] indexed:0 read:6min 2026-09-29 Β· β€”