← Back to Blog
Tutorials & How-Tos

Docling PDF RAG: Chunk Tables Without Losing Their Source

A table value is useful only when its headers, units and conditions survive retrieval. This Docling walkthrough shows how to inspect table serialization, budget contextualized embedding text and retain source references without promising exact cell citations.

Written by Hamza Diaz
October 3, 202610 min read51 views

Start with the question your PDF table must answer

For inspectable Docling PDF RAG ingestion, preserve table meaning alongside source references. Start by inspecting the extracted structure, then budget the contextualized embedding text, and keep document-item references with every available provenance entry. Those locations help a reviewer inspect the source. They are not automatic proof of an exact cell citation.

A hypothetical supplier specification, not a client case

Imagine a supplier specification with component names, operating limits, units, and conditional footnotes. The retrieval question is simple on the surface: Which operating limit applies to this component under the stated conditions?

This is a hypothetical example, not an Optijara client story. Pulling back a plausible number is not good enough. The retrieved passage needs the right row label, column heading, unit, and condition. A readable export is a useful starting point, but it does not prove the chunk still carries those relationships.

The scope here is narrow on purpose: design an inspectable table-chunk record that connects embedding text to document items and available source locations. This is not a complete RAG deployment, and it is not a promise that answers will be correct. Before changing the model, inspect whether ingestion preserved the table evidence needed to answer the question.

Keep the original PDF and its structured document

Retain the permitted PDF revision beside its structured JSON. DoclingDocument represents tables, document hierarchy, layout information where available, and provenance. A plain text export is not a substitute for retaining the structured document and its item references.

Basic conversion is already covered in resources such as Simon Willison's hands-on Docling note. This walkthrough focuses on what to preserve between conversion and indexing. The procedures and examples below are proposed patterns, not reported test results.

Inspect the extracted table before choosing its serialization

Check headers, cells, units, and reading order

Use the hypothetical specification throughout: an ordinary row, a row with a long conditions cell, and a table that continues on another page. Compare the PDF with the structured extraction before splitting anything.

Check whether each value still belongs to the intended row and column. Inspect merged headers, units in captions, footnote markers, and reading order. If extraction puts a condition under the wrong component, splitting the resulting text into smaller chunks does not itself correct that association.

The conversion path produces a document that can be exported as structured JSON. The outline below is pseudocode, not a verified installation recipe. Record installed docling, docling-core, and tokenizer package versions before adapting the official examples. A version shown on a documentation site is not your installed version.

convert the permitted PDF using DocumentConverter
retain the returned DoclingDocument
export the structured document as JSON
store the JSON beside the exact PDF revision
inspect extracted tables against the original pages

Choose a representation for the retrieval question

The advanced serialization example documents MarkdownTableSerializer as an alternative to the default table representation. Compare the two by asking whether the component, limit, and condition remain understandable. Pretty output is secondary.

What inspection revealsSuggested next actionWhat to keep
Readable cells and clear header relationshipsCompare default and Markdown serializationBoth candidate texts for inspection
Broken cell boundaries or misplaced labelsInvestigate extraction settings before chunkingOriginal PDF and problematic table item
Essential visual relationships missing from textConsider a page-image retrieval routePage identity and permitted source image

For the last case, our visual document retrieval guide explains the distinct role of page-based retrieval.

The advanced example also discusses traversing OCR text nested under pictures. That option may recover otherwise omitted content, but it can add noise. Treat it as a choice to test, not a default rule for every PDF.

Use HybridChunker and inspect the text you will actually embed

Align the tokenizer with the embedding route

Docling's native chunking path operates directly on DoclingDocument. Exporting to Markdown first is optional. HybridChunker refines hierarchical chunks with token-aware splitting and compatible-peer merging.

The hybrid chunking example distinguishes chunk.text from chunker.contextualize(chunk). Keep both. In this design, the contextualized string is the embedding payload. The raw string remains useful for inspecting what the chunk contains before added context.

Align the tokenizer with the intended embedding model, then inspect the exact serialized payload. Counting only raw text misses context added afterward. Also account for any prefixes or wrappers introduced by your embedding integration.

flowchart TD A[Permitted PDF revision] --> B[DoclingDocument and retained JSON] B --> C[Inspect extracted table structure] C --> D[Choose table serialization] D --> E[HybridChunker] E --> F[Raw and contextualized text] B --> G[Item references and available provenance] F --> H[Application chunk record] G --> H H --> I[Inspect before indexing]

Retain table headers without assuming every row fits

The researched chunking documentation lists repeat_table_header=True and omit_header_on_overflow=False as defaults. These are documented settings, not defaults verified against an installed package here. Check your version's API before configuring them.

Repeated headers help a split table retain context, but headers also consume space. In the hypothetical specification, a long conditions cell may leave a row difficult to fit without losing meaning. Inspect that row's contextualized output instead of assuming the configured budget guarantees acceptance by your embedding endpoint.

The official advanced example displays contextualized outputs above its shown token setting. That is a reason to inspect your own payload. It is not evidence of a diagnosed library bug or a universal overflow behavior.

Do not silently truncate a condition to make the request fit. Flag it for another representation or explicit review. Also separate two problems that often get mixed together: splitting an already detected table, and joining separate page fragments. Header repetition does not prove continuity.

Build a chunk record that carries its source references

Keep document identity beside both text representations

Use an application-owned record to connect the embedding payload with its source. The field names below are a proposed design, not a built-in Docling schema.

document_id identifies the source in your application. source_version identifies the retained revision. chunk_id distinguishes the particular ingestion output. Store chunk_text and embedding_text separately, plus serialization_settings and tokenizer_revision.

Do not reuse an old chunk identity after switching serializers or source revisions. Record the chunking settings and embedding model revision in your ingestion manifest too. Otherwise, a later operator cannot tell whether different retrieval behavior followed a document change or a representation change.

This illustrative JSON describes the intended record contract. Nulls are placeholders, not observed source values or a record ready for indexing.

{
  "document_id": null,
  "source_version": null,
  "chunk_id": null,
  "chunk_text": null,
  "embedding_text": null,
  "doc_item_refs": [],
  "item_provenance": [],
  "serialization_settings": {},
  "tokenizer_revision": null,
  "review_status": "not_checked"
}

Retain every available item provenance entry

The advanced example exposes contributing items through chunk.meta.doc_items. The document reference describes item references and provenance fields including page_no, bbox, and charspan. Preserve the association between each item reference and its provenance entries.

A proposed record-construction procedure is:

for each chunk returned by the configured native chunker:
    retain chunk.text as chunk_text
    retain chunker.contextualize(chunk) as embedding_text
    for each item in chunk.meta.doc_items:
        retain the item's self_ref
        retain every available entry in the item's prov list
        associate each entry with that self_ref
        mark missing provenance explicitly
    attach document identity, revision and ingestion settings
    validate references against the retained structured document

Avoid selecting only prov[0]. Keep all available locations for each relevant item. An empty provenance list should remain an explicit missing-location condition, not become a fabricated page reference.

Resolve each self_ref against the exact retained structured document, not the latest conversion of a similarly named file. Preserve bounding-box coordinate conventions when passing locations to a viewer.

There is a precision boundary. A referenced table item can cover more content than the current split chunk. Its location does not automatically isolate the cell supporting an answer. Preserving provenance makes inspection possible; it does not establish that retrieval selected the right condition or that generation used it correctly.

Handle page continuations and ambiguous citations explicitly

A continued table needs more than a repeated header

Return to the hypothetical specification's continuation page. Compare headers, units, component identity, and footnotes. Check whether the first visible row is new or continues a conditions cell from the preceding page.

Discussion 704 documents practitioners dealing with page-spanning tables and continued rows. It is evidence of a practical problem, not authoritative proof that a current version always lacks multipage support.

If you add application-specific joining, retain the original fragments and their references beside the derived table. Record the joining rule and leave unresolved continuity visible. Matching column names alone should not authorize a silent merge.

A matching value is not an exact citation

Suppose the same short value appears in several cells. A string match does not identify which component and condition support the answer. Before highlighting a cell, require verified alignment between the answer, the relevant row and column, and the source location.

Discussion 4321 raises this granularity problem. Its cite_sources and confidence_scores interfaces are proposals, not verified shipped APIs to copy into an implementation.

When only item-level evidence is available, label the reference honestly as a table or page reference. Keep ambiguity visible, or route the question for review, instead of presenting a precise-looking highlight without support.

Common mistakes and operational limits

Do not confuse confidence grades with correct table cells

Common mistakes have concrete corrections:

  • Embedding bare rows: inspect labels, headers, units, and conditions in the final payload.
  • Counting only chunk.text: tokenize the contextualized text actually submitted.
  • Discarding references or keeping only the first provenance entry: preserve item associations and every available location.
  • Silently joining continuations: retain fragments and document the continuity decision.
  • Treating a source box as proof: separate location accuracy from answer correctness.

The researched confidence documentation recommends mean_grade and low_grade over internal numerical scores and marks table_score as unimplemented. Do not invent a table-confidence threshold or treat a document grade as certification of a particular cell.

Local processing still requires privacy and resource planning

Docling's advanced options document TableFormer mode and cell-matching controls. Investigate those when extraction is wrong, but do not assume tuning guarantees a repair. Additional conversion attempts and manual inspection have implementation and resource costs.

Local processing, initial model downloads, and explicitly enabled remote services are separate concerns. Prefetch required models and review configuration before expecting offline operation. Embeddings, storage, and logging require their own data-handling review. A local converter does not make the entire application local-only.

Apply source permissions to retained PDFs, extracted text, and chunk records. Review access controls before indexing sensitive conditions or notes. Set document resource limits, record package and model revisions, and decide how superseded source versions will be retired without making existing citations unresolvable.

Before indexing: a practical verification checklist

Compare representations on the same permitted PDF

The following is a proposed, unexecuted procedure. Compare default and Markdown table serialization on the same permitted PDF and retrieval question. Include ordinary rows, oversized rows, and page continuations. Record observations instead of assuming a winner.

CheckEvidence to recordHold indexing when
Environment and source identityInstalled versions, PDF revision and retained JSONThe source or configuration cannot be identified
Table extractionComparison of headers, units, rows and notes with the pageCell relationships are wrong
Embedding payloadExact contextualized text and tokenizer-length resultRequired context is missing or payload exceeds the route's limit
Reference integrityResolution of item references and retained provenanceReferences fail to resolve or missing locations are concealed
Continuation handlingOriginal fragments and any documented joining decisionRow continuity remains uncertain
Answer evaluationRetrieved passage, proposed answer and supporting conditionThe answer applies the wrong row or condition

Keep the document revision, questions, and review rules fixed during that comparison. Our guide to documenting evaluation protocols explains why results need their test conditions, not just a score.

Decide what is ready and what still needs review

If extraction is wrong, revisit conversion. If context is missing, revisit serialization or chunking. If evidence is ambiguous, preserve that uncertainty instead of issuing a more precise citation.

A valid record is not necessarily a relevant retrieval result. Our guide to recall, relevance and answer correctness explains why those outcomes need separate evaluation.

Before indexing, you should be able to open the retained revision, resolve the chunk's item references, and inspect the exact embedding text. That gives the next stage something concrete to evaluate, including a clear account of what the source locations do not prove.

Key Takeaways

  • 1Inspect table structure before chunking; splitting text does not itself repair incorrect cell relationships.
  • 2Use Docling's native chunking path on DoclingDocument and compare table serialization choices on the same source.
  • 3Inspect and tokenize the contextualized embedding payload, not only chunk.text.
  • 4Keep source revisions, item references and every available provenance entry beside both text representations.
  • 5Treat header repetition and joining page-spanning table fragments as separate operations.
  • 6Available provenance supports inspection but does not guarantee exact cell citations or correct answers.

Conclusion

Inspect structure before chunking, preserve source references beside embedding text, and be honest about what provenance can prove. A useful Docling ingestion record lets a reviewer see both the representation sent for embedding and the source material behind it, including unresolved gaps. If your team needs help designing document ingestion and retrieval around its own PDFs, Optijara offers AI consulting.

Frequently Asked Questions

Do I need to export a PDF to Markdown before using Docling's HybridChunker?

No. Native Docling chunkers operate directly on DoclingDocument. Markdown export is optional; table serialization determines the representation used within native chunking. See https://docling-project.github.io/docling/concepts/chunking/.

How do I keep table headers, and does that reconstruct page-spanning tables?

Inspect repeat_table_header and the overflow settings, then check your version's contextualized output. Repeating headers does not prove separate page fragments were correctly joined. Verify continued rows, units and notes separately. See https://docling-project.github.io/docling/concepts/chunking/.

Should I embed chunk.text or the result of chunker.contextualize(chunk)?

This tutorial uses chunker.contextualize(chunk) for embedding and keeps chunk.text for inspection. Tokenize the exact submitted text, including integration-added prefixes, with the intended embedding tokenizer. See https://docling-project.github.io/docling/_generated/examples/hybrid_chunking/.

Does Docling provenance give every RAG answer an exact table-cell citation?

No. Item provenance can cover more than a split chunk or value, and locations may be missing. Preserve available entries; precise cell highlighting requires additional verified alignment. Location alone does not prove answer correctness. See https://docling-project.github.io/docling/reference/docling_document/.

Can Docling process sensitive PDFs locally?

Yes, local execution is documented. Offline use also requires available models and deliberate configuration. Review remote-service settings and the separate embedding, storage and logging path before calling the whole application local-only. See https://docling-project.github.io/docling/usage/advanced_options/.

Sources

Share this article

Hamza Diaz

Written by

Hamza Diaz

Hamza Diaz is the founder of Optijara, where he builds practical AI agents, automation systems, and Copilot workflows for service businesses. He writes about AI operations, agent strategy, and real-world implementation for teams that want usable systems instead of hype.