Before using an AI agent’s account of a document, check the representation it read. A text extraction can contain the headline and every number while losing which table heading applies to each number. A page image can preserve the layout while leaving small footnotes unreadable at the supplied resolution.
This reference is for editors and builders who need to know whether a source supports an extracted passage. Its sixteen selected checks concern information inside a document, especially PDFs. They do not measure any model’s reading accuracy. The useful question is whether the specific source feature survived the path from document to agent.
- 01Check relationships.A value without its heading, unit or qualifier can be transcribed correctly and interpreted wrongly.
- 02Record the representation.State whether the agent saw text, images, structure or a combination, with page coverage.
- 03Preserve unresolved passages.Mark inaccessible or ambiguous content before turning an extraction into a confident claim.
01 — Inspect the source feature that carries the claimInspect the source feature that carries the claim #
Use the table on the pages that matter to the decision. Text layer means machine-readable characters stored in the file. OCR, or optical character recognition, creates text from an image. Neither term promises that paragraph order, table associations or small print survived correctly. The final column identifies what to compare before using the passage.
| Complete selected editorial classification; informed by the primary sources discussed below. As of September 7, 2026. | ||
|---|---|---|
| Feature and group | What can be lost | Evidence check |
| --- | --- | --- |
| Text: Scanned paragraph | An image may have no usable text layer. | Compare recognized words with the page before quoting. |
| Text: Multiple columns | Extraction may interleave separate passages. | Follow each column’s sentence order on the rendered page. |
| Text: Repeated page furniture | Headers and page numbers can enter body text. | Separate recurring page elements from the author’s argument. |
| Text: Symbols and signs | A small symbol can change a value or relationship. | Inspect minus signs, inequality signs and superscripts in context. |
| Tables: Column headings | A value can lose the heading that defines it. | Associate each claim-bearing cell with all applicable headers. |
| Tables: Row headings | Nearby values can attach to the wrong entity. | Trace the full row label, including indented subcategories. |
| Tables: Spanning cells | One label may govern several rows or columns. | Recover the span before flattening the table into records. |
| Tables: Continued table | A page break can hide a header or continuation. | Join the relevant pages and distinguish repeated headers from data. |
| Qualifiers: Footnote marker | The marker and note can become separated. | Locate the corresponding note and preserve its scope. |
| Qualifiers: Units and bases | A heading can carry the unit or comparison base. | Record units, period and denominator with the extracted value. |
| Qualifiers: Appendix reference | The main text can depend on material elsewhere. | Open the referenced appendix before adopting its qualification. |
| Qualifiers: Coverage boundary | Only a page range or excerpt may be supplied. | Record inspected pages and mark unseen sections explicitly. |
| Images: Chart axes | The scale or baseline can be absent from text. | Read axes and units; do not infer precise values from an unclear plot. |
| Images: Legend mapping | Color or pattern can identify a series. | Match the claim to the legend in the supplied rendering. |
| Images: Alternative description | Alt text can summarize meaning without all detail. | Compare its scope with the specific visual claim you need. |
| Images: Diagram relationship | An arrow or spatial grouping can carry meaning. | Inspect connections and direction rather than reading labels alone. |
02 — What the document guidance establishesWhat the document guidance establishes #
W3C PDF3 explains the relationship between logical reading order and a tagged PDF’s structure. PDF6 addresses table markup that preserves row and column relationships. Those are accessibility techniques, not tests of an AI model, but they identify information a flat text dump can fail to communicate.
W3C PDF7 distinguishes scanned images of text from actual text and includes checking OCR output. PDF1 explains text alternatives that convey an image’s meaning. An alternative description can help interpretation without being a complete numerical transcription of a chart.
Our table applies these distinctions to agent reading. It does not claim WCAG conformance, universal PDF behavior or that a particular tool uses document tags. Inspect the tool’s actual output. If a renderer supplies only page images, the existence of tags in the original does not establish that the agent received them.
03 — Reconstruct one table before trusting its summaryReconstruct one table before trusting its summary #
Consider a hypothetical report with columns for current customers and new customers. A footnote excludes trials from only one column. An extraction containing both totals and the footnote text may still join the exclusion to the wrong population. The numbers are present; the relationship is the missing evidence.
Choose a claim-bearing cell and record its row label, every applicable column heading, unit and note marker. Compare that record with the rendered page. If the table continues over a page boundary, inspect both pages. A repeated header can be mistaken for data, while a missing header can leave later values detached from their meaning.
Do not infer that all other cells passed because one cell did. State the scope of the check. For a decision depending on the whole table, inspect the complete relevant table or obtain a source format that preserves its structure.
04 — Keep a reading coverage recordKeep a reading coverage record #
Attach a short record to the extraction: file and revision, page range supplied, representation used, sections inspected, passages unresolved and any external attachments not opened. The record should make omissions visible. “Read the report” is too broad if the tool returned selected search snippets or only the first pages.
A footnote marker without its note is a useful stop signal. Look for the matching note on the page, at the end of the section or in an appendix. If it cannot be located, retain the narrower claim that the available text supports. Do not invent an exclusion because a familiar reporting convention would normally use one.
Our citation-checking reference begins once the relevant source passage is available. This document check comes earlier: it asks whether the agent had the necessary passage and relationships to assess the claim at all.
05 — Resolve uncertainty at the missing layerResolve uncertainty at the missing layer #
If words are missing, inspect the image or rerun extraction with a suitable method. If relationships are missing, recover the table or page structure. If a chart’s values cannot be read accurately, look for its accompanying data rather than estimating a precise number from a small image. Keep extracted text and your correction distinguishable. In a hypothetical note, “footnote manually checked on page 12” tells the next reviewer more than silently inserting a qualifier. The source-independence reference explains why these different representations remain views of one evidence origin.
For audio-derived material, use the separate transcript verification reference. Speaker attribution and missing sound require checks this PDF-oriented table does not cover. A combined research task can need both records without pretending that one checklist proves the entire source set.
- Scope
- 16 document features across 4 kinds of reading evidence. The complete selected reference appears above; no claim of exhaustive coverage.
- As-of date
- September 7, 2026. Actual source collection and review date; assigned publication is September 7, 2026.
- Collection
- Read four W3C PDF techniques covering reading order, tables, text alternatives and OCR. Select claim-bearing features and define a comparison check for each. Group rows by the information that must survive extraction.
- Counting
- Each row is one selected editorial case, assigned to its displayed group. Chart widths use 45 SVG units per entry. Group sizes describe this reference, not a measured distribution.
- Sources and interpretation
- W3C techniques explain document accessibility. Applying them to agent evidence is our interpretation. No model, parser or customer document was tested. Selected checks do not cover every format or document feature.
- Exclusions
- No vendor census, model benchmark, search-volume estimate, measured savings or failure rate. Worked examples are hypothetical; no customer operations were tested.
- Gaps and limitations
- UNVERIFIED means the required evidence was not inspected, was inaccessible or remains ambiguous after inspection. A selected case can overlap others in practice; preserve the specific claim and its uncertainty.
06 — DecisionWhat to do next #
Verify the passage in the form that carries its meaning.
Pick the source feature on which the decision depends, inspect its representation and keep the unresolved parts visible. A narrower, traceable extraction is more useful than a broad claim to have read everything.
For implementation support, explore our AI transformation services.