cd /news/artificial-intelligence/evaluating-ocr-on-the-community-memo… · home › topics › artificial-intelligence › article
[ARTICLE · art-146955] src=ztoz.blog ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Evaluating OCR on the Community Memory Corpus

A study evaluating optical character recognition on the Community Memory corpus of roughly 5,500 scanned 1980s Berkeley bulletin-board printouts found that current general-purpose LLMs outperform specialized OCR engines, with Gemini 3.8 surprisingly degrading versus Gemini 3.7 while Qwen3 and Gemini 3.7 showed similar price/performance ratios. The evaluation covered five models — Document AI, Gemini 3.7, Gemini 3.8, Mistral OCR, and Qwen3 — and concluded that research tasks using the OCR output must be robust to at least four character errors per line, or one word error per line. The Community Memory printouts, retained by co-founder Lee Felsenstein and later gifted to the Computer History Museum, were scanned and released as PDFs whose embedded OCR text is poor quality.

read22 min views1 publishedOct 7, 2026

The Community Memory corpus consists of approximately 5,500 scanned printouts of bulletin board-style posts made by Berkeley residents in the 1980s. Since residents posted messages through publicly accessible terminals, the postings reflect a potentially distinct and broader population than other computer-facilitated networks of the time (e.g. Usenet, The WELL). In advance of a potential project to translate the scans into an indexed, machine-readable form, we evaluate the optical character recognition (OCR) performance and economy of five models (Document AI, Gemini 3.7 and 3.8, Mistral OCR, and Qwen3) on a sample of the corpus. We find that current generation general purpose LLMs perform better than specialized OCR engines and while Qwen3 and Gemini 3.7 have similar price/performance ratios, Gemini 3.8 is surprisingly a degradation to Gemini 3.7. We conclude that research tasks using the OCR output will need to be robust to at least four character errors per line or one word error per line.

Community Memory was a series of public access social networks. The organizers placed terminals in public spaces which allowed reading and posting of messages, similar to a bulletin board. Community members could use the systems anonymously and without payment. The second generation (1984-1989) of Community Memory used Soroc IQ-120 CRT terminals, ran on a Plexus P/40 minicomputer under Unix Version 7, and connected over 2400bps modems (Felsenstein 2024, pgs 135-136). The second generation terminals were placed in three co-op markets around Berkeley and one at the La Peña Cultural Center (Felsenstein 2024, pg 139).

Community members posted messages on a wide variety of subjects, including buying and selling products, advertising special events, and politics and organizing.

Messages within the system would age out over a period of time. As they were deleted from the system, the messages were printed out for archival purposes. The printer was likely shared with the Resource One organization. Based on changes in fonts and the use of mixed case in output, the line printer was likely upgraded from the one used in the first generation system (1973-1975). The printouts were retained by Lee Felsenstein, one of the founders of Community Memory, and were later gifted to the Computer History Museum (CHM). CHM has scanned the print-outs (about 5,500 pages worth of posted messages) and made them available as PDFs for research.

Individual pages measure about 65 lines long and contain monospaced, mixed case text. Although the scans are high-quality, the printer was not. Below is an example post, demonstrating the standard structure for the entries.

Community Memory included a tagging system for posts and the developers experimented with various schemes, but unfortunately the message tags for the second generation system seem to be lost. Additionally, although some posts have non-empty “Comment On” and “Comment By” fields, the printouts do not include post ids so reconstructing message threads seems impossible.

The PDFs include embedded text but the quality of the OCR is poor. The above post has the following representation in the file, hence our interest in providing a better translation:

i Title: Moving Aaie Sunday October 14th
Comment On :
f..ommen.t By
IMesSag¢:
I MC)VINC, SALR:
Monks furniture, couch and chair;
wood/hive c~rdu:coy: 595.00
~
H OlJ.SP.hO Id furnishings, kitchen utensils,
tools, plants beautiful, books, etc.
Or.toher 14, Sunday; 10 A.M.- 4 P.M.
291$ Deakin St. Berkeley Apt. 3.
Between Russell and Ashby, 2 blocks
west of 'Telegraph. 848-4906.
~A.tathor: Anonymous
Site Entered: 166'8247408
;Rate Entered: OCT 12, 1984
Date Expires: DEC 1, 1984

As the Community Memory terminals could be used anonymously, were placed in public spaces, and often unattended, they attracted unserious users who posted junk messages or attempted to spam messages by tagging them with inappropriate topics. Below is an example junk entry in the evaluation sample that has a very low agreement between the models.

Although some posts include extended narratives and “ASCII art,” most are fairly concise and direct.

The printed page limited lines to 72 characters, but the median line was shorter than 40 characters (see figure below). The number of words in non-empty lines was also quite low, with a median of five, although this is biased by the message formatting structure.

We selected a random sample of pages from the corpus and created ground truth representations of each page (10 pages, 649 total lines). We invoked the models with a single page per call. For Qwen3, the AWS Bedrock endpoint did not accept PDFs so we translated the page into PNGs using pdftoppm. For the general purposes LLMs, we used a prompt derived from the one used in the OmniDocBench benchmark (OmniDocBench 2025):

    You are an AI assistant specialized in converting PDF images to plain text format. Please follow these 
    instructions for the conversion:

    1. Text Processing:
        - Accurately recognize all text content in the PDF image without guessing or inferring.
        - Convert the recognized text into plain text format.
        - Maintain the original document structure, including headings, paragraphs, lists, etc.

    2. Output Format:
        - Ensure the output text document has a clear structure with appropriate line breaks between elements.
        - For complex layouts, try to maintain the original document's structure and format as closely as possible.

    Please strictly follow these guidelines to ensure accuracy and consistency in the conversion. Your task is to 
    accurately convert the content of the PDF image into plain text format without adding any extra explanations
    or comments.

We evaluated five models (see table below). Document AI and Mistral OCR are both models specifically for extraction and document processing, while Gemini and Qwen3 are general purpose LLMs. The estimated costs are based on their standard commercial pricing and projected based on reported token usage from our sample.

Table: Models Evaluated

Model (Short) Type Model (Long) Vendor Year Released ~ $/1k pages Notes
Document AI OCR pretrained-ocr-v2.1-2024-08-07 2024 $1.50 (1)
Gemini 3.7 General gemini-3.7-flash 2026 $15.97 ± $3.48 (2)
Gemini 3.8 General gemini-3.8-flash 2026 $35.30 ± $23.70 (3)
Mistral OCR OCR mistral-ocr-4-1 Mistral 2026 $4.00 (4)
Qwen3 General qwen.qwen3-vl-235b-a22b Alibaba 2026 $11.35 ± $0.09 (5)

We aligned the lines of the output of each model to the ground truth by finding the minimal Levenshtein distance of the pages using a 1,1,1 cost function. By tracing the edit distance matrix (see Algorithm X in (Wagner and Fischer 1974) for one potential implementation), we projected each ground truth line into the corresponding comparable or output line.

We also evaluated a majority vote with tiebreaker ensemble approach (3-way and 4-way) on the model outputs, but found that the ensemble only improved the CERiws metric. For the other metrics, performance was worse or no better than a single model. (Lund 2014, Section 2.6.7) includes an overview of the literature and discusses when ensembles might break down.

Our hypothesis was that the OCR specific models would be inferior to the general LLMs in character-based metrics, since they would not use semantic context as evidence to improve their predictions, but that OCR models would be superior to general LLMs in word-based metrics as they would not “fix” the user generated text based on spelling rules or language patterns. Our hypothesis was based partially on our prior digitization efforts where LLMs predicted text to match examples widely found in their training data, even if the image disagreed significantly (example OREGON), as well as other case studies (e.g. (Vesalainen 2026) where Qwen2.5 normalized historical text).

We report normalized distance variants of the familiar OCR metrics (Rice 1996):

$$ CER = \frac{s + d + i}{max(|A|,|B|)} $$

$$ CER_{iws} = \frac{s + d + i}{max(|A_{iws}|,|B_{iws}|)} $$

$$ WER = \frac{s + i}{max(|A_{words}|, |B_{words}|)} $$

Errors are considered edits, with $s$ being the number of substitution edits, $d$ the the number of deletes, and $i$ the number of insertions. For character based metrics, $|A|$ is the length of the A string and $|B|$ is the length of the B or reference string. For word based metrics, $A$ and $B$ are based on words, which are contiguous series of alphanumeric characters, delimited by punctuation or whitespace. Words are further normalized in case, so ABC, aBc, and abc are all considered equal. WER does not consider deletes (from the comparison to reference list); this biases the metric towards recall-based applications. By using normalized distance, these metrics are naturally constrained to the zero to one (inclusive) range.

$$ Precision = \frac{|A_{words} \cap B_{words}|}{|A_{words}|} $$

$$ Recall = \frac{|A_{words} \cap B_{words}|}{|B_{words}|} $$

Within many of the tables, we have scaled the results from [0, 1] to [0, 100] for clarity. Notably, some practitioners have redefined the metrics to use 100 words as the basis, matching the same effect.

While Document AI performed best as measured by CER (see section below for discussion why), Gemini 3.7 outperformed the rest of the models as measured by CERiws and WER (see tables below). However, the confidence intervals for the three general LLMs show significant overlap. For the CER, CERiws, and WER metrics, lower scores are better, while for precision and recall, higher scores are better.

Table: Average (Mean) Character and Word Metrics per Line Evaluation Results (Scaled 0-100)

Model CER CERiws WER
Document AI 8.42 ± 1.14 1.49 ± 0.53 4.95 ± 1.43
Gemini 3.7 11.62 ± 2.16 0.51 ± 0.34 2.34 ± 1.03
Gemini 3.8 12.00 ± 2.21 0.58 ± 0.36 2.59 ± 1.12
Mistral OCR 33.22 ± 2.91 2.37 ± 0.99 4.89 ± 1.69
Qwen3 15.51 ± 2.42 0.80 ± 0.40 3.27 ± 1.18

Table: Average (Mean) Word List Metrics per Line Evaluation Results (Scaled 0-100)

Model Precision Recall
Document AI 94.8 ± 1.4 95.5 ± 1.4
Gemini 3.7 97.4 ± 1.1 97.6 ± 1.0
Gemini 3.8 97.1 ± 1.2 97.3 ± 1.0
Mistral OCR 96.4 ± 1.4 96.5 ± 1.3
Qwen3 96.3 ± 1.3 96.5 ± 1.2

Since errors in short lines will have greater impact on the averages, we also present word metrics by filtering out any lines with fewer than five words. This filter removes most of the lines containing database “structure” and leaves lines with narratives and richer content.

Table: Mean Word List for Lines with ≥5 Words Evaluation Results (Scaled 0-100)

Model Precision Recall
Document AI 95.9 ± 1.1 96.6 ± 1.1
Gemini 3.7 98.8 ± 0.5 98.9 ± 0.5
Gemini 3.8 98.8 ± 0.5 98.8 ± 0.5
Mistral OCR 97.7 ± 1.1 97.6 ± 1.2
Qwen3 98.0 ± 0.6 98.0 ± 0.6

Table: Ensemble (Majority Voting) Results showing Improvement (Scaled 0-100)

Model CERiws
document_ai,gemini-3.7-flash,gemini-3.8-flash^gemini-3.7-flash 0.505 ± 0.337
document_ai,gemini-3.7-flash,mistral-ocr-4-1^gemini-3.7-flash 0.505 ± 0.337
gemini-3.7-flash,gemini-3.8-flash,mistral-ocr-4-1^gemini-3.7-flash 0.507 ± 0.337
gemini-3.7-flash,gemini-3.8-flash,qwen3^gemini-3.7-flash 0.508 ± 0.337
gemini-3.7-flash,mistral-ocr-4-1,qwen3^gemini-3.7-flash 0.495 ± 0.336

The model strings show the list of models used in the ensemble, with the model after the hat being the tie-breaker. This list has been filtered to only show ensembles that are an improvement or equal to the best single model.

Of the five models, only Document AI returns results with individual lines and their locations within the page. In fact, consumers must interpret the geometric data as the extracted text is not organized in a natural “left to right, top to bottom” ordering for English documents. See the appendix for the algorithm and heuristics we used to project Document AI’s output to a plain text file.

None of the other models attempt to render the text file with high fidelity horizontal and vertical positioning. Thus, the model outputs commonly include whitespace-related errors as the figures illustrate.

Mistral’s CER score is further negatively impacted by it interpreting the structure of some of the source documents as containing a table and outputting a Markdown table. The additional characters in the output for the table format necessarily inflate the character-related error rate. Inspecting the API documentation (overview, endpoint), we believe there are ways to control the table output format (Markdown versus HTML) but no means to disable table extraction.

When whitespace is ignored, all of the models perform similarly, although Mistral’s injection of tables into certain outputs degrades its score. The distribution plots indicate that errors tend to be independent and rare.

Since lines tend to have few words, the WER plot shows stratification (e.g. at 0.2 or 1 word in five). Word alignment issues (such as splitting a word into two after intepreting a character as punctuation) cause some lines to have high error rates.

In examining the false positive and false negative word predictions, the model errors seem to be random. The table below lists words that a model predicted (but was not in the ground truth) or that the model missed where the count is greater than one. Although Document AI appears to struggle with anonymous (found 45 times in the ground truth sample), counts by words are quite low. Mistral OCR’s false negative words overlap with terms used to define the database structure, which might impact use cases involving database reconstruction. Finally, the overlap in false positive words between four of the five models may suggest errors in the dataset’s version of the ground truth.

Model False Positive Words (n>1) False Negative Words (n>1)
document_ai i:4, corament:3, 3:3, fer:3, 1:2, anonymo115:2, jones:2, wilson:2, 45:2, company:2 anonymous:7, i:4, by:3, h:3, feb:3, on:2, author:2, q:2, fbraihgjdfqrhiw6tryirtui:2, r:2
gemini-3.7-flash j:3, a:2, 1668247408:2, 8ujh4hgs:2, feb:2 bujh4hgs:2, fer:2
gemini-3.8-flash a:2, 1668247408:2, j:2, 8ujh4hgs:2, feb:2 bujh4hgs:2, fer:2
mistral-ocr-4-1 a:2, 1668247408:2, j:2, e:2, 8ujh4hgs:2, feb:2 entered:4, fer:3, title:3, bujh4hgs:2, is:2, on:2, comment:2, by:2, message:2, author:2
qwen3 a:2, on:2, 1668247408:2, 1:2, fbraihgjdfqrihw6tryirtui:2, 8ujh4hgs:2, gooh:2, feb:2 fbraihgjdfqrhiw6tryirtui:2, bujh4hgs:2, goob:2, fer:2

The economics of OCR are use-case dependent and at the scale of the Community Memory corpus, any human labor costs almost certainly swamp the API costs. However, if we focus on the WER metric, and compare percentage improvement to price increases, using Document AI as the baseline, we obtain the table below.

Model WER % Price % Price:WER Ratio
Baseline (Document AI) 0% 0% -
Gemini 3.7 53% 965% 18
Gemini 3.8 48% 2253% 47
Mistral OCR 1% 167% 137
Qwen3 34% 657% 19

Mistral OCR performs only marginally better than Document AI, but costs more than twice as much, so the Price to WER ratio indicates it is a poor trade-off. (This analysis, of course, is ignoring other metrics or other complementary factors.) Gemini 3.7 and Qwen3 have similar ratios. Even though Gemini 3.8 and 3.7 performed similarly, Gemini 3.8’s price increase (driven by the variability in thinking tokens), renders it a poor choice for this use case.

Ensembles, or combining multiple model’s outputs (or intermediate outputs) together in some fashion, is a common method to boost quality. (Lund 2014, section 2.6) lists seven types of ensemble methods and we evaluated one, voting (ibid, section 2.6.7). As we only found one metric improved with a voting ensemble, one of the possible causes is that our models lacked diversity. Without trying to measure model diversity directly (see (Kuncheva and Whitaker 2003) for some potential approaches), we instead counted the different voting results between ensembles instances. In the table below, we have applied bold formatting to the ensemble that beat the best model and italics to ensembles that matched the best performing single model (CERiws metric). All of the result’s confidence values heavily overlap.

As Gemini 3.7 was the best performing model, it may be expected it is the “tie breaker” within each highly performing ensemble. The best ensemble also has the most “diversity,” with models from Google, Mistral, and Alibaba and mean of 0.495. However, the straight Google ensemble has a mean of 0.505, statistically identical given their intervals of 0.336 and 0.337, respectively. The effect is too weak to measure.

Table: Count of Voting Results per Ensemble (649 total votes)

Ensemble Unanimous Majority Tie Broke
document_ai,gemini-3.7-flash,gemini-3.8-flash^document_ai 61 409 179
document_ai,gemini-3.7-flash,gemini-3.8-flash^gemini-3.7-flash 61 409 179
document_ai,gemini-3.7-flash,gemini-3.8-flash^gemini-3.8-flash 61 409 179
document_ai,gemini-3.7-flash,mistral-ocr-4-1^document_ai 23 186 440
document_ai,gemini-3.7-flash,mistral-ocr-4-1^gemini-3.7-flash 23 186 440
document_ai,gemini-3.7-flash,mistral-ocr-4-1^mistral-ocr-4-1 23 186 440
document_ai,gemini-3.7-flash,qwen3^document_ai 42 341 266
document_ai,gemini-3.7-flash,qwen3^gemini-3.7-flash 42 341 266
document_ai,gemini-3.7-flash,qwen3^qwen3 42 341 266
document_ai,gemini-3.8-flash,mistral-ocr-4-1^document_ai 21 202 426
document_ai,gemini-3.8-flash,mistral-ocr-4-1^gemini-3.8-flash 21 202 426
document_ai,gemini-3.8-flash,mistral-ocr-4-1^mistral-ocr-4-1 21 202 426
document_ai,gemini-3.8-flash,qwen3^document_ai 41 331 277
document_ai,gemini-3.8-flash,qwen3^gemini-3.8-flash 41 331 277
document_ai,gemini-3.8-flash,qwen3^qwen3 41 331 277
document_ai,mistral-ocr-4-1,qwen3^document_ai 18 202 429
document_ai,mistral-ocr-4-1,qwen3^mistral-ocr-4-1 18 202 429
document_ai,mistral-ocr-4-1,qwen3^qwen3 18 202 429
gemini-3.7-flash,gemini-3.8-flash,mistral-ocr-4-1^gemini-3.7-flash 157 327 165
gemini-3.7-flash,gemini-3.8-flash,mistral-ocr-4-1^gemini-3.8-flash 157 327 165
gemini-3.7-flash,gemini-3.8-flash,mistral-ocr-4-1^mistral-ocr-4-1 157 327 165
gemini-3.7-flash,gemini-3.8-flash,qwen3^gemini-3.7-flash 326 167 156
gemini-3.7-flash,gemini-3.8-flash,qwen3^gemini-3.8-flash 326 167 156
gemini-3.7-flash,gemini-3.8-flash,qwen3^qwen3 326 167 156
gemini-3.7-flash,mistral-ocr-4-1,qwen3^gemini-3.7-flash 141 259 249
gemini-3.7-flash,mistral-ocr-4-1,qwen3^mistral-ocr-4-1 141 259 249
gemini-3.7-flash,mistral-ocr-4-1,qwen3^qwen3 141 259 249
gemini-3.8-flash,mistral-ocr-4-1,qwen3^gemini-3.8-flash 144 261 244
gemini-3.8-flash,mistral-ocr-4-1,qwen3^mistral-ocr-4-1 144 261 244
gemini-3.8-flash,mistral-ocr-4-1,qwen3^qwen3 144 261 244

Given our measurements of model quality, we would like to determine if any, or all, of the models “pass” some quality threshold. Selecting between models that meet the threshold would then be a straight-forward economic decision. If they do not meet the threshold, or are marginal, then additional research and engineering is called for, or even adopting a different approach entirely. If such a threshold does not exist, then we can only characterize the impact on historical research and future efforts may improve the situation.

If we recast our metrics as the number of expected errors within a median line, we can plot the relation by varying CER and WER (see figures below). With the best performing model having a CER around 0.08, we can expect three or fewer errors within a line. Improving the CER yields only marginal benefits, so we need to tailor our methods to be robust to a small number of errors. In the median line with five words, a WER of 0.05, which is worse than any of the models measured, will produce >77% of lines correctly and >98% of lines with zero or one error. In contrast, a WER of 0.02, which is better than any of the models measured, will produce >90% of lines correctly and >99% of lines with zero or one error. Word-based methods thus need to be robust to single word errors, but on a line basis, there is only marginal improvement to be robust to two or more word errors.

Do these robustness requirements map onto typical digital research tasks? Using interviews with practicing historians, (Traub 2015) proposes a taxonomy of four digital research tasks:

For the first three of four tasks, the paper discusses the impact of precision and recall metrics on the task and where characterization of performance assists in understanding the results of an analysis. This approach cannot be extended to task four which involves such diverse and specialized techniques that defy summarization. While the paper discusses the tasks in terms of their relevance to quantitative metrics, the paper does not present any quantitative models or thresholds.

For tasks one through three, which map to typically information retrieval efforts, matching terms with an edit or slop distance of one or two should increase recall (at the cost of false positives) to cover for the majority of the expected error. Alternatively, spell correction or language correction may improve the text with higher precision, at the loss of some authenticity as actual spelling errors in the ground truth are lost.

Task four covers tasks too specialized to characterize generally. However, while the type of error matters, (Singh 2024) measured performance of LLMs under various errors or pertubations of the input text and found LLMs were robust to multiple OCR-like errors. Furthermore, some tasks that initially seem sensitive to character errors may not be using different methods. (Kerry 2026) found evidence that LLMs used to classify ASCII art did not project the art into a visual space and thus were resistant to whitespace errors. Thus, if errors are well-understood, researchers continue to find ways to study datasets robustly.

(Felsenstein 2024) Felsenstein, Lee. 2024. Me and My Big Ideas: Counterculture, Social Media, and the Future. 1st ed. FelsenSigns. https://felsensigns.com/books/.

(Kerry 2026) Luo, Kerry, Michael Fu, Joshua Peguero, et al. 2026. “ASCIIBench: Evaluating Language-Model-Based Understanding of Visually-Oriented Text.” Version 2. Preprint, arXiv, September 24. https://doi.org/10.48550/ARXIV.2512.04125.

(Kuncheva and Whitaker 2003) Kuncheva, Ludmila I., and Christopher J. Whitaker. “Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy.” Machine learning 51, no. 2 (2003): 181-207. https://link.springer.com/content/pdf/10.1023/A:1022859003006.pdf

(Lund 2014) Lund, William B. 2014. “Ensemble Methods for Historical Machine-Printed Document Recognition”. Theses and Dissertations. 4024. https://scholarsarchive.byu.edu/etd/4024

(OmniDocBench 2025) Ouyang, Linke, Yuan Qu, Hongbin Zhou, et al. 2025. “OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations.” arXiv:2412.07626. Preprint, arXiv, March 25. https://doi.org/10.48550/arXiv.2412.07626.

(Rice 1996) Rice, Stephen Vincent. 1996. “Measuring the Accuracy of Page-Reading Systems.” University of Nevada, Las Vegas. https://doi.org/10.25669/HFA8-0CQV.

(Singh 2024) Singh, Ayush, Navpreet Singh, and Shubham Vatsal. 2024. “Robustness of Large Language Models to Perturbations in Text.” Version 2. Preprint, arXiv. https://doi.org/10.48550/ARXIV.2407.08989.

(Traub 2015) Traub, Myriam C., Jacco Van Ossenbruggen, and Lynda Hardman. 2015. “Impact Analysis of OCR Quality on Research Tasks in Digital Archives.” In Research and Advanced Technology for Digital Libraries, edited by Sarantos Kapidakis, Cezary Mazurek, and Marcin Werla, vol. 9316. Lecture Notes in Computer Science. Springer International Publishing. https://doi.org/10.1007/978-3-319-24592-8_19.

(Vesalainen 2026) Vesalainen, Ari, Eetu Mäkelä, Laura Ruotsalainen, and Mikko Tolonen. 2026. “Error Patterns in Historical OCR: A Comparative Analysis of TrOCR and a Vision-Language Model.” arXiv:2602.14524. Preprint, arXiv, February 16. https://doi.org/10.48550/arXiv.2602.14524.

(Wagner and Fischer 1974) Wagner, Robert A., and Michael J. Fischer. 1974. “The String-to-String Correction Problem.” Journal of the ACM 21 (1): 168–73. https://doi.org/10.1145/321796.321811.

#!/usr/bin/env python3

import dataclasses
import grapheme
import json
import statistics

@dataclasses.dataclass
class Bounds:
    corners: list[tuple[int, int]]

    @property
    def width(self) -> int:
        return max(c[0] for c in self.corners) - min(c[0] for c in self.corners)

    @property
    def height(self) -> int:
        return max(c[1] for c in self.corners) - min(c[1] for c in self.corners)

    @property
    def top_left(self) -> tuple[int, int]:
        """return top left (min x, min y) coordinate of bounds if bounds were a rectangle"""
        min_x = min(c[0] for c in self.corners)
        min_y = min(c[1] for c in self.corners)
        return (min_x, min_y)

@dataclasses.dataclass
class Line:
    bounds: Bounds
    text: str

    @property
    def grapheme_len(self) -> int:
        """Number of user visible characters or graphemes within the text"""
        return grapheme.length(self.text)

@dataclasses.dataclass
class Page:
    height: int
    width: int
    lines: list[Line]

def from_documentai_json(documentai_json: dict) -> list[Page]:
    pages: list[Page] = []

    text = documentai_json["text"]

    for page in documentai_json["pages"]:
        height = page["dimension"]["height"]
        width = page["dimension"]["width"]

        lines: list[Line] = []
        for line in page["lines"]:
            corners = [(v.get('x', 0), v.get('y', 0)) for v in line["layout"]["boundingPoly"]["vertices"]]
            substring = ''.join([text[int(s.get('startIndex', 0)):int(s['endIndex'])] for s in line["layout"]["textAnchor"]["textSegments"]])
            substring = substring.strip()  # removes trailing newline (at least)
            lines.append(Line(Bounds(corners), substring))
        pages.append(Page(height, width, lines))

    return pages

def render_page_to_text(page: Page) -> str:
    """
    Render a Page to a text, inferring the grid system from the data within the page's lines. This assumes largely
    monospace fonts, similar to that of a computer print-out.

    :param page: Page of lines
    :return: a rendered page as a string
    """
    if not page.lines:
        return ""

    heights = [line.bounds.height for line in page.lines]
    widths = [line.bounds.width / max(line.grapheme_len, 1) for line in page.lines]
    cell_height = statistics.median(heights)
    cell_width = statistics.median(widths)


    grid: dict[int, dict[int, str]] = {}
    margin_left = int(page.width // cell_width)

    for line in page.lines:
        top_left = line.bounds.top_left
        grid_y = int(top_left[1] // cell_height)
        grid_x = int(top_left[0] // cell_width)

        margin_left = min(margin_left, grid_x)

        if grid_y not in grid:
            grid[grid_y] = {}
        grid[grid_y][grid_x] = line.text

    render: list[str] = []

    for line_idx in range(0, int(page.height // cell_height)):
        if line_idx not in grid:
            render.append("")
        else:
            render.append(render_line_to_text(grid[line_idx], margin_left))

    return "\n".join(render)

def render_line_to_text(text_starts: dict[int, str], margin_left: int) -> str:
    """
    Render a monospace line to text, starting text at least a number of cells given by the key. Text starts at
    index margin_left.

    :param text_starts: dict of position, measured from the left side of the page, to a string of text
    :param margin_left: starting position of text (subtracts from each text start position)
    :return: rendered line
    """
    if not text_starts:
        return ''
    else:
        line = ''
        last_len = 0
        for start_idx in sorted(text_starts.keys()):
            pos = start_idx - margin_left
            if last_len < pos:
                line += ' ' * (pos - last_len)
            line += text_starts[start_idx]
            last_len = pos + grapheme.length(text_starts[start_idx])
        return line
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @community memory 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evaluating-ocr-on-th…] indexed:0 read:22min 2026-10-07 · —