← Back to the text

How the text was made

The four gospels of Pallis’s 1910 New Testament, converted from page scans to a searchable text by optical character recognition (OCR). Status: October 2026.

Source

Η Νέα Διαθήκη κατά το Βατικανό χερόγραφο, μεταφρασμένη από τον Αλέξ. Πάλλη. Μέρος πρώτο, third and fourth thousand, Liverpool: The Liverpool Booksellers’ Co., 1910. The text was made from the digitised copy in the library of the University of Crete: 274 page scans, of which printed pages 1–255 hold the gospels. One leaf (pages 224–225, John 8:51–9:21) is missing from that copy.

The book is hard for standard OCR. It is set in an unaccented unicase Greek type with a lunate sigma (ϲ), the pages curve towards the spine, and the language is Pallis’s demotic with its own spelling. Tesseract’s stock Greek model got about one character in five wrong.

Steps

  1. Page images. Each page was extracted from the PDF at 300 dpi, leaving out the library seal, which sits on the scans as a separate layer.
  2. Clean-up. The paper tone was flattened, the page deskewed, enlarged to 450 dpi and binarised (Sauvola method).
  3. Straightening and lines. The curl near the spine bends lines by up to a full line height. Each page was cut into narrow vertical strips, the shift between neighbouring strips measured, and the page remapped so that the lines run straight. It was then split into 8,123 single-line images.
  4. Ground truth. In a browser-based review tool, lines were corrected by hand against their images, following fixed conventions (below). 946 lines were corrected this way.
  5. Training. Tesseract 5’s best Greek model (tessdata_best ell) was fine-tuned on the corrected lines. This ran in four rounds: after each round the remaining lines were read again with the new model, which made the next corrections faster. The final model, v4, was trained on 754 lines.
  6. Testing. Six whole pages (192 lines) were held back from training and used only to measure accuracy.
  7. Encoding. The text was exported as TEI P5 XML, keeping every printed page and line, the running heads, paragraphs (from the first-line indent) and the gospel divisions. Each line records whether it was checked by hand or comes straight from the model.
  8. Verse numbers. Pallis has only his own section numbers and chapter numbers. To add verse numbers, his text was aligned with the verses of Codex Vaticanus (GA 03), the manuscript he translated, using the INTF/NTVMR transcription. A dynamic-programming alignment combines word similarity, verse length and punctuation, with each chapter anchored at Pallis’s chapter number. On a hand-made test set (Mark 1–2), 87% of verse starts are placed exactly.

Accuracy

0.32%character errors of the final model on the held-back pages
~22%character errors of Tesseract’s stock Greek model
946lines checked by hand (913 in the gospel text)
3,710verse numbers placed

In practice about three characters in a thousand are wrong in the unchecked lines, most often confusions of similar letters, punctuation and numbers. Faint or damaged pages fare worse. In the reader, View → Mark unchecked OCR lines shows which lines have not been checked.

Transcription conventions

The text follows the print and is not normalised. The print has no accents or breathings, so none are added; a printed diaeresis (ϊ, ϋ) is kept. The lunate sigma is written σ inside a word and ς at its end. Letter case follows letter size in the unicase type, so capitals appear only where the print has taller letters. Quotation marks are “ ”, elision is ’ (κι’ ο), and line-end hyphens are kept in the line view.

Tools

Python (OpenCV, scikit-image, SciPy, lxml, RapidFuzz), Tesseract 5 with tesstrain, and a small review web app written for the project. The work was done with the help of Claude, an AI model by Anthropic.

The GA 03 transcription is by the Institut für Neutestamentliche Textforschung, Münster (NTVMR), licensed CC BY 4.0. Pallis’s translation (1910) is in the public domain.