hatchmoment. scored by care · not by stars

ru-petr1708-OCR

OCR tool for Russian pre‑1918 orthography using Tesseract

The project supplies a custom Tesseract traineddata model and helper scripts to correctly recognize Russian books printed between 1708 and 1918, handling the unique orthography of that era. It includes Python utilities for image normalization and OCR invocation, as well as tools to build auxiliary language models. Designed for historians, linguists, and anyone digitizing old Russian literature, it outperforms generic Russian OCR models on this niche corpus. By bundling the model and preprocessing steps, it offers a ready‑to‑use solution without needing extensive OCR tuning.

ocrru-petr1708russianrussian-languagerussian-ocrtesseracttesseract-ocr
View on GitHub →

slavenica/ru-petr1708-OCR