grabar-open-data

steps-re/grabar-open-data

Bedrossian's Armenian-English dictionary (1875-79) with a page image behind every entry, plus an OCR agreement benchmark over eleven Armenian imprints 1668-1905. CC BY 4.0.

Stars: 0 Forks: 0 License: CC-BY-4.0 Data

Summary

This repository provides three structured datasets focused on Classical Armenian (Grabar) texts. The primary dataset is a machine-digitized version of Bedrossian's Armenian-English dictionary (1875-79) with 34,557 headwords, each linked to its original scanned page image for verification. The second dataset is a list of 4,329 dictionary headwords potentially missing from the popular Calfa Armenian dictionary database, serving as a gap analysis. The third dataset is an OCR agreement benchmark comparing Calfa's open Tesseract model against manual transcriptions across 2,053 pages of historical Armenian imprints (1668-1905). The repository emphasizes data provenance, open licensing (CC BY 4.0), and correction via linked page images.

Similar Projects