Datasets
Dataset
NayanaOCRBench_Synthetic
by Cognitive-Lab
NayanaOCRBench · Synthetic Held-out evaluation split of the NayanaOCR 2026 pipeline, 22 languages, ~770 pages per language. Every page is the same source document re-typeset in each language with the layout preserved, so the set is parallel across languages…
CC-BY-NC-4.0image-to-text 0 010K < n < 100K rows
published 17 Sept 2026
Preview
The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.
Cite this
CC-BY-NC-4.0imagetextCite
Tags
task_categories:image-to-texttask_categories:image-text-to-textlanguage:enlanguage:hilanguage:bnlanguage:mrlanguage:gulanguage:palanguage:orlanguage:knlanguage:talanguage:telanguage:mllanguage:salanguage:zhlanguage:jalanguage:kolanguage:delanguage:frlanguage:eslanguage:itlanguage:rulanguage:arlanguage:thlicense:cc-by-nc-4.0size_categories:10K<n<100Kformat:parquetmodality:imagemodality:textlibrary:datasetslibrary:pandaslibrary:polarslibrary:mlcroissantregion:usocrdocument-aimultilingualindicbenchmarknayana