Datasets
Dataset
nayanabench-rendered
by ranjanhr1
NayanaBench Rendered Dataset Dataset Description This dataset contains 2D rendered text images for document analysis and OCR tasks across 22 languages. Each language split contains document images with text rendered in bounding boxes using authentic fonts.
No licenseimage-to-text 20 01K < n < 10K rows
published 29 Nov 2025
Preview
The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.
Cite this
No licenseimagetextCite
Tags
task_categories:image-to-texttask_categories:object-detectionmultilinguality:multilinguallanguage:enlanguage:hilanguage:knlanguage:talanguage:telanguage:mrlanguage:palanguage:bnlanguage:orlanguage:mllanguage:gulanguage:salanguage:jalanguage:kolanguage:zhlanguage:delanguage:frlanguage:itlanguage:rulanguage:arlanguage:eslanguage:thsize_categories:1K<n<10Kformat:parquetmodality:imagemodality:textlibrary:datasetslibrary:pandaslibrary:mlcroissantlibrary:polarsregion:usocrtext-renderingmultilingualdocument-analysis