Datasets

Dataset

nayanabench-rendered

by ranjanhr1

NayanaBench Rendered Dataset Dataset Description This dataset contains 2D rendered text images for document analysis and OCR tasks across 22 languages. Each language split contains document images with text rendered in bounding boxes using authentic fonts.

No licenseimage-to-text 20 01K < n < 10K rows

published 29 Nov 2025

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

No licenseimagetextCite

Tags

task_categories:image-to-texttask_categories:object-detectionmultilinguality:multilinguallanguage:enlanguage:hilanguage:knlanguage:talanguage:telanguage:mrlanguage:palanguage:bnlanguage:orlanguage:mllanguage:gulanguage:salanguage:jalanguage:kolanguage:zhlanguage:delanguage:frlanguage:itlanguage:rulanguage:arlanguage:eslanguage:thsize_categories:1K<n<10Kformat:parquetmodality:imagemodality:textlibrary:datasetslibrary:pandaslibrary:mlcroissantlibrary:polarsregion:usocrtext-renderingmultilingualdocument-analysis