Datasets

Dataset

BhartiOCR

by Faizaniqbal

BhartiOCR A Multilingual Synthetic OCR and Document Benchmark for Indian Languages 13. 25M image-text pairs · 23 languages · 12 writing systems · 2,650 WebDataset shards BhartiOCR is a large-scale synthetic dataset and benchmark designed for research in mul…

Apache-2.0image-to-text 1370 010M < n < 100M rows

published 9 Sept 2026

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

Apache-2.0imageCite

Tags

task_categories:image-to-textlanguage:hilanguage:urlanguage:bnlanguage:talanguage:mrlanguage:telanguage:gulanguage:knlanguage:mllanguage:orlanguage:palanguage:aslanguage:nelanguage:salanguage:satlanguage:mnilanguage:sdlanguage:brxlanguage:bholanguage:kslanguage:gomlanguage:mailanguage:doilicense:apache-2.0size_categories:10M<n<100Mmodality:imagelibrary:webdatasetregion:usocrmultilingual-ocrdocument-aisynthetic-datacomputer-visionindian-languagespan-indicwebdatasetdocument-understandingimage-textoptical-character-recognition