Datasets
Dataset
BhartiOCR
by Faizaniqbal
BhartiOCR A Multilingual Synthetic OCR and Document Benchmark for Indian Languages 13. 25M image-text pairs · 23 languages · 12 writing systems · 2,650 WebDataset shards BhartiOCR is a large-scale synthetic dataset and benchmark designed for research in mul…
Apache-2.0image-to-text 1370 010M < n < 100M rows
published 9 Sept 2026
Preview
The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.
Cite this
Apache-2.0imageCite
Tags
task_categories:image-to-textlanguage:hilanguage:urlanguage:bnlanguage:talanguage:mrlanguage:telanguage:gulanguage:knlanguage:mllanguage:orlanguage:palanguage:aslanguage:nelanguage:salanguage:satlanguage:mnilanguage:sdlanguage:brxlanguage:bholanguage:kslanguage:gomlanguage:mailanguage:doilicense:apache-2.0size_categories:10M<n<100Mmodality:imagelibrary:webdatasetregion:usocrmultilingual-ocrdocument-aisynthetic-datacomputer-visionindian-languagespan-indicwebdatasetdocument-understandingimage-textoptical-character-recognition