Datasets

Dataset

BharatSetu-22Indic-1M

by Omarrran

BharatSetu 22-Indic 1M BharatSetu 22-Indic 1M is a private wide parallel translation dataset containing 984,764 English source rows translated into 22 Indic target languages. The Hugging Face Dataset Viewer is configured to read the Parquet shards under dat…

No licensetranslation 1 0100K < n < 1M rows

published 18 Sept 2026

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

No licensetextCite

Tags

task_categories:translationlanguage:enlanguage:aslanguage:bnlanguage:brxlanguage:doilanguage:gulanguage:hilanguage:knlanguage:kslanguage:koklanguage:mailanguage:mllanguage:mnilanguage:mrlanguage:nelanguage:orlanguage:palanguage:salanguage:satlanguage:sdlanguage:talanguage:telanguage:ursize_categories:100K<n<1Mformat:parquetmodality:textlibrary:datasetslibrary:dasklibrary:polarslibrary:mlcroissantregion:ustranslationmultilingualindic-languagesenglish-to-indicsarvam-translate