Datasets
Dataset
samanantar
by Racrobot1927
Dataset Card for Samanantar Dataset Summary Samanantar is the largest publicly available parallel corpora collection for Indic language: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu. The corpus has 49.
CC-BY-NC-4.0text-generation 16 010M < n < 100M rows
published 26 Mar 2026
Preview
The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.
Cite this
CC-BY-NC-4.0textCite
Tags
task_categories:text-generationtask_categories:translationannotations_creators:no-annotationlanguage_creators:foundmultilinguality:translationsource_datasets:originallanguage:enlanguage:aslanguage:bnlanguage:gulanguage:hilanguage:knlanguage:mllanguage:mrlanguage:orlanguage:palanguage:talanguage:telicense:cc-by-nc-4.0size_categories:10M<n<100Mformat:parquetmodality:textlibrary:datasetslibrary:dasklibrary:polarslibrary:mlcroissantarxiv:2104.05596region:usconditional-text-generation