Datasets

Dataset

samanantar

by Racrobot1927

Dataset Card for Samanantar Dataset Summary Samanantar is the largest publicly available parallel corpora collection for Indic language: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu. The corpus has 49.

CC-BY-NC-4.0text-generation 16 010M < n < 100M rows

published 26 Mar 2026

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

CC-BY-NC-4.0textCite

Tags

task_categories:text-generationtask_categories:translationannotations_creators:no-annotationlanguage_creators:foundmultilinguality:translationsource_datasets:originallanguage:enlanguage:aslanguage:bnlanguage:gulanguage:hilanguage:knlanguage:mllanguage:mrlanguage:orlanguage:palanguage:talanguage:telicense:cc-by-nc-4.0size_categories:10M<n<100Mformat:parquetmodality:textlibrary:datasetslibrary:dasklibrary:polarslibrary:mlcroissantarxiv:2104.05596region:usconditional-text-generation