Datasets
Dataset
samanantar
by ai4bharat
Dataset Card for Samanantar Dataset Summary Samanantar is the largest publicly available parallel corpora collection for Indic language: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu. The corpus has 49.
CC-BY-NC-4.0text-generation 2513 4410M < n < 100M rows
published 2 Mar 2022
Preview
First 5 rows of or / train, from the Hugging Face dataset viewer.
| idx | src | tgt |
|---|---|---|
| 0 | The Congress, however, has also not announced its candidates so far. | ଅଥଚ ବଡ଼ଚଣାର କଂଗ୍ରେସ ପ୍ରାର୍ଥୀ ଆଜି ପର୍ଯ୍ୟନ୍ତ ଘୋଷଣା କରାଯାଇପାରି ନାହିଁ। |
| 1 | Modi said, hitting out at Naidu. | ମୋଦିଙ୍କୁ ଆକ୍ଷେପ କରି ଗଡକରି ଏହା କହିଛନ୍ତି । |
| 2 | The government cannot waive it off. | ସରକାର ଚାହିଲେ ବି ଏହାକୁ ଏଡ଼ାଇ ପାରିବେ ନାହିଁ। |
| 3 | Then add two cups of water. | ତା’ପରେ ସେଥିରେ ଦୁଇ କପ୍ ଗରମ ପାଣି ଢାଳନ୍ତୁ। |
| 4 | Tension prevailed as locals and family members demanded compensation to the next of kin of the deceased. | ମୃତକଙ୍କ ପରିବାରକୁ ସେସୁ ଓ ସମ୍ପୃକ୍ତ ଠିକା ସଂସ୍ଥା ପକ୍ଷରୁ କ୍ଷତିପୂରଣ ଦେବା ଦାବିରେ ସ୍ଥାନୀୟ ଲୋକେ ଘଟଣାସ୍ଥଳରେ ଆନ୍ଦୋଳନ କରିବାରୁ ଉତ୍ତେଜନା ପ୍ରକାଶ ପାଇଥିଲା। |
Cite this
CC-BY-NC-4.0textCite
Tags
task_categories:text-generationtask_categories:translationannotations_creators:no-annotationlanguage_creators:foundmultilinguality:translationsource_datasets:originallanguage:enlanguage:aslanguage:bnlanguage:gulanguage:hilanguage:knlanguage:mllanguage:mrlanguage:orlanguage:palanguage:talanguage:telicense:cc-by-nc-4.0size_categories:10M<n<100Mformat:parquetmodality:textlibrary:datasetslibrary:dasklibrary:mlcroissantlibrary:polarsarxiv:2104.05596region:usconditional-text-generation