Datasets

Dataset

COILD-MT-Corpus

by coild-dataset

COILD-MT-Corpus is a collection of sentence-aligned parallel text corpora across 20 language pairs, covering Hindi- and Tamil-centric pairings with other Indian languages. Each pair is stored as two line-aligned plain-text files (one sentence per line, same…

CC-BY-NC-4.0translation 3 01M < n < 10M rows

published 10 Sept 2026

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

CC-BY-NC-4.0Cite

Tags

task_categories:translationlanguage:hilanguage:aslanguage:bnlanguage:brxlanguage:doilanguage:gomlanguage:gulanguage:kslanguage:mailanguage:mrlanguage:mnilanguage:nelanguage:orlanguage:palanguage:satlanguage:sdlanguage:telanguage:urlanguage:talanguage:knlanguage:mllicense:cc-by-nc-4.0size_categories:1M<n<10Mregion:usmachine-translationparallel-corpusindic-languageslow-resourcemultilingual