Datasets
Dataset
COILD-MT-Corpus
by coild-dataset
COILD-MT-Corpus is a collection of sentence-aligned parallel text corpora across 20 language pairs, covering Hindi- and Tamil-centric pairings with other Indian languages. Each pair is stored as two line-aligned plain-text files (one sentence per line, same…
CC-BY-NC-4.0translation 3 01M < n < 10M rows
published 10 Sept 2026
Preview
The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.
Cite this
CC-BY-NC-4.0Cite
Tags
task_categories:translationlanguage:hilanguage:aslanguage:bnlanguage:brxlanguage:doilanguage:gomlanguage:gulanguage:kslanguage:mailanguage:mrlanguage:mnilanguage:nelanguage:orlanguage:palanguage:satlanguage:sdlanguage:telanguage:urlanguage:talanguage:knlanguage:mllicense:cc-by-nc-4.0size_categories:1M<n<10Mregion:usmachine-translationparallel-corpusindic-languageslow-resourcemultilingual