Datasets
Dataset
ccmatrix
by xezpeleta
CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WEB We show that margin-based bitext mining in LASER's multilingual sentence space can be applied to monolingual corpora of billions of sentences to produce high quality aligned translation…
No licensetranslation 89 0100M < n < 1B rows
published 19 Feb 2024
Preview
The dataset viewer doesn't support this dataset because it runs arbitrary python code. Please open a discussion in the discussion tab if you think this is an error and tag @lhoestq and @severo.
Cite this
No licenseCite
Tags
task_categories:translationannotations_creators:foundlanguage_creators:foundmultilinguality:multilingualsource_datasets:originallanguage:aflanguage:amlanguage:arlanguage:astlanguage:azlanguage:belanguage:bglanguage:bnlanguage:brlanguage:calanguage:ceblanguage:cslanguage:cylanguage:dalanguage:delanguage:ellanguage:enlanguage:eolanguage:eslanguage:etlanguage:eulanguage:falanguage:filanguage:frlanguage:fylanguage:galanguage:gdlanguage:gllanguage:halanguage:helanguage:hilanguage:hrlanguage:hulanguage:hylanguage:id