Datasets

Dataset

ccmatrix

by xezpeleta

CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WEB We show that margin-based bitext mining in LASER's multilingual sentence space can be applied to monolingual corpora of billions of sentences to produce high quality aligned translation…

No licensetranslation 89 0100M < n < 1B rows

published 19 Feb 2024

Preview

The dataset viewer doesn't support this dataset because it runs arbitrary python code. Please open a discussion in the discussion tab if you think this is an error and tag @lhoestq and @severo.

Cite this

No licenseCite

Tags

task_categories:translationannotations_creators:foundlanguage_creators:foundmultilinguality:multilingualsource_datasets:originallanguage:aflanguage:amlanguage:arlanguage:astlanguage:azlanguage:belanguage:bglanguage:bnlanguage:brlanguage:calanguage:ceblanguage:cslanguage:cylanguage:dalanguage:delanguage:ellanguage:enlanguage:eolanguage:eslanguage:etlanguage:eulanguage:falanguage:filanguage:frlanguage:fylanguage:galanguage:gdlanguage:gllanguage:halanguage:helanguage:hilanguage:hrlanguage:hulanguage:hylanguage:id