Datasets

Dataset

roots_indic-or_indic_nlp_corpus

by bigscience-data

ROOTS Subset: rootsindic-orindicnlpcorpus Indic NLP Corpus Dataset uid: indicnlpcorpus Description The IndicNLP corpus is a largescale, general-domain corpus containing 2. 7 billion words for 10 Indian languages from two language families.

CC-BY-NC-4.0other 7 01M < n < 10M rows

published 18 May 2022

Preview

The dataset does not exist, or is not accessible without authentication (private or gated). Please check the spelling of the dataset name or retry with authentication.

Cite this

CC-BY-NC-4.0textCite

Tags

language:orlicense:cc-by-nc-4.0size_categories:1M<n<10Mformat:parquetmodality:textlibrary:datasetslibrary:dasklibrary:mlcroissantlibrary:polarsregion:us