Datasets

Dataset

ccnews

by kareenamehta

This dataset is the result of processing all WARC files in the CCNews Corpus, from the beginning (2016) to June of 2024. The data has been cleaned and deduplicated, and language of articles have been detected and added.

No licensetext-classification 1200 1100M < n < 1B rows

published 18 Mar 2026

Preview

Parquet error: Scan size limit exceeded: attempted to read 1977184565 bytes, limit is 300000000 bytes Make sure that 1. the Parquet files contain a page index to enable random access without loading entire row groups2. otherwise use smaller row-group sizes when serializing the Parquet files

Cite this

No licenseimagetextCite

Tags

task_categories:text-classificationtask_categories:question-answeringtask_categories:text-generationlanguage:multilinguallanguage:aflanguage:amlanguage:arlanguage:aslanguage:azlanguage:belanguage:bglanguage:bnlanguage:brlanguage:bslanguage:calanguage:cslanguage:cylanguage:dalanguage:delanguage:ellanguage:enlanguage:eolanguage:eslanguage:etlanguage:eulanguage:falanguage:filanguage:frlanguage:fylanguage:galanguage:gdlanguage:gllanguage:gulanguage:halanguage:helanguage:hilanguage:hrlanguage:hulanguage:hylanguage:id