Datasets
Dataset
ccnews
by kareenamehta
This dataset is the result of processing all WARC files in the CCNews Corpus, from the beginning (2016) to June of 2024. The data has been cleaned and deduplicated, and language of articles have been detected and added.
No licensetext-classification 1200 1100M < n < 1B rows
published 18 Mar 2026
Preview
Parquet error: Scan size limit exceeded: attempted to read 1977184565 bytes, limit is 300000000 bytes Make sure that 1. the Parquet files contain a page index to enable random access without loading entire row groups2. otherwise use smaller row-group sizes when serializing the Parquet files
Cite this
No licenseimagetextCite
Tags
task_categories:text-classificationtask_categories:question-answeringtask_categories:text-generationlanguage:multilinguallanguage:aflanguage:amlanguage:arlanguage:aslanguage:azlanguage:belanguage:bglanguage:bnlanguage:brlanguage:bslanguage:calanguage:cslanguage:cylanguage:dalanguage:delanguage:ellanguage:enlanguage:eolanguage:eslanguage:etlanguage:eulanguage:falanguage:filanguage:frlanguage:fylanguage:galanguage:gdlanguage:gllanguage:gulanguage:halanguage:helanguage:hilanguage:hrlanguage:hulanguage:hylanguage:id