Browser
Odia datasets.
Live from Hugging Face — every dataset tagged for Odia. Parallel corpora, speech, classification, instruction-tuning, each with its size, license, and a citation.
★ Featuredrotates every Monday
Text Generation
oscar
@oscar-corpus
The Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.
524 208
Text Generation
cc100
@statmt
This corpus is an attempt to recreate the dataset used for training XLM-R. This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages (indicated by rom).
2.0k 107
Also worth a look
Other
Global-MMLU-Lite
@CohereLabs
Releases: Version 3. 0 (May 2026): GMMLU Lite 3.
10K < n < 100K rows 41.1k 43
Apache-2.0textCite
Showing 30 of 466 datasets