Datasets

Dataset

ipa-dict

by omneity-labs

IPA Dict Dataset This dataset is a comprehensive collection of word pronunciations in the International Phonetic Alphabet (IPA) for multiple languages. It is based on the open-source ipa-dict repository.

MITother 79 21M < n < 10M rows

published 25 Mar 2026

Preview

First 5 rows of default / train, from the Hugging Face dataset viewer.

textipalang
آئلaːʔilar
آبaːbar
آبaːbaar
آباءaːbaːʔar
آبادaːbaːdar

Cite this

MITtextCite

Tags

language:arlanguage:delanguage:enlanguage:eolanguage:eslanguage:falanguage:filanguage:frlanguage:islanguage:jalanguage:jamlanguage:kmlanguage:kolanguage:malanguage:nblanguage:nllanguage:orlanguage:ptlanguage:rolanguage:svlanguage:swlanguage:ttslanguage:vilanguage:yuelanguage:zhlicense:mitsize_categories:1M<n<10Mformat:parquetmodality:textlibrary:datasetslibrary:pandaslibrary:polarslibrary:mlcroissantregion:uslinguisticspronunciationipa