"Suuline eesti keel arvudes" [1, 2] on projekt, mille eesmärk on pakkuda suulise eesti keele kohta baasstatistikat (keeleliste üksuste sagedusi ja pikkuseid), mida seni on olnud saadaval ainult kirjakeele kohta. Andmeallikatena kasutatakse olemasolevaid aegjoondusega suulise eesti keele korpuseid ning projekti käigus loodud suuremahulist automaatselt transkribeeritud ja morfoloogiliselt annoteeritud raadio ja taskuhäälingu korpust. Käesolev kollektsioon koondab loodud korpuste tekste ning kõigist kasutatud korpustest tuletatud sagedusandmestikke.

Projektimeeskond: Pärtel Lippus (vastutav täitja, Tartu Ülikool), Tanel Alumäe (Tallinna Tehnikaülikool), Liina Lindström (Tartu Ülikool), Kaidi Lõo (Tartu Ülikool), Anton Malmi (Tartu Ülikool), Maarja-Liisa Pilvik (Tartu Ülikool), Siim Orasmaa (Tartu Ülikool), Aleksei Kelli (Tartu Ülikool)

The aim of the project "Basic statistics of spoken Estonian" is to provide frequency lists and other basic statistics of spoken Estonian that have only been available based on written language. The data comes from manually annotated time-alligned speech corpora. In addition to analysing systematically collected and manually annotated speech corpora the project team has collected a larger speech corpus of radio talk shows and podcasts that have been automatically annotated using ASR and NLP tools. The speech corpora are used for creating lists of phoneme, morpheme, word, n-gram, and collocation frequencies. This collection consists of the text transcriptions of the speech corpora collected within the project and the frequency data sets.

Project team: Pärtel Lippus (PI, University of Tartu), Tanel Alumäe (TalTech), Liina Lindström (University of Tartu), Kaidi Lõo (University of Tartu), Anton Malmi (University of Tartu), Maarja-Liisa Pilvik (University of Tartu), Siim Orasmaa (University of Tartu), Aleksei Kelli (University of Tartu)

1] "Riiklik programm: Eesti keel ja kultuur digiajastul (EKKD)" projekt EKKD93 "Suuline eesti keel arvudes"(1.01.2022−31.12.2022)

[2] "Riiklik programm: Eesti keel ja kultuur digiajastul (EKKD)" projekt EKKD117 "Suuline eesti keel arvudes II" (1.01.2023−31.12.2023)

Featured Dataverses

In order to use this feature you must have at least one published or linked dataverse.

Publish Dataverse

Are you sure you want to publish your dataverse? Once you do so it must remain published.

Publish Dataverse

This dataverse cannot be published because the dataverse it is in has not been published.

Delete Dataverse

Are you sure you want to delete your dataverse? You cannot undelete this dataverse.

Advanced Search

11 to 20 of 67 Results
Comma Separated Values - 64.0 MB - MD5: c994d352db036f0e9ee33f0896e615c7
wordform frequency
Comma Separated Values - 443.3 MB - MD5: c4dab5d770b6b7e5a5cca1ddbf3efd06
wordform trigrams
Comma Separated Values - 274.9 MB - MD5: c52832c20b9fb03ed79fd36e06c3eea1
wordform trigrams
Comma Separated Values - 625.2 KB - MD5: fb4dd1573ac6bbf12d37068b7ed03d2c
phoneme bigrams
Comma Separated Values - 4.4 MB - MD5: ae29df126a5e58f7c091f3f195497e88
phoneme trigrams
Comma Separated Values - 5.5 MB - MD5: d5c6160cb81c0e4108ce04ee21c0ddd1
phoneme trigrams
Comma Separated Values - 208.9 KB - MD5: e5b90a1c157c4f2447e70c3f9ce09487
phoneme frequency
Comma Separated Values - 710 B - MD5: e796b820eba1f282175f65118c3e27fc
phoneme frequency
Comma Separated Values - 82.3 KB - MD5: 25844fed5a104f85f23460a10f0d7809
consonant frequency
Comma Separated Values - 1.5 MB - MD5: 5ededf39191d6a8a2c8b5c7512de81bd
lemma bigrams
Add Data

Sign up or log in to create a dataverse or add a dataset.

Share Dataverse

Share this dataverse on your favorite social media networks.

Link Dataverse
Reset Modifications

Are you sure you want to reset the selected metadata fields? If you do this, any customizations (hidden, required, optional) you have done will no longer appear.