Published June 19, 2020 | Version v1
Dataset Open

SemMdf - Semantic Database for Moksha

Description

This SQLite database contains Moksha lemmas and their frequencies in a big corpus. The lemmas are linked to each other based on the syntactic relations they have had in the corpus. Also, the frequency of a syntactic relation between two words is recorded. This means that it is possible to see how frequently for example the word for a dog has appeared with a subject relation with the verb for bark. These database is translated from SemFi by using Giellatekno XML dictionaries. For a detailed description of the structure, see https://www.kaggle.com/mikahama/semfi-finnish-semantics-with-syntactic-relations An easy programmatic interface is provided in UralicNLP: https://github.com/mikahama/uralicNLP/wiki/Semantics-(SemFi,-SemUr) Cite as Hämäläinen, Mika. (2018). Extracting a Semantic Database with Syntactic Relations for Finnish to Boost Resources for Endangered Uralic Languages. In The Proceedings of Logic and Engineering of Natural Language Semantics 15 (LENLS15)

Files

Files (400.1 MB)

Name Size Download all
Checksum: md5:500b8539a17e438f94d4a36697f8586c

PID: http://hdl.handle.net/11304/44140733-0444-4891-aefc-9d2ba680342b
400.1 MB Download

Additional details

Identifiers

B2SHARE Legacy Record ID
98682546da13404a841d7bb7278e63a3

CLARIN metadata

Language Code
mdf
Resource Type
Other