![]() |
OLAC Record oai:www.clarin.si:11356/1181 |
Metadata | ||
Title: | Training corpus ssj500k 2.1 | |
Bibliographic Citation: | http://hdl.handle.net/11356/1181 | |
Creator: | Krek, Simon | |
Dobrovoljc, Kaja | ||
Erjavec, Tomaž | ||
Može, Sara | ||
Ledinek, Nina | ||
Holz, Nanika | ||
Zupan, Katja | ||
Gantar, Polona | ||
Kuzman, Taja | ||
Čibej, Jaka | ||
Arhar Holdt, Špela | ||
Kavčič, Teja | ||
Škrjanec, Iza | ||
Marko, Dafne | ||
Jezeršek, Lucija | ||
Zajc, Anja | ||
Date (W3CDTF): | 2018-03-16T17:57:20Z | |
Date Available: | 2018-03-16T17:57:20Z | |
Description: | The ssj500k training corpus contains about 500,000 tokens manually annotated on the levels of tokenisation, sentence segmentation, morphosyntactic tagging, and lemmatisation. About half of the corpus is also manually annotated with syntactic dependencies, named entities, and verbal multiword expressions. About a quarter of the corpus is annotated with semantic role labels. The annotations of the ssj500k corpus follow (1) the MULTEXT-East V5 morphosyntactic specifications for Slovene, http://nl.ijs.si/ME/V5/msd/, (2) the JOS dependency schema, http://nl.ijs.si/jos/bib/jos-skladnja-navodila.pdf, (3) the Janes annotation guidelines for Slovenian named entities, http://nl.ijs.si/janes/wp-content/uploads/2017/09/SlovenianNER-eng-v1.1.pdf, and (4) the Guidelines of the PARSEME shared task on verbal multiword expressions, http://parsemefr.lif.univ-mrs.fr/parseme-st-guidelines/1.1/ The vocabulary of (1) and (2) is provided in the back element and (3) and (4) in the teiHeader of the TEI encoded corpus. The semantic role labels are also documented in the teiHeader. | |
Identifier (URI): | http://hdl.handle.net/11356/1181 | |
Is Replaced By (URI): | http://hdl.handle.net/11356/1210 | |
Language: | Slovenian | |
Language (ISO639): | slv | |
Publisher: | Centre for Language Resources and Technologies, University of Ljubljana | |
Replaces (URI): | http://hdl.handle.net/11356/1165 | |
Rights: | Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) | |
https://creativecommons.org/licenses/by-nc-sa/4.0/ | ||
Subject: | tagging | |
dependency treebank | ||
parsing | ||
named entities | ||
tokenisation | ||
manual annotation | ||
TEI | ||
verbal multiword expressions | ||
semantic role labelling | ||
Type: | corpus | |
Type (DCMI): | Text | |
Type (OLAC): | primary_text | |
OLAC Info |
||
Archive: | Slovenian language resource repository CLARIN.SI | |
Description: | http://www.language-archives.org/archive/clarin.si | |
GetRecord: | OAI-PMH request for OLAC format | |
GetRecord: | Pre-generated XML file | |
OAI Info |
||
OaiIdentifier: | oai:www.clarin.si:11356/1181 | |
DateStamp: | 2019-01-26 | |
GetRecord: | OAI-PMH request for simple DC format | |
Search Info | ||
Citation: | Krek, Simon; Dobrovoljc, Kaja; Erjavec, Tomaž; Može, Sara; Ledinek, Nina; Holz, Nanika; Zupan, Katja; Gantar, Polona; Kuzman, Taja; Čibej, Jaka; Arhar Holdt, Špela; Kavčič, Teja; Škrjanec, Iza; Marko, Dafne; Jezeršek, Lucija; Zajc, Anja. 2018. Centre for Language Resources and Technologies, University of Ljubljana. | |
Terms: | area_Europe country_SI dcmi_Text iso639_slv olac_primary_text |