OLAC Record

Title:Training corpus ssj500k 2.2
Bibliographic Citation:http://hdl.handle.net/11356/1210
Creator:Krek, Simon
Dobrovoljc, Kaja
Erjavec, Tomaž
Može, Sara
Ledinek, Nina
Holz, Nanika
Zupan, Katja
Gantar, Polona
Kuzman, Taja
Čibej, Jaka
Arhar Holdt, Špela
Kavčič, Teja
Škrjanec, Iza
Marko, Dafne
Jezeršek, Lucija
Zajc, Anja
Date (W3CDTF):2019-01-26T20:37:28Z
Date Available:2019-01-26T20:37:28Z
Description:The ssj500k training corpus contains about 500,000 tokens manually annotated on the levels of tokenisation, sentence segmentation, morphosyntactic tagging, and lemmatisation. About half of the corpus is also manually annotated with syntactic dependencies, named entities, and verbal multiword expressions. About a quarter of the corpus is annotated with semantic role labels. The morphosyntactic tags and syntactic dependencies are included both in the JOS/MULTEXT-East framework, as well as in the framework of Universal Dependencies. The annotations of the ssj500k corpus follow (1) the MULTEXT-East V6 morphosyntactic specifications for Slovene, http://nl.ijs.si/ME/V6/msd/, (2) the JOS dependency schema, http://nl.ijs.si/jos/bib/jos-skladnja-navodila.pdf, the Universal Dependencies morphosyntactic specifications and syntactic dependencies for Slovene-SSJ, https://universaldependencies.org/, (4) the Janes annotation guidelines for Slovenian named entities, http://nl.ijs.si/janes/wp-content/uploads/2017/09/SlovenianNER-eng-v1.1.pdf, and (5) the Guidelines of the PARSEME shared task on verbal multiword expressions, http://parsemefr.lif.univ-mrs.fr/parseme-st-guidelines/1.1/ The vocabulary of (1) and (2) is provided in the back element and (3), (4), and (5) in the teiHeader of the TEI encoded corpus. The semantic role labels are also documented in the teiHeader. In contrast to the previous version 2.1, this version corrects various errata in spacing and text metadata and adds UD morphological and (where it was possible to do so automatically) dependency annotations to the corpus. Note that the UD annotations are not included in the vertical file.
Identifier (URI):http://hdl.handle.net/11356/1210
Language (ISO639):slv
Publisher:Centre for Language Resources and Technologies, University of Ljubljana
Replaces (URI):http://hdl.handle.net/11356/1181
Rights:Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)
Subject:part-of-speech tagging
dependency treebank
named entities
manual annotation
verbal multiword expressions
semantic role labelling
Type (DCMI):Text
Type (OLAC):primary_text


Archive:  Slovenian language resource repository CLARIN.SI
Description:  http://www.language-archives.org/archive/clarin.si
GetRecord:  OAI-PMH request for OLAC format
GetRecord:  Pre-generated XML file

OAI Info

OaiIdentifier:  oai:www.clarin.si:11356/1210
DateStamp:  2019-10-10
GetRecord:  OAI-PMH request for simple DC format

Search Info

Citation: Krek, Simon; Dobrovoljc, Kaja; Erjavec, Tomaž; Može, Sara; Ledinek, Nina; Holz, Nanika; Zupan, Katja; Gantar, Polona; Kuzman, Taja; Čibej, Jaka; Arhar Holdt, Špela; Kavčič, Teja; Škrjanec, Iza; Marko, Dafne; Jezeršek, Lucija; Zajc, Anja. 2019. Centre for Language Resources and Technologies, University of Ljubljana.
Terms: area_Europe country_SI dcmi_Text iso639_slv olac_primary_text

Up-to-date as of: Fri Jan 10 9:22:54 EST 2020