Title:Training corpus ssj500k 2.0
Bibliographic Citation:http://hdl.handle.net/11356/1165
Creator:Krek, Simon
Dobrovoljc, Kaja
Erjavec, Tomaž
Može, Sara
Ledinek, Nina
Holz, Nanika
Zupan, Katja
Gantar, Polona
Kuzman, Taja
Date (W3CDTF):2017-11-23T21:36:56Z
Date Available:2017-11-23T21:36:56Z
Description:The ssj500k training corpus contains about 500,000 tokens manually annotated on the levels of tokenisation, sentence segmentation, morphosyntactic tagging, and lemmatisation. About half of the corpus is also manually annotated with syntactic dependencies, named entities, and verbal multiword expressions. The annotations of the ssj500k corpus follow (1) the MULTEXT-East V5 morphosyntactic specifications for Slovene, http://nl.ijs.si/ME/V5/msd/, (2) the JOS dependency schema, http://nl.ijs.si/jos/bib/jos-skladnja-navodila.pdf, (3) the Janes Annotation guidelines for Slovenian named entities, http://nl.ijs.si/janes/wp-content/uploads/2017/09/SlovenianNER-eng-v1.1.pdf, and the Guidelines of the PARSEME shared task on verbal multiword expressions, http://parsemefr.lif.univ-mrs.fr/parseme-st-guidelines/1.0/ The vocabulary of (1) and (2) is provided in the back element and (3) and (4) in the teiHeader of the TEI encoded corpus.
Identifier (URI):http://hdl.handle.net/11356/1165
Is Replaced By (URI):http://hdl.handle.net/11356/1181
Language (ISO639):slv
Publisher:Centre for Language Resources and Technologies, University of Ljubljana
Replaces (URI):http://hdl.handle.net/11356/1052
Rights:Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)
dependency treebank
named entities
manual annotation
verbal multiword expressions
Type (DCMI):Text
Type (OLAC):primary_text


Archive:  Slovenian language resource repository CLARIN.SI
Description:  http://www.language-archives.org/archive/clarin.si
OaiIdentifier:  oai:www.clarin.si:11356/1165
DateStamp:  2018-05-28
Citation: Krek, Simon; Dobrovoljc, Kaja; Erjavec, Tomaž; Može, Sara; Ledinek, Nina; Holz, Nanika; Zupan, Katja; Gantar, Polona; Kuzman, Taja. 2017. Centre for Language Resources and Technologies, University of Ljubljana.
Terms: area_Europe country_SI dcmi_Text iso639_slv olac_primary_text

