OLAC Record oai:www.ldc.upenn.edu:LDC2024T08 |
Metadata | ||
Title: | RST Continuity Corpus | |
Access Rights: | Licensing Instructions for Subscription & Standard Members, and Non-Members: http://www.ldc.upenn.edu/language-resources/data/obtaining | |
Bibliographic Citation: | Das, Debopam, and Markus Egg. RST Continuity Corpus LDC2024T08. Web Download. Philadelphia: Linguistic Data Consortium, 2024 | |
Contributor: | Das, Debopam | |
Egg, Markus | ||
Date (W3CDTF): | 2024 | |
Date Issued (W3CDTF): | 2024-10-15 | |
Description: | *Introduction* RST Continuity Corpus was developed at Åbo Akademi University and Humboldt-Universität zu Berlin and contains annotations for continuity dimensions added to RST Discourse Treebank (LDC2002T07). RST Discourse Treebank is a collection of English news texts from the Penn Treebank annotated for rhetorical relations under the RST (Rhetorical Structure Theory) framework. In RST Continuity Corpus, the relations are annotated for the seven continuity dimensions: time, space, reference, action, perspective, modality, and speech act. The relations are also annotated for polarity, order of segments, nuclearity, and context. *Data* The source data consists of 1,009 relations from 217 Wall Street Journal texts annotated in RST Discourse Treebank for five relation types: causal, contrastive, conditional, elaboration and temporal. Annotation was performed using the UAM CorpusTool, version 2.8.16 or later. Files are presented as UTF-8 encoded XML and plain text. The corpus is divided into four sub-directories as described in the README file. *Samples* Please view the following samples: * Continuity Sample * Additional Parameters Sample *Updates* None at this time. | |
Extent: | Corpus size: 16245 KB | |
Identifier: | LDC2024T08 | |
https://catalog.ldc.upenn.edu/LDC2024T08 | ||
ISLRN: 183-361-437-399-8 | ||
DOI: 10.35111/jfbf-gn90 | ||
Language: | English | |
Language (ISO639): | eng | |
License: | LDC User Agreement for Non-Members: https://catalog.ldc.upenn.edu/license/ldc-non-members-agreement.pdf | |
Medium: | Distribution: Web Download | |
Publisher: | Linguistic Data Consortium | |
Publisher (URI): | https://www.ldc.upenn.edu | |
Relation (URI): | https://catalog.ldc.upenn.edu/docs/LDC2024T08 | |
Rights Holder: | Portions © 1987-1989 Dow Jones & Company, Inc., © 2024 Depobam Das, © 2024 Markus Egg, © 1995, 1999, 2002, 2015, 2024 Trustees of the University of Pennsylvania | |
Type (DCMI): | Text | |
Type (OLAC): | primary_text | |
OLAC Info |
||
Archive: | The LDC Corpus Catalog | |
Description: | http://www.language-archives.org/archive/www.ldc.upenn.edu | |
GetRecord: | OAI-PMH request for OLAC format | |
GetRecord: | Pre-generated XML file | |
OAI Info |
||
OaiIdentifier: | oai:www.ldc.upenn.edu:LDC2024T08 | |
DateStamp: | 2024-10-15 | |
GetRecord: | OAI-PMH request for simple DC format | |
Search Info | ||
Citation: | Das, Debopam; Egg, Markus. 2024. Linguistic Data Consortium. | |
Terms: | area_Europe country_GB dcmi_Text iso639_eng olac_primary_text |