OLAC Record: Greybeard

OLAC Record
oai:www.ldc.upenn.edu:LDC2013S05

Metadata

Title: Greybeard

Access Rights: Licensing Instructions for Subscription & Standard Members, and Non-Members: http://www.ldc.upenn.edu/language-resources/data/obtaining

Bibliographic Citation: Brandschain, Linda, and David Graff. Greybeard LDC2013S05. Web Download. Philadelphia: Linguistic Data Consortium, 2013

Contributor: Brandschain, Linda

Graff, David

Date (W3CDTF): 2013

Date Issued (W3CDTF): 2013-06-17

Description: *Introduction* Greybeard was developed by the Linguistic Data Consortium (LDC) and is comprised of approximately 590 hours of English telephone conversation speech collected in October and November 2008 by LDC. The goal was to record new telephone conversations among subjects who had participated in one or more previous LDC telephone collections, from Switchboard-1 (1991) through the Mixer studies (2006). A total of 172 subjects were enrolled in the Greybeard collection, all of whom had participated in one of the following: * Switchboard-1 (LDC97S62) 1991-1992: 2 subjects * Switchboard-2 (LDC98S75, LDC99S79, LDC2002S06) 1996-1997: 16 subjects * Mixer 1 and 2 2003-2005: 103 subjects * Mixer 3 2006: 51 subjects Most Greybeard participants completed 12 calls. Some subjects completed up to 24 calls. Calls were made or received via an automatic operator system at LDC which connected two participants and announced a topic for discussion. *Data* This releases consists of 4680 calls -- the complete set of calls recorded during the Greybeard collection (1098 calls) as well as all calls from the legacy collections that involved the Greybeard speakers. The audio from each call was captured digitally by the operator system and stored in a separate file as raw mu-law sample data. As the recordings were uploaded daily from the robot operator to network disk storage, automated processes reformatted the audio into a 2-channel SPHERE-format file for each conversation and queued the recordings for manual audit to verify speaker identification and to check other aspects of the recording. Auditors provided impressionistic judgments on overall audio quality, presence of background noise and cross-channel echo and any other technical difficulty with the call, in addition to confirming the speaker-ID on each channel. These auditor decisions are provided in the call_info tables, described in more detail in the included documentation. For this release, each 2-channel recording was converted from SPHERE to MS-WAV file format and compressed using FLAC. All audio files are 2-channel, 8 KHz, 16-bit PCM sample data, in FLAC-compressed form (http://flac.sourceforge.net). When uncompressed, they have MS-WAV/RIFF headers. *Samples* Please listen to the following audio sample. *Updates* None at this time.

Extent: Corpus size: 20320096 KB

Format: Sampling Rate: 8000

Sampling Format: pcm

Identifier: LDC2013S05

https://catalog.ldc.upenn.edu/LDC2013S05

ISBN: 1-58563-645-2

ISLRN: 854-216-857-102-2

DOI: 10.35111/tahq-9n25

Language: English

Language (ISO639): eng

License: LDC User Agreement for Non-Members: https://catalog.ldc.upenn.edu/license/ldc-non-members-agreement.pdf

Medium: Distribution: Web Download

Publisher: Linguistic Data Consortium

Publisher (URI): https://www.ldc.upenn.edu

Relation (URI): https://catalog.ldc.upenn.edu/docs/LDC2013S05

Rights Holder: Portions © 2008, 2013 Trustees of the University of Pennsylvania

Type (DCMI): Sound

Type (OLAC): primary_text

OLAC Info

Archive: The LDC Corpus Catalog

Description: http://www.language-archives.org/archive/www.ldc.upenn.edu

GetRecord: OAI-PMH request for OLAC format

GetRecord: Pre-generated XML file

OAI Info

OaiIdentifier: oai:www.ldc.upenn.edu:LDC2013S05

DateStamp: 2020-11-30

GetRecord: OAI-PMH request for simple DC format

Search Info
Citation: Brandschain, Linda; Graff, David. 2013. Linguistic Data Consortium.
Terms: area_Europe country_GB dcmi_Sound iso639_eng olac_primary_text

http://www.language-archives.org/item.php/oai:www.ldc.upenn.edu:LDC2013S05
Up-to-date as of: Wed Oct 29 7:01:24 EDT 2025

Metadata
Title:		Greybeard
Access Rights:		Licensing Instructions for Subscription & Standard Members, and Non-Members: http://www.ldc.upenn.edu/language-resources/data/obtaining
Bibliographic Citation:		Brandschain, Linda, and David Graff. Greybeard LDC2013S05. Web Download. Philadelphia: Linguistic Data Consortium, 2013
Contributor:		Brandschain, Linda
Contributor:		Graff, David
Date (W3CDTF):		2013
Date Issued (W3CDTF):		2013-06-17
Description:		Introduction Greybeard was developed by the Linguistic Data Consortium (LDC) and is comprised of approximately 590 hours of English telephone conversation speech collected in October and November 2008 by LDC. The goal was to record new telephone conversations among subjects who had participated in one or more previous LDC telephone collections, from Switchboard-1 (1991) through the Mixer studies (2006). A total of 172 subjects were enrolled in the Greybeard collection, all of whom had participated in one of the following: * Switchboard-1 (LDC97S62) 1991-1992: 2 subjects * Switchboard-2 (LDC98S75, LDC99S79, LDC2002S06) 1996-1997: 16 subjects * Mixer 1 and 2 2003-2005: 103 subjects * Mixer 3 2006: 51 subjects Most Greybeard participants completed 12 calls. Some subjects completed up to 24 calls. Calls were made or received via an automatic operator system at LDC which connected two participants and announced a topic for discussion. Data This releases consists of 4680 calls -- the complete set of calls recorded during the Greybeard collection (1098 calls) as well as all calls from the legacy collections that involved the Greybeard speakers. The audio from each call was captured digitally by the operator system and stored in a separate file as raw mu-law sample data. As the recordings were uploaded daily from the robot operator to network disk storage, automated processes reformatted the audio into a 2-channel SPHERE-format file for each conversation and queued the recordings for manual audit to verify speaker identification and to check other aspects of the recording. Auditors provided impressionistic judgments on overall audio quality, presence of background noise and cross-channel echo and any other technical difficulty with the call, in addition to confirming the speaker-ID on each channel. These auditor decisions are provided in the call_info tables, described in more detail in the included documentation. For this release, each 2-channel recording was converted from SPHERE to MS-WAV file format and compressed using FLAC. All audio files are 2-channel, 8 KHz, 16-bit PCM sample data, in FLAC-compressed form (http://flac.sourceforge.net). When uncompressed, they have MS-WAV/RIFF headers. Samples Please listen to the following audio sample. Updates None at this time.
Extent:		Corpus size: 20320096 KB
Format:		Sampling Rate: 8000
Format:		Sampling Format: pcm
Identifier:		LDC2013S05
		https://catalog.ldc.upenn.edu/LDC2013S05
		ISBN: 1-58563-645-2
		ISLRN: 854-216-857-102-2
		DOI: 10.35111/tahq-9n25
Language:		English
Language (ISO639):		eng
License:		LDC User Agreement for Non-Members: https://catalog.ldc.upenn.edu/license/ldc-non-members-agreement.pdf
Medium:		Distribution: Web Download
Publisher:		Linguistic Data Consortium
Publisher (URI):		https://www.ldc.upenn.edu
Relation (URI):		https://catalog.ldc.upenn.edu/docs/LDC2013S05
Rights Holder:		Portions © 2008, 2013 Trustees of the University of Pennsylvania
Type (DCMI):		Sound
Type (OLAC):		primary_text
OLAC Info
Archive:		The LDC Corpus Catalog
Description:		http://www.language-archives.org/archive/www.ldc.upenn.edu
GetRecord:		OAI-PMH request for OLAC format
GetRecord:		Pre-generated XML file
OAI Info
OaiIdentifier:		oai:www.ldc.upenn.edu:LDC2013S05
DateStamp:		2020-11-30
GetRecord:		OAI-PMH request for simple DC format
Search Info
Citation:		Brandschain, Linda; Graff, David. 2013. Linguistic Data Consortium.
Terms:		area_Europe country_GB dcmi_Sound iso639_eng olac_primary_text