SPCS Speech Corpus

Modipa, T. I.; Davel, M. H.; De Wet, F.

Title	SPCS Speech Corpus
Description	Broadband speech corpus of approximately 10 hours and the corresponding transcriptions. The development process of the corpus involved the recording and transcribing of radio broadcasts. The transcriptions were used to generate the Sepedi code-switched prompts to re-record speech from multiple speakers. The following sub-directories are found in this directory: Audio: Audio files for all the recorded code-switched speech Transcriptions: The corresponding orthographic transcriptions Metadata: Information about the speakers and the transcriptions Documentation: The directory structure and the Sepedi prompt list
Contact name	Ulrike Janke
Contact email	ulrike.must@gmail.com
Publisher(s)	Council for Scientific and Industrial Research; North-West University
License	Creative Commons Attribution 2.5 South Africa license: https://creativecommons.org/licenses/by/2.5/legalcode
Language(s)	English; Sepedi
Author(s)	Modipa, T. I.; Davel, M. H.; De Wet, F.
Subject	Sepedi; Sesotho sa Leboa; code-switching; orthographic transcription; English
Citation	T. I. Modipa, M. H. Davel, F. De Wet, "Implications of Sepedi/English code switching for ASR systems", Pattern Recognition Association of South Africa, pp. 112-117, 2015
URI	https://hdl.handle.net/20.500.12185/530
Media type	Speech
Media category	Orthographic transcribed broadband speech corpus
Primary collection	Resource Catalogue
Secondary collection	Resource Index
ISO639 code	eng; nso
Submit date	2020-04-21T10:15:12Z
Date available	2020-04-21T10:15:12Z
Date created	2015-11-25

Files in this item

Name:: spcs.tar.gz
Size:: 867.0Mb
Format:: Unknown
MD5:: 2c2b367ba1811e4024b52b95bf85acf8

Download

This item appears in the following Collection(s)

Resource Catalogue [335]
A collection of language resources available for download from the RMA of SADiLaR. The collection mostly consists of resources developed with funding from the Department of Arts and Culture.
Resource Index [386]
A collection of language resource metadata mostly collected during the NHN funded technology audit of 2009, as well as the SADiLaR technology audit of 2018. Not all resources in this collection are available for download.

Show simple item record

SPCS Speech Corpus

Files in this item

License agreement

This item appears in the following Collection(s)