Skip to main navigation Skip to search Skip to main content

Development of a taiwanese speech and text corpus

  • Tzu Yu Liao
  • , Ren Yuan Lyu
  • , Ming Tat Ko
  • , Yuang Chin Chiang
  • , Jyh Shing Roger Jang
  • Academia Sinica - Institute of Information Science
  • National Tsing Hua University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The main goal of this paper is to develop a large scale Taiwanese corpus. In the mean time, we try to establish a successful model for the computational linguistic research on other minority Taiwanese languages such as Haka. In this paper, we will build a Taiwanese speech corpus. The source of speech corpus is Taiwanese dramas and news from TV stations. The goal of the corpus is 200 hours speech material with annotation.

Original languageEnglish
Title of host publicationProceedings of the 24th Conference on Computational Linguistics and Speech Processing, ROCLING 2012
Pages102-111
Number of pages10
StatePublished - 2012
Event24th Conference on Computational Linguistics and Speech Processing, ROCLING 2012 - Chung-Li, Taiwan
Duration: 21 09 201222 09 2012

Publication series

NameProceedings of the 24th Conference on Computational Linguistics and Speech Processing, ROCLING 2012

Conference

Conference24th Conference on Computational Linguistics and Speech Processing, ROCLING 2012
Country/TerritoryTaiwan
CityChung-Li
Period21/09/1222/09/12

Keywords

  • Corpus
  • Speech recognition
  • Taiwanese
  • Transcription

Fingerprint

Dive into the research topics of 'Development of a taiwanese speech and text corpus'. Together they form a unique fingerprint.

Cite this