跳至主導覽 跳至搜尋 跳過主要內容

Improving Multi-Criteria Chinese Word Segmentation through Learning Sentence Representation

  • Chun Yi Lin
  • , Ying Jia Lin
  • , Yi Ting Li
  • , Chia Jen Yeh
  • , Ching Wen Yang
  • , Hung Yu Kao
  • National Cheng Kung University

研究成果: 圖書/報告稿件的類型會議稿件同行評審

4 引文 斯高帕斯(Scopus)

摘要

Recent Chinese word segmentation (CWS) models have shown competitive performance with pre-trained language models' knowledge. However, these models tend to learn the segmentation knowledge through in-vocabulary words rather than understanding the meaning of the entire context. To address this issue, we introduce a context-aware approach that incorporates unsupervised sentence representation learning over different dropout masks into the multi-criteria training framework. We demonstrate that our approach reaches state-of-the-art (SoTA) performance on F1 scores for six of the nine CWS benchmark datasets and out-of-vocabulary (OOV) recalls for eight of nine. Further experiments discover that substantial improvements can be brought with various sentence representation objectives.

原文英語
主出版物標題Findings of the Association for Computational Linguistics
主出版物子標題EMNLP 2023
發行者Association for Computational Linguistics (ACL)
頁面12756-12763
頁數8
ISBN(電子)9798891760615
DOIs
出版狀態已出版 - 2023
對外發佈
事件2023 Findings of the Association for Computational Linguistics: EMNLP 2023 - Hybrid, 新加坡
持續時間: 06 12 202310 12 2023

出版系列

名字Findings of the Association for Computational Linguistics: EMNLP 2023

Conference

Conference2023 Findings of the Association for Computational Linguistics: EMNLP 2023
國家/地區新加坡
城市Hybrid
期間06/12/2310/12/23

文獻附註

Publisher Copyright:
© 2023 Association for Computational Linguistics.

指紋

深入研究「Improving Multi-Criteria Chinese Word Segmentation through Learning Sentence Representation」主題。共同形成了獨特的指紋。

引用此