摘要
Recent Chinese word segmentation (CWS) models have shown competitive performance with pre-trained language models' knowledge. However, these models tend to learn the segmentation knowledge through in-vocabulary words rather than understanding the meaning of the entire context. To address this issue, we introduce a context-aware approach that incorporates unsupervised sentence representation learning over different dropout masks into the multi-criteria training framework. We demonstrate that our approach reaches state-of-the-art (SoTA) performance on F1 scores for six of the nine CWS benchmark datasets and out-of-vocabulary (OOV) recalls for eight of nine. Further experiments discover that substantial improvements can be brought with various sentence representation objectives.
| 原文 | 英語 |
|---|---|
| 主出版物標題 | Findings of the Association for Computational Linguistics |
| 主出版物子標題 | EMNLP 2023 |
| 發行者 | Association for Computational Linguistics (ACL) |
| 頁面 | 12756-12763 |
| 頁數 | 8 |
| ISBN(電子) | 9798891760615 |
| DOIs | |
| 出版狀態 | 已出版 - 2023 |
| 對外發佈 | 是 |
| 事件 | 2023 Findings of the Association for Computational Linguistics: EMNLP 2023 - Hybrid, 新加坡 持續時間: 06 12 2023 → 10 12 2023 |
出版系列
| 名字 | Findings of the Association for Computational Linguistics: EMNLP 2023 |
|---|
Conference
| Conference | 2023 Findings of the Association for Computational Linguistics: EMNLP 2023 |
|---|---|
| 國家/地區 | 新加坡 |
| 城市 | Hybrid |
| 期間 | 06/12/23 → 10/12/23 |
文獻附註
Publisher Copyright:© 2023 Association for Computational Linguistics.
指紋
深入研究「Improving Multi-Criteria Chinese Word Segmentation through Learning Sentence Representation」主題。共同形成了獨特的指紋。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver