Abstract
This paper proposes an intrusion detection method that integrates language models with textualized network flow features. The core idea is to transform the original numerical and categorical attributes of network traffic into textual sequences and utilize the semantic modeling capability of language models to map these features into a readable vector space. This approach enables the model to effectively capture contextual relationships among features, thereby maintaining strong generalization performance even when faced with previously unseen attacks. In the classification stage, we combine language embeddings with boosting-based ensemble learning, employing a progressive weighting strategy to emphasize hard-to-classify samples and reduce the risk of overfitting in individual models. Experiments conducted in the CSE-CIC-IDS2018 dataset demonstrate outstanding performance in multiple evaluation metrics, achieving an overall accuracy of 99.66%. The results confirm that the integration of textualized features with language embeddings can clearly enhance the effectiveness and generalization capability of malicious traffic detection and provide a robust and realistic approach for modern-day intrusion detection systems.
| Original language | English |
|---|---|
| Title of host publication | Advances in Natural Language Processing and Information Retrieval |
| Editors | Herwig Unger, Phayung Meesad |
| Publisher | Springer Science and Business Media Deutschland GmbH |
| Pages | 1001-1011 |
| Number of pages | 11 |
| ISBN (Print) | 9783032208965 |
| DOIs | |
| State | Published - 2026 |
| Event | 9th International Conference on Natural Language Processing and Information Retrieval, NLPIR 2025 - Fukuoka, Japan Duration: 12 12 2025 → 14 12 2025 |
Publication series
| Name | Lecture Notes in Networks and Systems |
|---|---|
| Volume | 1904 LNNS |
| ISSN (Print) | 2367-3370 |
| ISSN (Electronic) | 2367-3389 |
Conference
| Conference | 9th International Conference on Natural Language Processing and Information Retrieval, NLPIR 2025 |
|---|---|
| Country/Territory | Japan |
| City | Fukuoka |
| Period | 12/12/25 → 14/12/25 |
Bibliographical note
Publisher Copyright:© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
Keywords
- Boosting Ensemble Learning
- Intrusion Detection System
- Language Model
- Semantic Embedding
Fingerprint
Dive into the research topics of 'Language Model Embedding and Boosting Ensemble Learning for Malicious Intrusion Detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver