Skip to main navigation Skip to search Skip to main content

Language Model Embedding and Boosting Ensemble Learning for Malicious Intrusion Detection

  • Chang Gung University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

This paper proposes an intrusion detection method that integrates language models with textualized network flow features. The core idea is to transform the original numerical and categorical attributes of network traffic into textual sequences and utilize the semantic modeling capability of language models to map these features into a readable vector space. This approach enables the model to effectively capture contextual relationships among features, thereby maintaining strong generalization performance even when faced with previously unseen attacks. In the classification stage, we combine language embeddings with boosting-based ensemble learning, employing a progressive weighting strategy to emphasize hard-to-classify samples and reduce the risk of overfitting in individual models. Experiments conducted in the CSE-CIC-IDS2018 dataset demonstrate outstanding performance in multiple evaluation metrics, achieving an overall accuracy of 99.66%. The results confirm that the integration of textualized features with language embeddings can clearly enhance the effectiveness and generalization capability of malicious traffic detection and provide a robust and realistic approach for modern-day intrusion detection systems.

Original languageEnglish
Title of host publicationAdvances in Natural Language Processing and Information Retrieval
EditorsHerwig Unger, Phayung Meesad
PublisherSpringer Science and Business Media Deutschland GmbH
Pages1001-1011
Number of pages11
ISBN (Print)9783032208965
DOIs
StatePublished - 2026
Event9th International Conference on Natural Language Processing and Information Retrieval, NLPIR 2025 - Fukuoka, Japan
Duration: 12 12 202514 12 2025

Publication series

NameLecture Notes in Networks and Systems
Volume1904 LNNS
ISSN (Print)2367-3370
ISSN (Electronic)2367-3389

Conference

Conference9th International Conference on Natural Language Processing and Information Retrieval, NLPIR 2025
Country/TerritoryJapan
CityFukuoka
Period12/12/2514/12/25

Bibliographical note

Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.

Keywords

  • Boosting Ensemble Learning
  • Intrusion Detection System
  • Language Model
  • Semantic Embedding

Fingerprint

Dive into the research topics of 'Language Model Embedding and Boosting Ensemble Learning for Malicious Intrusion Detection'. Together they form a unique fingerprint.

Cite this