Hyperparameter optimization for phobert in sentiment analysis of Vietnamese tourism data

Abstract

In the field of Vietnamese Natural Language Processing (NLP), the pre-trained language model PhoBERT has demonstrated superior performance in text classification tasks due to its ability to capture contextual and linguistic characteristics effectively. This paper presents a comprehensive study on the impact of hyperparameters during the fine-tuning process of the PhoBERT model for sentiment analysis of Vietnamese tourism reviews. The experimental process focuses on evaluating the influence of key hyperparameters, including learning rate, batch size, number of training epochs, maximum sequence length, and dropout rate. The results indicate that appropriate hyperparameter selection significantly affects model performance, leading to improvements in both accuracy and F1-score compared to the default configuration. In particular, the learning rate and maximum sequence length are identified as the two most influential factors, reflecting the characteristics of Vietnamese online tourism reviews. This study not only proposes an optimal set of hyperparameters for the PhoBERT model but also addresses the lack of empirical studies on hyperparameter tuning for Vietnamese tourism-related datasets.

https://doi.org/10.26459/hueunijtt.v135i2A.8449
In-Press (Vietnamese)
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 International License.

Copyright (c) 2026 Array