ViGPTQA-state-of-the-art LLMs for vietnamese question answering : system overview, core models training, and evaluations
(2023) p.754-764- Abstract
- Large language models (LLMs) and their applications in low-resource languages (such as in Vietnamese) are limited due to lack of training data and benchmarking datasets. This paper introduces a practical real-world implementation of a question answering system for Vietnamese, called ViGPTQA, leveraging the power of LLM. Since there is no effective LLM in Vietnamese to date, we also propose, evaluate, and open-source an instruction-tuned LLM for Vietnamese, named ViGPT. ViGPT demonstrates exceptional performances, especially on real-world scenarios. We curate a new set of benchmark datasets that encompass both AI and human-generated data, providing a comprehensive evaluation framework for Vietnamese LLMs. By achieving state-of-the-art... (More)
- Large language models (LLMs) and their applications in low-resource languages (such as in Vietnamese) are limited due to lack of training data and benchmarking datasets. This paper introduces a practical real-world implementation of a question answering system for Vietnamese, called ViGPTQA, leveraging the power of LLM. Since there is no effective LLM in Vietnamese to date, we also propose, evaluate, and open-source an instruction-tuned LLM for Vietnamese, named ViGPT. ViGPT demonstrates exceptional performances, especially on real-world scenarios. We curate a new set of benchmark datasets that encompass both AI and human-generated data, providing a comprehensive evaluation framework for Vietnamese LLMs. By achieving state-of-the-art results and approaching other multilingual LLMs, our instruction-tuned LLM underscores the need for dedicated Vietnamese-specific LLMs. Our open-source model supports customized and privacy-fulfilled Vietnamese language processing systems. (Less)
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/record/e9a2135b-41c5-4145-96e6-6c1b90022303
- author
- Nguyen, Minh Thuan ; Tran, Khanh-Tung ; Nguyen, Vincent and Vu, Xuan-Son LU
- publishing date
- 2023
- type
- Chapter in Book/Report/Conference proceeding
- publication status
- published
- subject
- host publication
- Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing : industry Track - industry Track
- editor
- Wang, Mingxuan and Zitouni, Imed
- pages
- 754 - 764
- publisher
- Association for Computational Linguistics
- external identifiers
-
- scopus:85180767795
- DOI
- 10.18653/v1/2023.emnlp-industry.70
- language
- English
- LU publication?
- no
- id
- e9a2135b-41c5-4145-96e6-6c1b90022303
- date added to LUP
- 2026-02-10 23:56:07
- date last changed
- 2026-03-10 10:23:24
@inproceedings{e9a2135b-41c5-4145-96e6-6c1b90022303,
abstract = {{Large language models (LLMs) and their applications in low-resource languages (such as in Vietnamese) are limited due to lack of training data and benchmarking datasets. This paper introduces a practical real-world implementation of a question answering system for Vietnamese, called ViGPTQA, leveraging the power of LLM. Since there is no effective LLM in Vietnamese to date, we also propose, evaluate, and open-source an instruction-tuned LLM for Vietnamese, named ViGPT. ViGPT demonstrates exceptional performances, especially on real-world scenarios. We curate a new set of benchmark datasets that encompass both AI and human-generated data, providing a comprehensive evaluation framework for Vietnamese LLMs. By achieving state-of-the-art results and approaching other multilingual LLMs, our instruction-tuned LLM underscores the need for dedicated Vietnamese-specific LLMs. Our open-source model supports customized and privacy-fulfilled Vietnamese language processing systems.}},
author = {{Nguyen, Minh Thuan and Tran, Khanh-Tung and Nguyen, Vincent and Vu, Xuan-Son}},
booktitle = {{Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing : industry Track}},
editor = {{Wang, Mingxuan and Zitouni, Imed}},
language = {{eng}},
pages = {{754--764}},
publisher = {{Association for Computational Linguistics}},
title = {{ViGPTQA-state-of-the-art LLMs for vietnamese question answering : system overview, core models training, and evaluations}},
url = {{http://dx.doi.org/10.18653/v1/2023.emnlp-industry.70}},
doi = {{10.18653/v1/2023.emnlp-industry.70}},
year = {{2023}},
}