Lesson 17: NLP for Vietnamese — Challenges & Solutions
Vietnamese language characteristics: word segmentation (VnCoreNLP, underthesea), diacritics, compound words. PhoBERT, ViT5, BARTpho. Vietnamese dataset: VLSP, vietnews. Benchmark models on Vietnamese tasks. Best practices for multilingual NLP.
NLP from Basics to Advanced: Mastering Natural Language Processing
Part 6: NLP Production & Modern Trends
xdev.asia
Introduction
Vietnamese belongs to the group of analytic languages (isolated languages) — fundamentally different from English. Vietnamese NLP requires a deep understanding of language specificities and specialized tools.
1. Specific Challenges
1.1 Word Segmentation — Problem #1
Vietnamese does not separate words with spaces like English:
Tiếng Anh: "machine learning" → ["machine", "learning"] ← Rõ ràng
Tiếng Việt: "học sinh học sinh học" → ???
- "học_sinh / học / sinh_học" ← học sinh ĐI học môn sinh học
- "học / sinh_học / sinh_học" ← ???
from underthesea import word_tokenize
text = "Trường đại học Bách Khoa Hà Nội là trường đại học kỹ thuật hàng đầu"
tokens = word_tokenize(text)
print(tokens)
# ['Trường', 'đại_học', 'Bách_Khoa', 'Hà_Nội', 'là', 'trường',
# 'đại_học', 'kỹ_thuật', 'hàng_đầu']
1.2 Tone Marks
# 6 thanh điệu: ngang, sắc, huyền, hỏi, ngã, nặng
# "ma", "má", "mà", "mả", "mã", "mạ" — 6 từ hoàn toàn khác nghĩa!
# Vấn đề: user thường gõ không dấu
text_no_accent = "hoc sinh hoc sinh hoc"
# Cần accent restoration trước khi xử lý NLP
1.3 Challenge Comparison Table
Challenge
English
Vietnamese
Word boundary
Space
Need word segmentation
Morphology
Inflection (run/ran/running)
No inflection
Tone/Accent
No
6 tones
Resources
A lot
Much less
Tokenizer efficiency
~1 token/word
~1.5-2 tokens/word (LLM)
2. Vietnamese NLP tool
2.1 underthesea
from underthesea import (
word_tokenize,
pos_tag,
ner,
classify,
sentiment,
)
text = "Nguyễn Phú Trọng làm việc tại Hà Nội, Việt Nam"
# Word segmentation
print(word_tokenize(text))
# POS Tagging
print(pos_tag(text))
# [('Nguyễn_Phú_Trọng', 'Np'), ('làm_việc', 'V'), ('tại', 'E'),
# ('Hà_Nội', 'Np'), (',', 'CH'), ('Việt_Nam', 'Np')]
# NER
print(ner(text))
# [('Nguyễn_Phú_Trọng', 'B-PER'), ..., ('Hà_Nội', 'B-LOC'), ...]
# Sentiment
print(sentiment("Sản phẩm này rất tốt"))
# positive
2.2 VnCoreNLP
from vncorenlp import VnCoreNLP
annotator = VnCoreNLP("VnCoreNLP-1.2.jar", annotators="wseg,pos,ner", max_heap_size='-Xmx2g')
text = "Trường Đại học Bách Khoa Hà Nội tuyển sinh năm 2026"
result = annotator.annotate(text)
for sentence in result['sentences']:
for word_info in sentence:
print(f" {word_info['form']:20s} | {word_info['posTag']:5s} | {word_info['nerLabel']}")
3. Pre-trained Models for Vietnamese
Model
Type
Base
Tasks
PhoBERT
Encoders
RoBERTa
Classification, NER, QA
BARTpho
Enc-Dec
BART
Summarization, generation
ViT5
Enc-Dec
T5
Summarization, translation
XLM-RoBERTa
Encoders
RoBERTa
Multilingual tasks
BGE-M3
Encoders
—
Multilingual embeddings
PhoBERT for Text Classification
from transformers import AutoTokenizer, AutoModelForSequenceClassification
# PhoBERT yêu cầu word segmentation TRƯỚC khi tokenize
from underthesea import word_tokenize
text = "Sản phẩm rất tốt và giao hàng nhanh"
segmented = word_tokenize(text, format="text")
# "Sản_phẩm rất tốt và giao_hàng nhanh"
tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")
model = AutoModelForSequenceClassification.from_pretrained(
"vinai/phobert-base-v2", num_labels=3
)
inputs = tokenizer(segmented, return_tensors="pt")
outputs = model(**inputs)
4. Vietnamese Datasets
Dataset
Tasks
Size
Source
VLSP 2016-2023
NER, SA, QA, WS
Varies
VLSP workshops
UIT-VSFC
Sentiments
16K reviews
UIT
vietnews
Summarization
150K articles
VietAI
PhoNER_COVID19
NER (COVID)
35K entities
VinAI
ViQuAD
Question Answering
23K QA pairs
UIT
5. Best Practices
Always word-segment before using PhoBERT
Unicode NFC normalize — especially important
Multilingual models (XLM-R, BGE-M3) are often better for zero-shot
Increase data with back-translation (EN→VI→EN) or LLM synthetic data
Summary
Aspect
Solution
Word segmentation
underthesea, VnCoreNLP
Classification/NER
PhoBERT v2
Summarization
ViT5, BARTpho
Embeddings
BGE-M3, multilingual-e5
Translation
NLLB, envit5
Next article
Lesson 18: NLP Pipeline Production — Bringing NLP models to production: serving, monitoring, CI/CD.