Chuyển đến nội dung chính

Lesson 17: NLP for Vietnamese — Challenges & Solutions

Vietnamese language characteristics: word segmentation (VnCoreNLP, underthesea), diacritics, compound words. PhoBERT, ViT5, BARTpho. Vietnamese dataset: VLSP, vietnews. Benchmark models on Vietnamese tasks. Best practices for multilingual NLP.

🧠 AI & ML — Lesson 16 Lesson 17: NLP for Vietnamese — Challenges & Solution

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 6: NLP Production & Modern Trends

xdev.asia

Introduction

Vietnamese belongs to the group of analytic languages (isolated languages) — fundamentally different from English. Vietnamese NLP requires a deep understanding of language specificities and specialized tools.


1. Specific Challenges

1.1 Word Segmentation — Problem #1

Vietnamese does not separate words with spaces like English:

Tiếng Anh: "machine learning"     → ["machine", "learning"]  ← Rõ ràng
Tiếng Việt: "học sinh học sinh học" → ???
  - "học_sinh / học / sinh_học"    ← học sinh ĐI học môn sinh học
  - "học / sinh_học / sinh_học"    ← ???
from underthesea import word_tokenize

text = "Trường đại học Bách Khoa Hà Nội là trường đại học kỹ thuật hàng đầu"
tokens = word_tokenize(text)
print(tokens)
# ['Trường', 'đại_học', 'Bách_Khoa', 'Hà_Nội', 'là', 'trường',
#  'đại_học', 'kỹ_thuật', 'hàng_đầu']

1.2 Tone Marks

# 6 thanh điệu: ngang, sắc, huyền, hỏi, ngã, nặng
# "ma", "má", "mà", "mả", "mã", "mạ" — 6 từ hoàn toàn khác nghĩa!

# Vấn đề: user thường gõ không dấu
text_no_accent = "hoc sinh hoc sinh hoc"
# Cần accent restoration trước khi xử lý NLP

1.3 Challenge Comparison Table

ChallengeEnglishVietnamese
Word boundarySpaceNeed word segmentation
MorphologyInflection (run/ran/running)No inflection
Tone/AccentNo6 tones
ResourcesA lotMuch less
Tokenizer efficiency~1 token/word~1.5-2 tokens/word (LLM)

2. Vietnamese NLP tool

2.1 underthesea

from underthesea import (
    word_tokenize,
    pos_tag,
    ner,
    classify,
    sentiment,
)

text = "Nguyễn Phú Trọng làm việc tại Hà Nội, Việt Nam"

# Word segmentation
print(word_tokenize(text))

# POS Tagging
print(pos_tag(text))
# [('Nguyễn_Phú_Trọng', 'Np'), ('làm_việc', 'V'), ('tại', 'E'),
#  ('Hà_Nội', 'Np'), (',', 'CH'), ('Việt_Nam', 'Np')]

# NER
print(ner(text))
# [('Nguyễn_Phú_Trọng', 'B-PER'), ..., ('Hà_Nội', 'B-LOC'), ...]

# Sentiment
print(sentiment("Sản phẩm này rất tốt"))
# positive

2.2 VnCoreNLP

from vncorenlp import VnCoreNLP

annotator = VnCoreNLP("VnCoreNLP-1.2.jar", annotators="wseg,pos,ner", max_heap_size='-Xmx2g')

text = "Trường Đại học Bách Khoa Hà Nội tuyển sinh năm 2026"
result = annotator.annotate(text)

for sentence in result['sentences']:
    for word_info in sentence:
        print(f"  {word_info['form']:20s} | {word_info['posTag']:5s} | {word_info['nerLabel']}")

3. Pre-trained Models for Vietnamese

ModelTypeBaseTasks
PhoBERTEncodersRoBERTaClassification, NER, QA
BARTphoEnc-DecBARTSummarization, generation
ViT5Enc-DecT5Summarization, translation
XLM-RoBERTaEncodersRoBERTaMultilingual tasks
BGE-M3Encoders—Multilingual embeddings

PhoBERT for Text Classification

from transformers import AutoTokenizer, AutoModelForSequenceClassification

# PhoBERT yêu cầu word segmentation TRƯỚC khi tokenize
from underthesea import word_tokenize

text = "Sản phẩm rất tốt và giao hàng nhanh"
segmented = word_tokenize(text, format="text")
# "Sản_phẩm rất tốt và giao_hàng nhanh"

tokenizer = AutoTokenizer.from_pretrained("vinai/phobert-base-v2")
model = AutoModelForSequenceClassification.from_pretrained(
    "vinai/phobert-base-v2", num_labels=3
)

inputs = tokenizer(segmented, return_tensors="pt")
outputs = model(**inputs)

4. Vietnamese Datasets

DatasetTasksSizeSource
VLSP 2016-2023NER, SA, QA, WSVariesVLSP workshops
UIT-VSFCSentiments16K reviewsUIT
vietnewsSummarization150K articlesVietAI
PhoNER_COVID19NER (COVID)35K entitiesVinAI
ViQuADQuestion Answering23K QA pairsUIT

5. Best Practices

  1. Always word-segment before using PhoBERT
  2. Unicode NFC normalize — especially important
  3. Multilingual models (XLM-R, BGE-M3) are often better for zero-shot
  4. Increase data with back-translation (EN→VI→EN) or LLM synthetic data

Summary

AspectSolution
Word segmentationunderthesea, VnCoreNLP
Classification/NERPhoBERT v2
SummarizationViT5, BARTpho
EmbeddingsBGE-M3, multilingual-e5
TranslationNLLB, envit5

Next article

Lesson 18: NLP Pipeline Production — Bringing NLP models to production: serving, monitoring, CI/CD.