Chuyển đến nội dung chính

Lesson 4: Bag of Words, TF-IDF & N-grams — Classical Method

Bag of Words model. TF-IDF weighting and mathematical intuition. N-grams for language modeling. CountVectorizer and TfidfVectorizer with scikit-learn. Pros and cons and when is it still effective?

🧠 AI & ML — Lesson 3 Lesson 4: Bag of Words, TF-IDF & N-grams — Classical Method

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 2: Language Representation — From BoW to Word Embeddings

xdev.asia

Introduction

Before Word2Vec and Transformer, NLP relied on simple but surprisingly effective word counting methods. Bag of Words and TF-IDF are still widely used in 2026 — especially when you need fast baselines or small data.


1. Bag of Words (BoW)

Ideas

Represent each document by frequency count vector of words, ignoring order.

from sklearn.feature_extraction.text import CountVectorizer

corpus = [
    "NLP rất thú vị",
    "Machine Learning rất hay",
    "NLP và Machine Learning bổ trợ nhau",
]

vectorizer = CountVectorizer()
X = vectorizer.fit_transform(corpus)

print(vectorizer.get_feature_names_out())
# ['learning', 'machine', 'và', 'nlp', 'nhau', 'bổ', 'hay', 'rất', 'thú', 'trợ', 'vị']

print(X.toarray())
# [[0, 0, 0, 1, 0, 0, 0, 1, 1, 0, 1],  ← "NLP rất thú vị"
#  [1, 1, 0, 0, 0, 0, 1, 1, 0, 0, 0],  ← "ML rất hay"
#  [1, 1, 1, 1, 1, 1, 0, 0, 0, 1, 0]]  ← "NLP và ML bổ trợ nhau"

Limitations

  • Out of order: "dog bites man" = "man bites dog"
  • Sparse matrix: Very sparse vector (most = 0)
  • Do not represent meaning: Synonyms have different vectors

2. TF-IDF (Term Frequency – Inverse Document Frequency)

Intuition

  • TF (Term Frequency): Words that appear a lot in the document → are important to that document
  • IDF (Inverse Document Frequency): Words appear in fewer documents → more important (discriminative)

$$TF\text{-}IDF(t, d) = TF(t, d) \times IDF(t) = \frac{f_{t,d}}{\sum_{t'} f_{t',d}} \times \log\frac{N}{n_t}$$

from sklearn.feature_extraction.text import TfidfVectorizer

corpus = [
    "NLP xử lý ngôn ngữ tự nhiên",
    "Machine Learning học từ dữ liệu",
    "NLP kết hợp Machine Learning để xử lý ngôn ngữ",
]

tfidf = TfidfVectorizer()
X = tfidf.fit_transform(corpus)

# TF-IDF values — từ "xử" và "lý" có IDF thấp vì xuất hiện nhiều docs
import pandas as pd
df = pd.DataFrame(X.toarray(), columns=tfidf.get_feature_names_out())
print(df.round(2))

When is TF-IDF still "good"?

  • Text search / Information Retrieval
  • Keyword extraction
  • Baseline classification with small datasets (< 10K samples)
  • Feature engineering combined with deep learning

3. N-grams

Ideas

Instead of considering individual words, consider sequence of n consecutive words:

NNameExample ("NLP is cool")
1Unigram"NLP", "very", "interesting", "tasteful"
2Bigram"NLP is very", "very interesting", "interesting"
3Trigram"NLP is very interesting", "very interesting"
from sklearn.feature_extraction.text import CountVectorizer

# Bigram + Unigram
vectorizer = CountVectorizer(ngram_range=(1, 2))
X = vectorizer.fit_transform(["NLP rất thú vị và hay"])
print(vectorizer.get_feature_names_out())
# ['hay', 'nlp', 'nlp rất', 'rất', 'rất thú', 'thú', 'thú vị', 'và', 'và hay', 'vị', 'vị và']

N-grams help BoW/TF-IDF capture part of word order — "interesting" has a different meaning than "interesting".


4. Application: Text Classification with TF-IDF

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report

# Dataset ví dụ
texts = [
    "GPU mới của NVIDIA rất mạnh",
    "Giá Bitcoin tăng vọt hôm nay",
    "Đội tuyển Việt Nam thắng 3-0",
    "Transformer architecture cải tiến NLP",
    "Chứng khoán phục hồi sau phiên giảm",
    "World Cup 2026 sẽ tổ chức tại 3 nước",
    # ... thêm data
]
labels = ["tech", "finance", "sports", "tech", "finance", "sports"]

# Pipeline
tfidf = TfidfVectorizer(ngram_range=(1, 2), max_features=5000)
X = tfidf.fit_transform(texts)

model = LogisticRegression()
model.fit(X, labels)

# Predict
new_text = ["Apple ra mắt iPhone mới"]
prediction = model.predict(tfidf.transform(new_text))
print(prediction)  # ['tech']

Summary

MethodAdvantagesLimitationsUse cases
BoWSimple, fastLoss of order, sparseBaseline, word count
TF-IDFConsider importanceLoss of order, sparseSearch, keywords, classification
N-gramsGetting the local contextVocab explosionCombined with BoW/TF-IDF

Next article

Lesson 5: Word Embeddings — Word2Vec, GloVe, FastText — Leap: representing words with dense meaningful vectors, where "king - man + woman ≈ queen".