Chuyển đến nội dung chính

Lesson 14: Named Entity Recognition (NER) — Entity Extraction

What is NER: entity types (PER, ORG, LOC, DATE). IOB/BIO tagging. CRF for sequence labeling. Fine-tune BERT for NER. spaCy NER training. Custom entity types for domain-specific (medical, legal). Evaluation: entity-level F1.

🧠 AI & ML — Lesson 13 Lesson 14: Named Entity Recognition (NER) — Extract Entities

NLP from Basics to Advanced: Mastering Natural Language Processing

Part 5: Applied NLP problems — Hands-on Projects

xdev.asia

Introduction

Named Entity Recognition (NER) — named entity recognition — is the problem of extracting entities (people, organizations, places, dates...) from text. NER is a core component in information extraction, knowledge graphs, and chatbots.


1. NER Basics

Common Entity Types

TagsMeaningExample
PERPersonNguyen Van A
ORGOrganizationFPT, Google
LOCLocationHanoi, California
DATEDate/TimeMarch 31, 2026
MONEYMonetary value1 million VND
MISCMiscellaneousCOVID-19

IOB Tagging

Text:   Nguyễn Văn A  làm  việc  tại  FPT   ở  Hà  Nội
Tags:   B-PER  I-PER  I-PER  O    O    O   B-ORG  O  B-LOC I-LOC
PrefixMeaning
B-Beginning of entities
I-Inside (continuation) of entity
OOutside (not an entity)

2. NER with Hugging Face

from transformers import pipeline

# Pre-trained NER
ner = pipeline("ner", model="dslim/bert-base-NER", grouped_entities=True)

text = "Elon Musk is the CEO of Tesla and SpaceX, based in Austin, Texas"
entities = ner(text)

for e in entities:
    print(f"  {e['word']:20s} | {e['entity_group']:5s} | {e['score']:.4f}")
# Elon Musk            | PER   | 0.9987
# Tesla                | ORG   | 0.9956
# SpaceX               | ORG   | 0.9934
# Austin               | LOC   | 0.9891
# Texas                | LOC   | 0.9923

3. Fine-tune BERT for Custom NER

from transformers import (
    AutoTokenizer,
    AutoModelForTokenClassification,
    Trainer,
    TrainingArguments,
    DataCollatorForTokenClassification,
)
from datasets import load_dataset

# Load NER dataset
dataset = load_dataset("conll2003")

# Label mapping
label_list = dataset["train"].features["ner_tags"].feature.names
id2label = {i: l for i, l in enumerate(label_list)}
label2id = {l: i for i, l in enumerate(label_list)}

# Tokenizer
tokenizer = AutoTokenizer.from_pretrained("bert-base-cased")

def tokenize_and_align_labels(examples):
    tokenized = tokenizer(
        examples["tokens"],
        truncation=True,
        is_split_into_words=True,
    )
    labels = []
    for i, label in enumerate(examples["ner_tags"]):
        word_ids = tokenized.word_ids(batch_index=i)
        label_ids = []
        prev_word_id = None
        for word_id in word_ids:
            if word_id is None:
                label_ids.append(-100)
            elif word_id != prev_word_id:
                label_ids.append(label[word_id])
            else:
                label_ids.append(-100)  # Subword tokens
            prev_word_id = word_id
        labels.append(label_ids)
    tokenized["labels"] = labels
    return tokenized

tokenized = dataset.map(tokenize_and_align_labels, batched=True)

# Model
model = AutoModelForTokenClassification.from_pretrained(
    "bert-base-cased",
    num_labels=len(label_list),
    id2label=id2label,
    label2id=label2id,
)

# Train
data_collator = DataCollatorForTokenClassification(tokenizer)

training_args = TrainingArguments(
    output_dir="./ner-model",
    num_train_epochs=3,
    per_device_train_batch_size=16,
    learning_rate=2e-5,
    evaluation_strategy="epoch",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized["train"],
    eval_dataset=tokenized["validation"],
    data_collator=data_collator,
    tokenizer=tokenizer,
)

trainer.train()

4. NER with spaCy

import spacy

# Load pre-trained
nlp = spacy.load("en_core_web_trf")  # Transformer-based

doc = nlp("Apple was founded by Steve Jobs in Cupertino, California")

for ent in doc.ents:
    print(f"  {ent.text:20s} | {ent.label_:10s} | {ent.start_char}-{ent.end_char}")
# Apple                | ORG        | 0-5
# Steve Jobs           | PERSON     | 22-32
# Cupertino            | GPE        | 36-45
# California           | GPE        | 47-57

5. Evaluation: Entity-level F1

from seqeval.metrics import classification_report, f1_score

# Predictions vs Ground truth (IOB format)
y_true = [["B-PER", "I-PER", "O", "B-ORG", "O"]]
y_pred = [["B-PER", "I-PER", "O", "B-ORG", "O"]]

print(classification_report(y_true, y_pred))
# Entity-level: chỉ tính đúng khi TOÀN BỘ entity đúng

⚠️ Entity-level F1 is different from token-level F1: must be entire span (B + I tokens) correct to be considered correct.


Summary

ApproachAdvantagesUse cases
Pre-trained pipelineFast, no trainingGeneral NER
Fine-tune BERTCustom entitiesDomain-specific
spaCyProduction-readyPipeline integration

Next article

Lesson 15: Question Answering — Building a smart question and answer system: extractive QA with BERT and retrieval-augmented QA.