簡介
命名實體識別 (NER) — 命名實體識別 — 是從文本中提取實體(人、組織、地點、日期...)的問題。 NER 是資訊擷取、知識圖和聊天機器人的核心元件。
1. NER 基礎知識
常見實體類型
| 標籤 | 意義 | 範例 |
|---|---|---|
| PER | 人 | 阮文A |
| 組織 | 組織 | FPT、Google |
| 地點 | 地點 | 加州河內 |
| 日期 | 日期/时间 | 2026 年 3 月 31 日 |
| 金钱 | 货币价值 | 100万越南盾 |
| 其他 | 雜項 | 新型冠狀病毒 (COVID-19) |
IOB 標記
Text: Nguyễn Văn A làm việc tại FPT ở Hà Nội
Tags: B-PER I-PER I-PER O O O B-ORG O B-LOC I-LOC
| 前綴 | 意義 |
|---|---|
| B- | 實體的開始 |
| 我- | 實體內部(延續) |
| 哦 | 外部(非實體) |
2. NER 擁抱臉
from transformers import pipeline
# Pre-trained NER
ner = pipeline("ner", model="dslim/bert-base-NER", grouped_entities=True)
text = "Elon Musk is the CEO of Tesla and SpaceX, based in Austin, Texas"
entities = ner(text)
for e in entities:
print(f" {e['word']:20s} | {e['entity_group']:5s} | {e['score']:.4f}")
# Elon Musk | PER | 0.9987
# Tesla | ORG | 0.9956
# SpaceX | ORG | 0.9934
# Austin | LOC | 0.9891
# Texas | LOC | 0.9923
3. 微調 BERT 以實作自訂 NER
from transformers import (
AutoTokenizer,
AutoModelForTokenClassification,
Trainer,
TrainingArguments,
DataCollatorForTokenClassification,
)
from datasets import load_dataset
# Load NER dataset
dataset = load_dataset("conll2003")
# Label mapping
label_list = dataset["train"].features["ner_tags"].feature.names
id2label = {i: l for i, l in enumerate(label_list)}
label2id = {l: i for i, l in enumerate(label_list)}
# Tokenizer
tokenizer = AutoTokenizer.from_pretrained("bert-base-cased")
def tokenize_and_align_labels(examples):
tokenized = tokenizer(
examples["tokens"],
truncation=True,
is_split_into_words=True,
)
labels = []
for i, label in enumerate(examples["ner_tags"]):
word_ids = tokenized.word_ids(batch_index=i)
label_ids = []
prev_word_id = None
for word_id in word_ids:
if word_id is None:
label_ids.append(-100)
elif word_id != prev_word_id:
label_ids.append(label[word_id])
else:
label_ids.append(-100) # Subword tokens
prev_word_id = word_id
labels.append(label_ids)
tokenized["labels"] = labels
return tokenized
tokenized = dataset.map(tokenize_and_align_labels, batched=True)
# Model
model = AutoModelForTokenClassification.from_pretrained(
"bert-base-cased",
num_labels=len(label_list),
id2label=id2label,
label2id=label2id,
)
# Train
data_collator = DataCollatorForTokenClassification(tokenizer)
training_args = TrainingArguments(
output_dir="./ner-model",
num_train_epochs=3,
per_device_train_batch_size=16,
learning_rate=2e-5,
evaluation_strategy="epoch",
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["validation"],
data_collator=data_collator,
tokenizer=tokenizer,
)
trainer.train()
4. 使用 spaCy 进行 NER
import spacy
# Load pre-trained
nlp = spacy.load("en_core_web_trf") # Transformer-based
doc = nlp("Apple was founded by Steve Jobs in Cupertino, California")
for ent in doc.ents:
print(f" {ent.text:20s} | {ent.label_:10s} | {ent.start_char}-{ent.end_char}")
# Apple | ORG | 0-5
# Steve Jobs | PERSON | 22-32
# Cupertino | GPE | 36-45
# California | GPE | 47-57
5.評估:實體層級F1
from seqeval.metrics import classification_report, f1_score
# Predictions vs Ground truth (IOB format)
y_true = [["B-PER", "I-PER", "O", "B-ORG", "O"]]
y_pred = [["B-PER", "I-PER", "O", "B-ORG", "O"]]
print(classification_report(y_true, y_pred))
# Entity-level: chỉ tính đúng khi TOÀN BỘ entity đúng
⚠️ 實體級 F1 與令牌級 F1 不同:必須整個範圍(B + I 令牌)正確才能被視為正確。
總結
| 方法 | 优势 | 使用案例 |
|---|---|---|
| 預訓練管道 | 快速,無需訓練 | 一般NER |
| 微调 BERT | 自定义实体 | 特定领域 |
| 斯帕西 | 生产就绪 | 管道集成 |
下一篇文章
第 15 課:問答 — 建立智慧問答系統:使用 BERT 進行萃取 QA 和檢索增強式 QA。