Healthcare data is unlike any data you have ever worked with. This article explains why — and how to work with it properly from day one.
1. DICOM — Medical imaging standard
DICOM (Digital Imaging and Communications in Medicine) is an international standard for storing and transmitting medical images. Launched in 1993, today all hospitals around the world use DICOM.
1.1. DICOM file structure
DICOM isn't just an image — it's image + metadata:
File DICOM (.dcm)
├── File Meta Information
│ ├── Transfer Syntax UID (cách encode dữ liệu pixel)
│ └── SOP Class UID (loại hình ảnh: CT, MRI, X-ray...)
│
└── Data Set (hàng trăm tags)
├── Patient Info
│ ├── (0010,0010) PatientName = "Nguyen^Van^A"
│ ├── (0010,0020) PatientID = "BN-2024-001"
│ ├── (0010,0030) PatientBirthDate = "19650315"
│ └── (0010,0040) PatientSex = "M"
│
├── Study Info
│ ├── (0020,000D) StudyInstanceUID
│ ├── (0008,0020) StudyDate = "20240115"
│ └── (0008,1030) StudyDescription = "CHEST AP"
│
├── Image Info
│ ├── (0028,0010) Rows = 2048
│ ├── (0028,0011) Columns = 2048
│ ├── (0028,0030) PixelSpacing = [0.175, 0.175] mm/pixel
│ ├── (0028,1050) WindowCenter = 40
│ ├── (0028,1051) WindowWidth = 400
│ └── (0028,0103) PixelRepresentation = 1 (signed integers)
│
└── Pixel Data
└── (7FE0,0010) PixelData = [raw 12-bit pixel values...]
1.2. Read DICOM with pydicom
import pydicom
import numpy as np
import matplotlib.pyplot as plt
# Đọc file DICOM
dcm = pydicom.dcmread("chest_xray.dcm")
# Xem metadata
print(f"Patient: {dcm.PatientName}")
print(f"Study Date: {dcm.StudyDate}")
print(f"Modality: {dcm.Modality}") # CR (X-ray), CT, MR, US...
print(f"Image size: {dcm.Rows} x {dcm.Columns}")
print(f"Pixel spacing: {dcm.PixelSpacing} mm/pixel")
# Lấy pixel array (giá trị thô — chưa phải HU units!)
raw_pixels = dcm.pixel_array
print(f"Pixel shape: {raw_pixels.shape}")
print(f"Pixel dtype: {raw_pixels.dtype}") # Thường là int16 cho CT
print(f"Min/Max values: {raw_pixels.min()} / {raw_pixels.max()}")
# Chuyển sang Hounsfield Units (HU) cho CT
# Công thức: HU = pixel_value * RescaleSlope + RescaleIntercept
def to_hounsfield_units(dcm_file):
pixels = dcm_file.pixel_array.astype(np.float32)
slope = float(dcm_file.RescaleSlope) if hasattr(dcm_file, 'RescaleSlope') else 1.0
intercept = float(dcm_file.RescaleIntercept) if hasattr(dcm_file, 'RescaleIntercept') else 0.0
return pixels * slope + intercept
hu_image = to_hounsfield_units(dcm)
# HU reference values:
# Air: -1000 HU | Lung: -500 HU | Fat: -100 HU
# Water/Soft tissue: 0-80 HU | Blood: 40-80 HU | Bone: 400-1000 HU
1.3. Windowing — The core of Medical Image Display
The screen only displays 256 gray levels, but the CT has ~4000 HU levels. Windowing select the HU range to view:
def apply_windowing(image_hu: np.ndarray, window_center: int, window_width: int) -> np.ndarray:
"""
Window Center (WC) và Window Width (WW) cho các mô khác nhau:
Phổi: WC=-600, WW=1500 (thấy cấu trúc phổi chi tiết)
Ổ bụng: WC=40, WW=400 (thấy gan, lách, thận)
Não: WC=40, WW=80 (thấy hematoma)
Xương: WC=400, WW=1800 (thấy gãy xương)
"""
lower = window_center - window_width / 2
upper = window_center + window_width / 2
# Clip và normalize về [0, 1]
windowed = np.clip(image_hu, lower, upper)
normalized = (windowed - lower) / (upper - lower)
return normalized
# Hiển thị cùng một CT scan với 3 window settings khác nhau
fig, axes = plt.subplots(1, 3, figsize=(15, 5))
windows = [
("Lung Window", -600, 1500),
("Soft Tissue", 40, 400),
("Bone Window", 400, 1800),
]
for ax, (title, wc, ww) in zip(axes, windows):
windowed = apply_windowing(hu_image, wc, ww)
ax.imshow(windowed, cmap='gray')
ax.set_title(f"{title}\nWC={wc}, WW={ww}")
ax.axis('off')
plt.tight_layout()
plt.savefig("windowing_comparison.png", dpi=150)
2. HL7 FHIR — Clinical data exchange standard
HL7 FHIR (Fast Healthcare Interoperability Resources) is the most modern standard for data exchange between healthcare systems. Not images — this is structured clinical data.
2.1. Core Resources
FHIR organizes data into "Resources" — each entity type has its own schema:
FHIR Resources (chỉ các resource quan trọng nhất cho AI)
│
├── Patient — thông tin bệnh nhân
├── Observation — vital signs, lab results, measurements
├── Condition — diagnoses, problems (mã ICD-10/11)
├── MedicationRequest— đơn thuốc
├── Procedure — phẫu thuật, can thiệp (mã CPT/ICD)
├── DiagnosticReport — báo cáo xét nghiệm, hình ảnh
├── Encounter — lần khám/nhập viện
└── AllergyIntolerance — dị ứng
2.2. Works with FHIR API
import requests
from datetime import datetime, timedelta
# FHIR R4 RESTful API
FHIR_SERVER = "https://r4.smarthealthit.org" # Public test server
def get_patient_observations(patient_id: str, code: str) -> list:
"""
Lấy lab results của bệnh nhân
code: LOINC code — chuẩn mã cho lab tests
Ví dụ LOINC codes:
- 2339-0: Glucose [Mass/volume] in Blood
- 4548-4: Hemoglobin A1c/Hemoglobin.total in Blood (HbA1c)
- 2160-0: Creatinine [Mass/volume] in Serum or Plasma
- 718-7: Hemoglobin [Mass/volume] in Blood
"""
response = requests.get(
f"{FHIR_SERVER}/Observation",
params={
"patient": patient_id,
"code": code,
"_sort": "-date", # Mới nhất trước
"_count": 100,
},
headers={"Accept": "application/fhir+json"}
)
response.raise_for_status()
bundle = response.json()
observations = []
for entry in bundle.get("entry", []):
obs = entry["resource"]
observations.append({
"date": obs.get("effectiveDateTime"),
"value": obs.get("valueQuantity", {}).get("value"),
"unit": obs.get("valueQuantity", {}).get("unit"),
"status": obs.get("status"),
})
return observations
# Ví dụ: lấy lịch sử HbA1c của bệnh nhân để dự đoán tiểu đường
hba1c_history = get_patient_observations("patient-123", "4548-4")
# [{"date": "2024-01-15", "value": 8.2, "unit": "%", "status": "final"},
# {"date": "2023-10-10", "value": 7.9, "unit": "%", "status": "final"}, ...]
def build_patient_ai_features(patient_id: str) -> dict:
"""Tổng hợp features từ FHIR cho ML model"""
# Latest vitals
bp_readings = get_patient_observations(patient_id, "55284-4") # Blood pressure
glucose = get_patient_observations(patient_id, "2339-0")
hba1c = get_patient_observations(patient_id, "4548-4")
creatinine = get_patient_observations(patient_id, "2160-0")
return {
"latest_systolic": bp_readings[0]["value"] if bp_readings else None,
"latest_glucose": glucose[0]["value"] if glucose else None,
"latest_hba1c": hba1c[0]["value"] if hba1c else None,
"hba1c_trend": (
hba1c[0]["value"] - hba1c[-1]["value"]
if len(hba1c) >= 2 else 0
),
"latest_creatinine": creatinine[0]["value"] if creatinine else None,
# ... thêm nhiều features khác
}
3. De-identification — Anonymization of medical data
3.1. Why de-identification is harder than you think
Common mistake: "Just remove the name and date of birth and you're done."
Fact: Re-identification attack can re-identify from:
- Age + gender + zip code combination → 87% of Americans can be identified (Sweeney, 2000)
- Rare diagnosis combination: "50 years old + salivary gland cancer + minor stroke" → almost unique
- Wearable data patterns → recognize people from their gait
- Medical image metadata (date taken, device)
3.2. HIPAA Safe Harbor — 18 identifiers to remove
# HIPAA Safe Harbor method: xóa 18 loại thông tin nhận dạng
HIPAA_PHI_IDENTIFIERS = [
"name", # Tên
"geographic_data", # Địa chỉ (chỉ giữ 3 digits zip code)
"dates", # Ngày (chỉ giữ năm — QUAN TRỌNG!)
"phone_numbers", # Số điện thoại
"fax_numbers", # Fax
"email_addresses", # Email
"ssn", # Số an sinh xã hội
"medical_record_numbers", # Mã hồ sơ bệnh án
"health_plan_numbers", # Số bảo hiểm
"account_numbers", # Số tài khoản
"certificate_numbers", # Số chứng chỉ
"vehicle_identifiers", # Biển số xe
"device_identifiers", # Serial number thiết bị
"urls", # Website URLs
"ip_addresses", # Địa chỉ IP
"biometric_identifiers", # Vân tay, giọng nói, mống mắt
"full_face_photos", # Ảnh mặt
"unique_identifiers", # Bất kỳ ID duy nhất nào khác
]
import re
from dateutil import parser as dateparser
def deidentify_clinical_note(text: str) -> str:
"""
De-identify text sử dụng NLP + regex
Đây là ví dụ đơn giản — thực tế dùng tools như:
- Microsoft Presidio (open source)
- AWS Comprehend Medical + AWS Macie
- Google Cloud Healthcare DLP API
"""
# Xóa các patterns thông thường với regex
patterns = {
# Tên (giả sử đại tự hóa)
r'\b[A-Z][a-z]+ [A-Z][a-z]+\b': '[NAME]',
# Số điện thoại VN
r'\b(0[3-9]\d{8}|\+84[3-9]\d{8})\b': '[PHONE]',
# Email
r'\b[\w.+-]+@[\w-]+\.[a-z]{2,}\b': '[EMAIL]',
# Ngày tháng
r'\b\d{1,2}[/-]\d{1,2}[/-]\d{2,4}\b': '[DATE]',
# Tuổi chính xác (giữ age range thay vì exact age)
r'\b(\d+)\s*(tuổi|years?\s*old)\b': lambda m: f"[AGE_{(int(m.group(1))//10)*10}s]",
}
deidentified = text
for pattern, replacement in patterns.items():
if callable(replacement):
deidentified = re.sub(pattern, replacement, deidentified, flags=re.IGNORECASE)
else:
deidentified = re.sub(pattern, replacement, deidentified, flags=re.IGNORECASE)
return deidentified
# Input: "Bệnh nhân Nguyễn Văn A, 65 tuổi, nhập viện ngày 15/01/2024. SĐT: 0912345678"
# Output: "Bệnh nhân [NAME], [AGE_60s], nhập viện ngày [DATE]. SĐT: [PHONE]"
3.3. De-identification for DICOM images
import pydicom
from pydicom.uid import generate_uid
def deidentify_dicom(input_path: str, output_path: str, patient_pseudoid: str):
"""
De-identify DICOM file theo DICOM Supplement 142 (Clinical Trial De-identification Profiles)
"""
dcm = pydicom.dcmread(input_path)
# Tags cần xóa hoặc thay thế (danh sách đầy đủ hơn nhiều trong thực tế)
tags_to_blank = [
(0x0010, 0x0010), # PatientName
(0x0010, 0x0030), # PatientBirthDate
(0x0010, 0x0040), # PatientSex — có thể giữ cho bias evaluation
(0x0008, 0x0080), # InstitutionName
(0x0008, 0x1070), # OperatorsName
(0x0008, 0x1048), # PhysiciansOfRecord
(0x0010, 0x1000), # OtherPatientIDs
]
for tag in tags_to_blank:
if tag in dcm:
del dcm[tag]
# Thay PatientID bằng pseudonymized ID
dcm.PatientID = patient_pseudoid
# Chỉ giữ năm, xóa tháng ngày
if hasattr(dcm, 'StudyDate'):
dcm.StudyDate = dcm.StudyDate[:4] + "0101" # Keep year only
# Tạo UIDs mới (break linkage giữa các study)
dcm.StudyInstanceUID = generate_uid()
dcm.SeriesInstanceUID = generate_uid()
dcm.SOPInstanceUID = generate_uid()
# QUAN TRỌNG: Xóa private tags (manufacturer-specific metadata)
# Có thể chứa thông tin nhận dạng không chuẩn
dcm.remove_private_tags()
dcm.save_as(output_path, write_like_original=False)
print(f"De-identified DICOM saved to {output_path}")
4. EHR Systems — Understand the actual data environment
4.1. Popular EHR systems
| HIS/EHR | Market | Export Format |
|---|---|---|
| Epic | 33% of US hospitals | FHIR R4, HL7 v2 |
| Oracle Health (Cerner) | 25% of US hospitals | FHIR R4, CCDs |
| VNPT-HIS | Many Vietnamese hospitals | HL7 v2, CSV |
| Bach Mai HIS Hospital | Bach Mai, VN | Proprietary + HL7 |
| Medisoft | Small clinic | CSV, custom |
4.2. Practical problems when working with EHR data
# Những điều bạn sẽ gặp khi nhận EHR data thực tế:
problems = {
"missing_data": """
Lab result thiếu: bệnh nhân không làm xét nghiệm hoặc chưa import
Vitals gap: không điều dưỡng nhập giữa các ca
MCAR vs MAR vs MNAR — missing pattern quan trọng hơn missing rate!
MCAR = Missing Completely At Random (có thể impute an toàn)
MAR = Missing At Random (cần cẩn thận)
MNAR = Missing Not At Random (bias nghiêm trọng — BN nặng thường miss labs)
""",
"inconsistent_units": """
Glucose: một chỗ ghi mmol/L, chỗ khác mg/dL (nhân/chia 18.016!)
Weight: kg vs lbs
Creatinine: mg/dL vs μmol/L
→ Cần unit normalization pipeline nghiêm túc
""",
"free_text_mess": """
"BN khó thở 3 ngày, SOB on exertion, DOE x 3d"
"Pt c/o dyspnea 3 days, worsening with activity"
→ Cùng một ý, viết 100 cách khác nhau
→ NLP cần xử lý: abbreviations, typos, bilingual mixing (VN+EN)
""",
"temporal_issues": """
Timezone: bệnh viện dùng server time khác nhau
Backdating: bác sĩ nhập lại sau 3 ngày
Overlapping encounters: emergency + inpatient cùng lúc
""",
"icd_coding_errors": """
ICD-10 được code bởi billing coder, không phải bác sĩ
Over-coding (để maximize reimbursement)
Under-coding (để tránh audit)
→ Training label quality phụ thuộc vào billing accuracy!
"""
}
5. Build a basic Medical Data Pipeline
from pathlib import Path
import pandas as pd
import pydicom
import torch
from torch.utils.data import Dataset
from torchvision import transforms
class MedicalDataPipeline:
"""
Pipeline đơn giản: DICOM files → ML-ready tensors
"""
def __init__(self, dicom_dir: str, labels_csv: str):
self.dicom_dir = Path(dicom_dir)
self.labels = pd.read_csv(labels_csv)
def load_and_preprocess(self, patient_id: str) -> np.ndarray:
"""Load DICOM, apply windowing, normalize cho model input"""
dcm_path = self.dicom_dir / f"{patient_id}.dcm"
dcm = pydicom.dcmread(dcm_path)
# Convert to HU
pixels = self._to_hu(dcm)
# Apply lung window (ví dụ cho chest X-ray AI)
pixels = apply_windowing(pixels, window_center=-600, window_width=1500)
# Resize to model input (224x224 cho ResNet)
from PIL import Image
img = Image.fromarray((pixels * 255).astype(np.uint8))
img = img.convert("RGB") # Grayscale → 3 channels (cho ImageNet pretrained)
img = img.resize((224, 224), Image.BILINEAR)
return np.array(img)
def _to_hu(self, dcm) -> np.ndarray:
pixels = dcm.pixel_array.astype(np.float32)
slope = float(getattr(dcm, 'RescaleSlope', 1.0))
intercept = float(getattr(dcm, 'RescaleIntercept', 0.0))
return pixels * slope + intercept
class ChestXrayDataset(Dataset):
def __init__(self, df: pd.DataFrame, dicom_dir: str, transform=None, label_cols: list = None):
self.df = df.reset_index(drop=True)
self.pipeline = MedicalDataPipeline(dicom_dir, None)
self.transform = transform
self.label_cols = label_cols or []
def __len__(self):
return len(self.df)
def __getitem__(self, idx):
row = self.df.iloc[idx]
image = self.pipeline.load_and_preprocess(row["patient_id"])
if self.transform:
image = self.transform(image)
labels = torch.FloatTensor([row[col] for col in self.label_cols])
return image, labels
6. Summary & Exercises
After this article, you should understand:
- ✅ DICOM structure: tags, pixel data, Hounsfield Units
- ✅ Windowing: why is it needed and how to implement it
- ✅ FHIR R4: Main resources, REST API
- ✅ De-identification: HIPAA Safe Harbor, 18 identifiers
- ✅ Practical issues of EHR data
Lesson 3 will go into the complete preprocessing pipeline for medical images — proper normalization, augmentation, and why flip vertical chest X-rays are clinically wrong.
Exercises
-
Download a DICOM sample file from TCIA (The Cancer Imaging Archive). Open with pydicom and print out the 10 most important tags.
-
Write functions
apply_windowingwith parameter validation — throws exception if window_width ≤ 0 or image is not HU range (-1024 to 3000). -
Access FHIR public test server (
https://r4.smarthealthit.org). Use Python requests to get the patient's list of Conditions87a339d0-8cae-418e-89c7-8651e6aab3c6. What are ICD codes?
2. Architecture & Principles
Core Architecture
# Example implementation
import torch
import torch.nn as nn
class ExampleModel(nn.Module):
def __init__(self, input_dim, output_dim):
super().__init__()
self.net = nn.Sequential(
nn.Linear(input_dim, 256),
nn.ReLU(),
nn.Dropout(0.2),
nn.Linear(256, 128),
nn.ReLU(),
nn.Linear(128, output_dim),
)
def forward(self, x):
return self.net(x)
3. Practice
Setup
pip install torch transformers datasets
Training Pipeline
# Training loop
model = ExampleModel(input_dim=768, output_dim=10)
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-4)
criterion = nn.CrossEntropyLoss()
for epoch in range(10):
for batch in train_loader:
optimizer.zero_grad()
outputs = model(batch["input"])
loss = criterion(outputs, batch["label"])
loss.backward()
optimizer.step()
4. Best Practices
| Aspect | Recommendation |
|---|---|
| Data | Quality over quantity |
| Model | Start simple, scale up |
| Training | Monitor loss curves |
| Evaluation | Use appropriate metrics |
Summary
| Concepts | Key Takeaway |
|---|---|
| Architecture | Suitable for the problem |
| Training | Careful hyperparameter tuning |
| Evaluation | Multiple metrics |