Chuyển đến nội dung chính

Lesson 2: Medical Data — DICOM, HL7 FHIR & Privacy

Medical data formats: DICOM images, HL7 FHIR. EHR systems. De-identification. HIPAA compliance. Data pipeline for medical AI.

🧠 AI & ML — Lesson 1 Lesson 2: Medical Data — DICOM, HL7 FHIR & Privacy

AI in Health & Healthcare: Real Battle Applications

Part 1: Medical AI Platform & Medical Data

xdev.asia

Healthcare data is unlike any data you have ever worked with. This article explains why — and how to work with it properly from day one.


1. DICOM — Medical imaging standard

DICOM (Digital Imaging and Communications in Medicine) is an international standard for storing and transmitting medical images. Launched in 1993, today all hospitals around the world use DICOM.

1.1. DICOM file structure

DICOM isn't just an image — it's image + metadata:

File DICOM (.dcm)
├── File Meta Information
│   ├── Transfer Syntax UID   (cách encode dữ liệu pixel)
│   └── SOP Class UID         (loại hình ảnh: CT, MRI, X-ray...)
│
└── Data Set (hàng trăm tags)
    ├── Patient Info
    │   ├── (0010,0010) PatientName     = "Nguyen^Van^A"
    │   ├── (0010,0020) PatientID       = "BN-2024-001"
    │   ├── (0010,0030) PatientBirthDate = "19650315"
    │   └── (0010,0040) PatientSex      = "M"
    │
    ├── Study Info
    │   ├── (0020,000D) StudyInstanceUID
    │   ├── (0008,0020) StudyDate       = "20240115"
    │   └── (0008,1030) StudyDescription = "CHEST AP"
    │
    ├── Image Info
    │   ├── (0028,0010) Rows            = 2048
    │   ├── (0028,0011) Columns         = 2048
    │   ├── (0028,0030) PixelSpacing    = [0.175, 0.175] mm/pixel
    │   ├── (0028,1050) WindowCenter    = 40
    │   ├── (0028,1051) WindowWidth     = 400
    │   └── (0028,0103) PixelRepresentation = 1 (signed integers)
    │
    └── Pixel Data
        └── (7FE0,0010) PixelData = [raw 12-bit pixel values...]

1.2. Read DICOM with pydicom

import pydicom
import numpy as np
import matplotlib.pyplot as plt

# Đọc file DICOM
dcm = pydicom.dcmread("chest_xray.dcm")

# Xem metadata
print(f"Patient: {dcm.PatientName}")
print(f"Study Date: {dcm.StudyDate}")
print(f"Modality: {dcm.Modality}")       # CR (X-ray), CT, MR, US...
print(f"Image size: {dcm.Rows} x {dcm.Columns}")
print(f"Pixel spacing: {dcm.PixelSpacing} mm/pixel")

# Lấy pixel array (giá trị thô — chưa phải HU units!)
raw_pixels = dcm.pixel_array
print(f"Pixel shape: {raw_pixels.shape}")
print(f"Pixel dtype: {raw_pixels.dtype}")  # Thường là int16 cho CT
print(f"Min/Max values: {raw_pixels.min()} / {raw_pixels.max()}")

# Chuyển sang Hounsfield Units (HU) cho CT
# Công thức: HU = pixel_value * RescaleSlope + RescaleIntercept
def to_hounsfield_units(dcm_file):
    pixels = dcm_file.pixel_array.astype(np.float32)
    slope = float(dcm_file.RescaleSlope) if hasattr(dcm_file, 'RescaleSlope') else 1.0
    intercept = float(dcm_file.RescaleIntercept) if hasattr(dcm_file, 'RescaleIntercept') else 0.0
    return pixels * slope + intercept

hu_image = to_hounsfield_units(dcm)

# HU reference values:
# Air: -1000 HU | Lung: -500 HU | Fat: -100 HU
# Water/Soft tissue: 0-80 HU | Blood: 40-80 HU | Bone: 400-1000 HU

1.3. Windowing — The core of Medical Image Display

The screen only displays 256 gray levels, but the CT has ~4000 HU levels. Windowing select the HU range to view:

def apply_windowing(image_hu: np.ndarray, window_center: int, window_width: int) -> np.ndarray:
    """
    Window Center (WC) và Window Width (WW) cho các mô khác nhau:
    
    Phổi:       WC=-600, WW=1500  (thấy cấu trúc phổi chi tiết)
    Ổ bụng:    WC=40,   WW=400   (thấy gan, lách, thận)
    Não:        WC=40,   WW=80    (thấy hematoma)
    Xương:      WC=400,  WW=1800  (thấy gãy xương)
    """
    lower = window_center - window_width / 2
    upper = window_center + window_width / 2

    # Clip và normalize về [0, 1]
    windowed = np.clip(image_hu, lower, upper)
    normalized = (windowed - lower) / (upper - lower)
    return normalized

# Hiển thị cùng một CT scan với 3 window settings khác nhau
fig, axes = plt.subplots(1, 3, figsize=(15, 5))

windows = [
    ("Lung Window", -600, 1500),
    ("Soft Tissue", 40, 400),
    ("Bone Window", 400, 1800),
]

for ax, (title, wc, ww) in zip(axes, windows):
    windowed = apply_windowing(hu_image, wc, ww)
    ax.imshow(windowed, cmap='gray')
    ax.set_title(f"{title}\nWC={wc}, WW={ww}")
    ax.axis('off')

plt.tight_layout()
plt.savefig("windowing_comparison.png", dpi=150)

2. HL7 FHIR — Clinical data exchange standard

HL7 FHIR (Fast Healthcare Interoperability Resources) is the most modern standard for data exchange between healthcare systems. Not images — this is structured clinical data.

2.1. Core Resources

FHIR organizes data into "Resources" — each entity type has its own schema:

FHIR Resources (chỉ các resource quan trọng nhất cho AI)
│
├── Patient          — thông tin bệnh nhân
├── Observation      — vital signs, lab results, measurements
├── Condition        — diagnoses, problems (mã ICD-10/11)
├── MedicationRequest— đơn thuốc
├── Procedure        — phẫu thuật, can thiệp (mã CPT/ICD)
├── DiagnosticReport — báo cáo xét nghiệm, hình ảnh
├── Encounter        — lần khám/nhập viện
└── AllergyIntolerance — dị ứng

2.2. Works with FHIR API

import requests
from datetime import datetime, timedelta

# FHIR R4 RESTful API
FHIR_SERVER = "https://r4.smarthealthit.org"  # Public test server

def get_patient_observations(patient_id: str, code: str) -> list:
    """
    Lấy lab results của bệnh nhân
    code: LOINC code — chuẩn mã cho lab tests
    
    Ví dụ LOINC codes:
    - 2339-0: Glucose [Mass/volume] in Blood
    - 4548-4: Hemoglobin A1c/Hemoglobin.total in Blood (HbA1c)
    - 2160-0: Creatinine [Mass/volume] in Serum or Plasma
    - 718-7:  Hemoglobin [Mass/volume] in Blood
    """
    response = requests.get(
        f"{FHIR_SERVER}/Observation",
        params={
            "patient": patient_id,
            "code": code,
            "_sort": "-date",      # Mới nhất trước
            "_count": 100,
        },
        headers={"Accept": "application/fhir+json"}
    )
    response.raise_for_status()
    bundle = response.json()

    observations = []
    for entry in bundle.get("entry", []):
        obs = entry["resource"]
        observations.append({
            "date": obs.get("effectiveDateTime"),
            "value": obs.get("valueQuantity", {}).get("value"),
            "unit": obs.get("valueQuantity", {}).get("unit"),
            "status": obs.get("status"),
        })

    return observations

# Ví dụ: lấy lịch sử HbA1c của bệnh nhân để dự đoán tiểu đường
hba1c_history = get_patient_observations("patient-123", "4548-4")
# [{"date": "2024-01-15", "value": 8.2, "unit": "%", "status": "final"},
#  {"date": "2023-10-10", "value": 7.9, "unit": "%", "status": "final"}, ...]

def build_patient_ai_features(patient_id: str) -> dict:
    """Tổng hợp features từ FHIR cho ML model"""
    # Latest vitals
    bp_readings = get_patient_observations(patient_id, "55284-4")  # Blood pressure
    glucose = get_patient_observations(patient_id, "2339-0")
    hba1c = get_patient_observations(patient_id, "4548-4")
    creatinine = get_patient_observations(patient_id, "2160-0")

    return {
        "latest_systolic": bp_readings[0]["value"] if bp_readings else None,
        "latest_glucose": glucose[0]["value"] if glucose else None,
        "latest_hba1c": hba1c[0]["value"] if hba1c else None,
        "hba1c_trend": (
            hba1c[0]["value"] - hba1c[-1]["value"]
            if len(hba1c) >= 2 else 0
        ),
        "latest_creatinine": creatinine[0]["value"] if creatinine else None,
        # ... thêm nhiều features khác
    }

3. De-identification — Anonymization of medical data

3.1. Why de-identification is harder than you think

Common mistake: "Just remove the name and date of birth and you're done."

Fact: Re-identification attack can re-identify from:

  • Age + gender + zip code combination → 87% of Americans can be identified (Sweeney, 2000)
  • Rare diagnosis combination: "50 years old + salivary gland cancer + minor stroke" → almost unique
  • Wearable data patterns → recognize people from their gait
  • Medical image metadata (date taken, device)

3.2. HIPAA Safe Harbor — 18 identifiers to remove

# HIPAA Safe Harbor method: xóa 18 loại thông tin nhận dạng
HIPAA_PHI_IDENTIFIERS = [
    "name",                   # Tên
    "geographic_data",        # Địa chỉ (chỉ giữ 3 digits zip code)
    "dates",                  # Ngày (chỉ giữ năm — QUAN TRỌNG!)
    "phone_numbers",          # Số điện thoại
    "fax_numbers",            # Fax
    "email_addresses",        # Email
    "ssn",                    # Số an sinh xã hội
    "medical_record_numbers", # Mã hồ sơ bệnh án
    "health_plan_numbers",    # Số bảo hiểm
    "account_numbers",        # Số tài khoản
    "certificate_numbers",    # Số chứng chỉ
    "vehicle_identifiers",    # Biển số xe
    "device_identifiers",     # Serial number thiết bị
    "urls",                   # Website URLs
    "ip_addresses",           # Địa chỉ IP
    "biometric_identifiers",  # Vân tay, giọng nói, mống mắt
    "full_face_photos",       # Ảnh mặt
    "unique_identifiers",     # Bất kỳ ID duy nhất nào khác
]

import re
from dateutil import parser as dateparser

def deidentify_clinical_note(text: str) -> str:
    """
    De-identify text sử dụng NLP + regex
    Đây là ví dụ đơn giản — thực tế dùng tools như:
    - Microsoft Presidio (open source)
    - AWS Comprehend Medical + AWS Macie
    - Google Cloud Healthcare DLP API
    """
    # Xóa các patterns thông thường với regex
    patterns = {
        # Tên (giả sử đại tự hóa)
        r'\b[A-Z][a-z]+ [A-Z][a-z]+\b': '[NAME]',
        # Số điện thoại VN
        r'\b(0[3-9]\d{8}|\+84[3-9]\d{8})\b': '[PHONE]',
        # Email
        r'\b[\w.+-]+@[\w-]+\.[a-z]{2,}\b': '[EMAIL]',
        # Ngày tháng
        r'\b\d{1,2}[/-]\d{1,2}[/-]\d{2,4}\b': '[DATE]',
        # Tuổi chính xác (giữ age range thay vì exact age)
        r'\b(\d+)\s*(tuổi|years?\s*old)\b': lambda m: f"[AGE_{(int(m.group(1))//10)*10}s]",
    }

    deidentified = text
    for pattern, replacement in patterns.items():
        if callable(replacement):
            deidentified = re.sub(pattern, replacement, deidentified, flags=re.IGNORECASE)
        else:
            deidentified = re.sub(pattern, replacement, deidentified, flags=re.IGNORECASE)

    return deidentified

# Input:  "Bệnh nhân Nguyễn Văn A, 65 tuổi, nhập viện ngày 15/01/2024. SĐT: 0912345678"
# Output: "Bệnh nhân [NAME], [AGE_60s], nhập viện ngày [DATE]. SĐT: [PHONE]"

3.3. De-identification for DICOM images

import pydicom
from pydicom.uid import generate_uid

def deidentify_dicom(input_path: str, output_path: str, patient_pseudoid: str):
    """
    De-identify DICOM file theo DICOM Supplement 142 (Clinical Trial De-identification Profiles)
    """
    dcm = pydicom.dcmread(input_path)

    # Tags cần xóa hoặc thay thế (danh sách đầy đủ hơn nhiều trong thực tế)
    tags_to_blank = [
        (0x0010, 0x0010),   # PatientName
        (0x0010, 0x0030),   # PatientBirthDate
        (0x0010, 0x0040),   # PatientSex — có thể giữ cho bias evaluation
        (0x0008, 0x0080),   # InstitutionName
        (0x0008, 0x1070),   # OperatorsName
        (0x0008, 0x1048),   # PhysiciansOfRecord
        (0x0010, 0x1000),   # OtherPatientIDs
    ]

    for tag in tags_to_blank:
        if tag in dcm:
            del dcm[tag]

    # Thay PatientID bằng pseudonymized ID
    dcm.PatientID = patient_pseudoid

    # Chỉ giữ năm, xóa tháng ngày
    if hasattr(dcm, 'StudyDate'):
        dcm.StudyDate = dcm.StudyDate[:4] + "0101"  # Keep year only

    # Tạo UIDs mới (break linkage giữa các study)
    dcm.StudyInstanceUID = generate_uid()
    dcm.SeriesInstanceUID = generate_uid()
    dcm.SOPInstanceUID = generate_uid()

    # QUAN TRỌNG: Xóa private tags (manufacturer-specific metadata)
    # Có thể chứa thông tin nhận dạng không chuẩn
    dcm.remove_private_tags()

    dcm.save_as(output_path, write_like_original=False)
    print(f"De-identified DICOM saved to {output_path}")

4. EHR Systems — Understand the actual data environment

4.1. Popular EHR systems

HIS/EHRMarketExport Format
Epic33% of US hospitalsFHIR R4, HL7 v2
Oracle Health (Cerner)25% of US hospitalsFHIR R4, CCDs
VNPT-HISMany Vietnamese hospitalsHL7 v2, CSV
Bach Mai HIS HospitalBach Mai, VNProprietary + HL7
MedisoftSmall clinicCSV, custom

4.2. Practical problems when working with EHR data

# Những điều bạn sẽ gặp khi nhận EHR data thực tế:

problems = {
    "missing_data": """
        Lab result thiếu: bệnh nhân không làm xét nghiệm hoặc chưa import
        Vitals gap: không điều dưỡng nhập giữa các ca
        MCAR vs MAR vs MNAR — missing pattern quan trọng hơn missing rate!
        
        MCAR = Missing Completely At Random (có thể impute an toàn)
        MAR = Missing At Random (cần cẩn thận)
        MNAR = Missing Not At Random (bias nghiêm trọng — BN nặng thường miss labs)
    """,

    "inconsistent_units": """
        Glucose: một chỗ ghi mmol/L, chỗ khác mg/dL (nhân/chia 18.016!)
        Weight: kg vs lbs
        Creatinine: mg/dL vs μmol/L
        → Cần unit normalization pipeline nghiêm túc
    """,

    "free_text_mess": """
        "BN khó thở 3 ngày, SOB on exertion, DOE x 3d"
        "Pt c/o dyspnea 3 days, worsening with activity"
        → Cùng một ý, viết 100 cách khác nhau
        → NLP cần xử lý: abbreviations, typos, bilingual mixing (VN+EN)
    """,

    "temporal_issues": """
        Timezone: bệnh viện dùng server time khác nhau
        Backdating: bác sĩ nhập lại sau 3 ngày
        Overlapping encounters: emergency + inpatient cùng lúc
    """,

    "icd_coding_errors": """
        ICD-10 được code bởi billing coder, không phải bác sĩ
        Over-coding (để maximize reimbursement)
        Under-coding (để tránh audit)
        → Training label quality phụ thuộc vào billing accuracy!
    """
}

5. Build a basic Medical Data Pipeline

from pathlib import Path
import pandas as pd
import pydicom
import torch
from torch.utils.data import Dataset
from torchvision import transforms

class MedicalDataPipeline:
    """
    Pipeline đơn giản: DICOM files → ML-ready tensors
    """
    def __init__(self, dicom_dir: str, labels_csv: str):
        self.dicom_dir = Path(dicom_dir)
        self.labels = pd.read_csv(labels_csv)

    def load_and_preprocess(self, patient_id: str) -> np.ndarray:
        """Load DICOM, apply windowing, normalize cho model input"""
        dcm_path = self.dicom_dir / f"{patient_id}.dcm"
        dcm = pydicom.dcmread(dcm_path)

        # Convert to HU
        pixels = self._to_hu(dcm)

        # Apply lung window (ví dụ cho chest X-ray AI)
        pixels = apply_windowing(pixels, window_center=-600, window_width=1500)

        # Resize to model input (224x224 cho ResNet)
        from PIL import Image
        img = Image.fromarray((pixels * 255).astype(np.uint8))
        img = img.convert("RGB")  # Grayscale → 3 channels (cho ImageNet pretrained)
        img = img.resize((224, 224), Image.BILINEAR)

        return np.array(img)

    def _to_hu(self, dcm) -> np.ndarray:
        pixels = dcm.pixel_array.astype(np.float32)
        slope = float(getattr(dcm, 'RescaleSlope', 1.0))
        intercept = float(getattr(dcm, 'RescaleIntercept', 0.0))
        return pixels * slope + intercept

class ChestXrayDataset(Dataset):
    def __init__(self, df: pd.DataFrame, dicom_dir: str, transform=None, label_cols: list = None):
        self.df = df.reset_index(drop=True)
        self.pipeline = MedicalDataPipeline(dicom_dir, None)
        self.transform = transform
        self.label_cols = label_cols or []

    def __len__(self):
        return len(self.df)

    def __getitem__(self, idx):
        row = self.df.iloc[idx]
        image = self.pipeline.load_and_preprocess(row["patient_id"])

        if self.transform:
            image = self.transform(image)

        labels = torch.FloatTensor([row[col] for col in self.label_cols])
        return image, labels

6. Summary & Exercises

After this article, you should understand:

  • ✅ DICOM structure: tags, pixel data, Hounsfield Units
  • ✅ Windowing: why is it needed and how to implement it
  • ✅ FHIR R4: Main resources, REST API
  • ✅ De-identification: HIPAA Safe Harbor, 18 identifiers
  • ✅ Practical issues of EHR data

Lesson 3 will go into the complete preprocessing pipeline for medical images — proper normalization, augmentation, and why flip vertical chest X-rays are clinically wrong.


Exercises

  1. Download a DICOM sample file from TCIA (The Cancer Imaging Archive). Open with pydicom and print out the 10 most important tags.

  2. Write functions apply_windowing with parameter validation — throws exception if window_width ≤ 0 or image is not HU range (-1024 to 3000).

  3. Access FHIR public test server (https://r4.smarthealthit.org). Use Python requests to get the patient's list of Conditions 87a339d0-8cae-418e-89c7-8651e6aab3c6. What are ICD codes?

2. Architecture & Principles

Core Architecture

# Example implementation
import torch
import torch.nn as nn

class ExampleModel(nn.Module):
    def __init__(self, input_dim, output_dim):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(input_dim, 256),
            nn.ReLU(),
            nn.Dropout(0.2),
            nn.Linear(256, 128),
            nn.ReLU(),
            nn.Linear(128, output_dim),
        )
    
    def forward(self, x):
        return self.net(x)

3. Practice

Setup

pip install torch transformers datasets

Training Pipeline

# Training loop
model = ExampleModel(input_dim=768, output_dim=10)
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-4)
criterion = nn.CrossEntropyLoss()

for epoch in range(10):
    for batch in train_loader:
        optimizer.zero_grad()
        outputs = model(batch["input"])
        loss = criterion(outputs, batch["label"])
        loss.backward()
        optimizer.step()

4. Best Practices

AspectRecommendation
DataQuality over quantity
ModelStart simple, scale up
TrainingMonitor loss curves
EvaluationUse appropriate metrics

Summary

ConceptsKey Takeaway
ArchitectureSuitable for the problem
TrainingCareful hyperparameter tuning
EvaluationMultiple metrics