
Introduction
Global health data is a huge treasure trove — billions of records from hospitals, clinics, health insurance. But each system is stored in a separate form: Hospital A's HIS (Hospital Information System) is completely different from hospital B, ICD-10 disease codes are not consistent with SNOMED CT, prescriptions are stored by trade name instead of international active ingredients.
OHDSI (Observational Health Data Sciences and Informatics — pronounced "Odyssey") was created to solve this problem.
1. What is OHDSI?
1.1 Definition
OHDSI is an international, open-source research program that aims to:
- Standardize observational health data into a common data model
- Develop reliable, reproducible analytical methods
- Allow multicenter studies WITHOUT sharing patient data
1.2 History
2008: OMOP (Observational Medical Outcomes Partnership)
→ Dự án FDA nghiên cứu tác dụng phụ thuốc trên dữ liệu thực tế
→ Phát triển Common Data Model (CDM)
2013: OMOP kết thúc → OHDSI ra đời
→ Kế thừa OMOP CDM, mở rộng thành cộng đồng mã nguồn mở
→ Mục tiêu: evidence-based medicine trên quy mô toàn cầu
2024: 800+ tổ chức, 100+ quốc gia
→ 1+ tỷ bản ghi bệnh nhân được chuẩn hóa
→ Hàng nghìn nghiên cứu được publish
2026: OHDSI tiếp tục mở rộng
→ OMOP CDM v5.4, v6.0 đang phát triển
→ Tích hợp AI/ML, genomics, wearable data
1.3 Three pillars of OHDSI
┌─────────────────────────────────────────────────────────┐
│ OHDSI Mission │
│ │
│ ┌─────────────┐ ┌──────────────┐ ┌──────────────────┐ │
│ │ Open │ │ Open │ │ Open │ │
│ │ Science │ │ Source │ │ Community │ │
│ │ │ │ │ │ │ │
│ │ Transparent │ │ Free tools │ │ Collaborative │ │
│ │ Reproducible│ │ Peer-reviewed│ │ 800+ orgs │ │
│ │ Published │ │ GitHub │ │ Global network │ │
│ └─────────────┘ └──────────────┘ └──────────────────┘ │
└─────────────────────────────────────────────────────────┘
2. OHDSI problem solved
2.1 Fragmentation — Data fragmentation
Bệnh viện A (HIS: eHospital) Bệnh viện B (HIS: Telehealth)
┌─────────────────────────┐ ┌─────────────────────────┐
│ Patient: MA_BN_001 │ │ Patient: BN-2024-00123 │
│ Diagnosis: I10 (ICD-10) │ │ Diagnosis: 401.1 (ICD-9)│
│ Drug: Amlor 5mg │ │ Drug: Amlodipine 5mg │
│ Lab: Glucose: 126 mg/dL │ │ Lab: Đường huyết: 7.0 │
│ Date: 01/03/2024 │ │ Date: 2024-03-01 │
└─────────────────────────┘ └─────────────────────────┘
→ Cùng 1 bệnh nhân, cùng 1 bệnh (tăng huyết áp), cùng 1 thuốc
→ Nhưng KHÔNG THỂ truy vấn chung vì format khác nhau hoàn toàn
2.2 Solution: OMOP CDM
ETL (Extract - Transform - Load)
Bệnh viện A ─────┐ ┌──────────────────────┐
├─── Transform ───→ │ OMOP CDM Database │
Bệnh viện B ─────┘ │ │
│ person_id: 12345 │
│ condition: 320128 │
│ (Essential HTN) │
│ drug: 1332419 │
│ (Amlodipine 5mg) │
│ measurement: 3004410 │
│ (Glucose 126 mg/dL)│
└──────────────────────┘
→ Cùng concept IDs cho cùng ý nghĩa y khoa
→ Cùng cấu trúc bảng → cùng SQL query
→ Có thể phân tích đa trung tâm
3. OHDSI Ecosystem Architecture
┌─────────────────────────────────────────────────────────────────┐
│ OHDSI Ecosystem │
│ │
│ ┌──── Data Standardization ─────────────────────────────────┐ │
│ │ │ │
│ │ [Athena] → Standardized Vocabularies │ │
│ │ [WhiteRabbit] → Scan source data │ │
│ │ [Rabbit-in-a-Hat] → Design ETL mapping │ │
│ │ [Usagi] → Map source codes → standard concepts │ │
│ │ │ │
│ └───────────────────────────────────────────────────────────┘ │
│ │ ETL │
│ ▼ │
│ ┌──── OMOP CDM Database ────────────────────────────────────┐ │
│ │ PostgreSQL / SQL Server / Oracle / Spark │ │
│ └───────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌──── Data Quality ────────────────────────────────────────┐ │
│ │ [ACHILLES] → Data characterization │ │
│ │ [Data Quality Dashboard]→ 1,500+ quality checks │ │
│ └───────────────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ ┌──── Analytics Platform ──────────────────────────────────┐ │
│ │ [WebAPI] → REST API backend (Spring Boot / Java) │ │
│ │ [ATLAS] → Web UI (JavaScript) cho phân tích │ │
│ │ [HADES] → R packages cho advanced analytics │ │
│ └───────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
3.1 Main tools
| Tools | Purpose | Language |
|---|---|---|
| Athena | Look up & download Standardized Vocabularies | Web app |
| WhiteRabbit | Scan source data, create profile report | Java |
| Rabbit-in-a-Hat | ETL mapping design (GUI) | Java |
| Usagi | Map source codes → OMOP concepts | Java |
| WebAPI | Backend REST API for ATLAS | Java/Spring Boot |
| ATLAS | Web-based analytics platform | JavaScript |
| ACHILLES | Data characterization & profiling | R |
| DQD | Data Quality Dashboard — 1,500+ checks | R |
| HADES | R packages for observational research | R |
4. End-to-End OHDSI Workflow
Step 1: Vocabulary Preparation
Athena → Download vocabularies (ICD-10, SNOMED, RxNorm, LOINC...)
│
Step 2: Source Data Profiling
WhiteRabbit → Scan source database → Generate scan report
│
Step 3: ETL Design
Rabbit-in-a-Hat → Design table & field mappings
Usagi → Map source codes → Standard Concepts
│
Step 4: ETL Execution
Custom ETL scripts (Python/SQL) → Load data into OMOP CDM
│
Step 5: Data Quality
ACHILLES → Characterize CDM data
DQD → Run 1,500+ quality checks
│
Step 6: Analytics
WebAPI + ATLAS → Cohort Definitions, Characterization,
Incidence Rates, Estimation, Prediction
HADES → Advanced R-based analytics
│
Step 7: Network Study
Package study → Distribute to sites → Collect aggregate results
5. OHDSI Network — Distributed Research
What's special about OHDSI: patient data NEVER leaves the site.
┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ Site A │ │ Site B │ │ Site C │
│ (500K pts) │ │ (1M pts) │ │ (200K pts) │
│ │ │ │ │ │
│ OMOP CDM │ │ OMOP CDM │ │ OMOP CDM │
│ WebAPI+ATLAS │ │ WebAPI+ATLAS │ │ WebAPI+ATLAS │
│ │ │ │ │ │
│ Run Study │ │ Run Study │ │ Run Study │
│ Package ───┐ │ │ Package ───┐ │ │ Package ───┐ │
└────────────│─┘ └────────────│─┘ └────────────│─┘
│ │ │
└───── Aggregate ───┴──── Results ───────┘
│
▼
┌─────────────────┐
│ Central Analysis│
│ (chỉ aggregate │
│ không có PII) │
└─────────────────┘
Why is it important?
- Comply with privacy regulations (HIPAA, GDPR, Vietnam Information Security Law)
- Each site retains full control over data
- Research on millions of multinational patients is still possible
6. Practical application
6.1 COVID-19 (OHDSI COVID-19 Study-a-thon)
Timeline:
- Tháng 3/2020: Đại dịch bùng phát
- Tháng 3/2020: OHDSI tổ chức Study-a-thon online
- 96 giờ: 300+ nhà nghiên cứu từ 30+ quốc gia
- Kết quả: Phân tích đặc điểm bệnh nhân COVID-19
trên 5+ triệu bệnh nhân từ nhiều quốc gia
→ Published trong PNAS (top-tier journal)
6.2 Drug Safety Surveillance
- Detect drug side effects on real data (RWD)
- Compare the drug group with the control group
- Example: Analyzing risk of myocarditis after vaccination
6.3 Application in Vietnam
- Standardize hospital HIS data into OMOP CDM
- Mapping ICD-10 Vietnam → SNOMED CT
- Epidemiological research on social insurance data
- Evaluate the effectiveness of multicenter treatment regimens
Summary
| Concept | Explanation |
|---|---|
| OHDSI | Global open source medical research community |
| OMOP CDM | Generic data model for observational medical data |
| Standardized Vocabularies | Standard vocabulary set (SNOMED, RxNorm, LOINC...) |
| Distributed Research | Multi-center analysis without data leaving the site |
| ETL | Source data conversion process → OMOP CDM |
Next article: OMOP Common Data Model — Structure, principles & Domain