1. Data Classification Framework for Healthcare

1.1. Why is it necessary to classify data?
Not all data needs the same level of protection. Data classification helps:
- Optimize security costs: Focus resources on the most important data
- Legal compliance: Apply correct controls according to regulatory requirements
- Reduce attack surface: Limit the scope of sensitive data
- Incident response: Prioritize handling when a breach occurs
1.2. Healthcare Data Classification Levels

| Level | Name | Example | Encryption | Access | Audit |
|---|---|---|---|---|---|
| 4 - RESTRICTED | Maximum restrictions | HIV/AIDS, mental health, genetics, addiction treatment, reproductive health | Required (AES-256) | Named individuals only | Full logging, real-time alerts |
| 3 - CONFIDENTIAL | Security | Medical records, tests, prescriptions, diagnostic imaging, health insurance | Required (AES-256) | Role-based (treating clinicians) | Full logging |
| 2 - INTERNAL | Internal | Appointment schedule, statistics (anonymous), medical staff, configuration | Recommended | Department-based | Standard logging |
| 1 - PUBLIC | Public | List of services, working hours, hospital contacts, health instructions | Not required | Public | Basic logging |
1.3. Data Classification in PostgreSQL Schema
-- Data classification metadata table
CREATE TABLE data_classification (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
schema_name VARCHAR(100) NOT NULL,
table_name VARCHAR(100) NOT NULL,
column_name VARCHAR(100) NOT NULL,
classification_level INTEGER NOT NULL CHECK (classification_level BETWEEN 1 AND 4),
classification_label VARCHAR(50) NOT NULL,
contains_phi BOOLEAN DEFAULT false,
encryption_required BOOLEAN DEFAULT false,
masking_rule VARCHAR(100),
retention_days INTEGER,
legal_basis TEXT,
created_at TIMESTAMPTZ DEFAULT NOW(),
updated_at TIMESTAMPTZ DEFAULT NOW()
);
-- Ví dụ classification cho patient table
INSERT INTO data_classification (schema_name, table_name, column_name,
classification_level, classification_label, contains_phi, encryption_required, masking_rule)
VALUES
('public', 'patients', 'id', 2, 'INTERNAL', false, false, NULL),
('public', 'patients', 'full_name', 3, 'CONFIDENTIAL', true, true, 'PARTIAL_MASK'),
('public', 'patients', 'date_of_birth', 3, 'CONFIDENTIAL', true, false, 'YEAR_ONLY'),
('public', 'patients', 'cccd_number', 3, 'CONFIDENTIAL', true, true, 'FULL_MASK'),
('public', 'patients', 'phone', 3, 'CONFIDENTIAL', true, true, 'PARTIAL_MASK'),
('public', 'patients', 'email', 3, 'CONFIDENTIAL', true, true, 'PARTIAL_MASK'),
('public', 'patients', 'address', 3, 'CONFIDENTIAL', true, true, 'CITY_ONLY'),
('public', 'patients', 'blood_type', 2, 'INTERNAL', false, false, NULL),
('public', 'patients', 'hiv_status', 4, 'RESTRICTED', true, true, 'FULL_MASK'),
('public', 'patients', 'insurance_number', 3, 'CONFIDENTIAL', true, true, 'PARTIAL_MASK');
2. Data Flow Mapping
2.1. PHI Data Flow in Microservices

2.2. Data Flow Documentation Template
| # | Data Element | Source | Destination | Transportation | Encryption | Classification |
|---|---|---|---|---|---|---|
| 1 | Patient Name | Portal | Patient Service | HTTPS/TLS 1.3 | In-transit + At-rest | L3 |
| 2 | Lab Results | Lab Instruments | Lab Service | HL7v2/MLLP over TLS | In-transit + At-rest | L3 |
| 3 | Diagnosis Code | Clinical Service | Billing Service | Kafka (SSL) | Application-level | L3 |
| 4 | HIV Status | Clinical Service | Clinical DB | JDBC/SSL | Column encryption | L4 |
| 5 | Audit Event | All Services | Audit Service | Kafka (SSL) | Event encryption | L2 |
| 6 | Appointments | Scheduling Service | Notification Service | Kafka (SSL) | In-transit | L2 |
3. Risk Assessment according to NIST SP 800-30
3.1. Risk Assessment Methodology

3.2. Threat Identification for Healthcare Microservices
| Threat Category | Threat | Threat Source |
|---|---|---|
| External | SQL Injection into Patient Service | Attacker |
| External | Ransomware encrypts database | Cybercriminal |
| External | MITM attack on API calls | Network attacks |
| External | Credential stuffing into Patient Portal | Bot networks |
| Internal | Unauthorized employee access to PHI | Insider |
| Internal | Database admin exports all patient data | Privileged users |
| Internal | Developer hardcode credentials | Negligent employees |
| Environmental | Database corruption due to hardware failure | Infrastructure |
| Environmental | Data loss due to natural disasters | Natural disasters |
| Supply Chain | Vulnerability in Quarkus dependency | Third-party |
3.3. Vulnerability Assessment
// Ví dụ: Checklist kiểm tra vulnerabilities trong Quarkus service
public class SecurityVulnerabilityChecklist {
// V1: SQL Injection - Sử dụng parameterized queries
// ❌ VULNERABLE
String badQuery = "SELECT * FROM patients WHERE name = '" + userInput + "'";
// ✅ SECURE
@NamedQuery(name = "Patient.findByName",
query = "SELECT p FROM Patient p WHERE p.name = :name")
List<Patient> findByName(@Param("name") String name);
// V2: Broken Authentication - Token validation
// ❌ VULNERABLE: Không verify token
String userId = jwt.getClaim("sub"); // Không verify expiration, issuer
// ✅ SECURE: Quarkus OIDC tự động verify
@Authenticated
@RolesAllowed("doctor")
public Response getPatient(UUID id) { ... }
// V3: Sensitive Data Exposure in Logs
// ❌ VULNERABLE
log.info("Patient created: " + patient.toString()); // Logs PHI!
// ✅ SECURE
log.info("Patient created: id={}", patient.getId()); // Only log ID
}
3.4. Risk Matrix

| Negligible (1) | Low (2) | Medium (3) | High (4) | Critical (5) | |
|---|---|---|---|---|---|
| Very High (5) | LOW | MEDIUM | HIGH | CRITICAL | CRITICAL |
| High (4) | LOW | MEDIUM | HIGH | HIGH | CRITICAL |
| Medium (3) | LOW | LOW | MEDIUM | HIGH | HIGH |
| Low (2) | LOW | LOW | LOW | MEDIUM | MEDIUM |
| Very Low (1) | LOW | LOW | LOW | LOW | MEDIUM |
4. Risk Register for Healthcare Microservices
4.1. Risk Register Template
| ID | Risk Description | Likelihood | Impact | Risk Level | Mitigation | Owner | Status |
|---|---|---|---|---|---|---|---|
| R001 | SQL Injection into Patient API | Medium (3) | Critical (5) | HIGH | Parameterized queries, input validation, WAF | Dev Team | Mitigation |
| R002 | Insider access PHI not authorized | High (4) | High (4) | HIGH | RBAC, RLS, Audit logging, DLP | Security Team | In Progress |
| R003 | Ransomware encrypts patient_db | Medium (3) | Critical (5) | HIGH | Immutable backups, network segmentation, EDR | Ops Team | Mitigation |
| R004 | Keycloak token theft | Medium (3) | High (4) | HIGH | Short-lived tokens, mTLS, DPoP | Dev Team | In Progress |
| R005 | PHI exposure in logs | High (4) | High (4) | HIGH | Log sanitization, PHI detection in CI/CD | Dev Team | Open |
| R006 | Unencrypted PHI in Kafka | Medium (3) | High (4) | HIGH | Application-level encryption, Kafka SSL | Dev Team | Open |
| R007 | Database backup theft | Low (2) | Critical (5) | MEDIUM | Encrypted backups, key management | Ops Team | Mitigation |
| R008 | API key/credential exposure | Medium (3) | High (4) | HIGH | Vault secrets management, no hardcoded secrets | All Teams | In Progress |
| R009 | DDoS on patient portal | Medium (3) | Medium (3) | MEDIUM | Rate limiting, WAF, CDN | Ops Team | Mitigation |
| R010 | Third-party dependency CVE | High (4) | Medium (3) | HIGH | Automated scanning, Dependabot, SBOM | Dev Team | Ongoing |
4.2. Risk Treatment Plan

- MITIGATE ← Preferred for HIGH risks: Implement controls, reduce likelihood/impact
- TRANSFER (Transfer): Cyber insurance, outsourcing to specialist provider
- ACCEPT (Accept) ← Only for LOW risks: Document risk acceptance, monitor
- AVOID (Avoid): Eliminate sources of risk, change architecture
5. Data Retention Policy
5.1. Retention Requirements for Vietnamese Health
| Data type | Storage time | Legal basis |
|---|---|---|
| Outpatient medical records | 10 years | Circular 46/2018/TT-BYT |
| Inpatient medical records | 20 years | Circular 46/2018/TT-BYT |
| Death medical records | 20 years | Circular 46/2018/TT-BYT |
| Test results | 10 years | Hospital regulations |
| Diagnostic imaging | 10 years | Hospital regulations |
| Audit logs | 6 years (HIPAA) | HIPAA §164.530(j) |
| Prescription | 5 years | Pharmacy Law |
| Consent records | Lifetime + 6 years | HIPAA / Decree 13/2023 |
5.2. Automated Retention in PostgreSQL
-- Partition strategy for data retention
CREATE TABLE audit_events (
id UUID DEFAULT gen_random_uuid(),
event_time TIMESTAMPTZ NOT NULL DEFAULT NOW(),
event_type VARCHAR(50) NOT NULL,
actor_id UUID NOT NULL,
resource_type VARCHAR(100) NOT NULL,
resource_id UUID,
action VARCHAR(20) NOT NULL,
outcome VARCHAR(20) NOT NULL,
details JSONB
) PARTITION BY RANGE (event_time);
-- Create monthly partitions
CREATE TABLE audit_events_2026_01 PARTITION OF audit_events
FOR VALUES FROM ('2026-01-01') TO ('2026-02-01');
CREATE TABLE audit_events_2026_02 PARTITION OF audit_events
FOR VALUES FROM ('2026-02-01') TO ('2026-03-01');
-- Automated partition management
-- Drop partitions older than retention period (6 years for HIPAA)
-- Archive to cold storage before dropping
6. Summary
In this lesson, we have:
- Develop a 4-level Data Classification Framework for medical data
- Create Data Flow Mapping for PHI via microservices architecture
- Perform Risk Assessment according to NIST SP 800-30 methodology
- Set up Risk Register with risk treatment plans
- Definition of Data Retention Policy according to Vietnamese regulations and HIPAA
Exercises
- Classify all tables/columns in your medical system database into 4 levels
- Draw Data Flow Diagram for 3 main use cases: registering for examination, recording test results, prescribing medicine
- Perform Risk Assessment and create Risk Register for at least 15 risks
- Develop a Data Retention Policy suitable for the organization
| ◀ Previous article | Next article ▶ |
|---|---|
| Lesson 2: Safe Microservices Architecture for Healthcare with Quarkus Stack | Lesson 4: Threat Modeling STRIDE/DREAD for Health Information System |