Giới thiệu
Agent có quyền hành động — nghĩa là agent sai có thể gây thiệt hại thực tế: xóa file, gửi email sai, leak data. Safety không phải nice-to-have — nó là requirement bắt buộc.
1. Threats
1.1 Prompt Injection
User cố tình đưa instructions ẩn để hijack agent behavior.
1.2 Tool Misuse
Agent gọi tool sai cách: DELETE thay vì SELECT, gửi email cho sai người.
1.3 Data Leakage
Agent vô tình expose sensitive data trong response.
2. Defense Layers
class GuardedAgent:
def run(self, user_input):
# Layer 1: Input validation
if self.detect_injection(user_input):
return "Suspicious input detected"
# Layer 2: Tool permission check
# Only allow approved tools
# Layer 3: Output filtering
output = self.agent.run(user_input)
output = self.filter_pii(output)
# Layer 4: Human approval for risky actions
if self.is_risky_action(output):
return self.request_human_approval(output)
return output
Tóm tắt
- Agent safety = input validation + tool permissions + output filtering + human approval
- Prompt injection là threat #1
- Never give agent DELETE/UPDATE permissions without approval
- Cost controls prevent runaway spending
- Human-in-the-loop cho high-stakes decisions
Bài tập
- Implement prompt injection detector
- Build permission system cho tools (read-only vs read-write)
- Implement PII filter (mask emails, phone numbers)
- Xây human-in-the-loop approval flow