Introduction
Agents have the right to act — meaning the wrong agent can cause real damage: deleting files, sending wrong emails, leaking data. Safety is not a nice-to-have — it is a mandatory requirement.
1. Threats
1.1 Prompt Injection
Users intentionally provide hidden instructions to hijack agent behavior.
1.2 Tool Misuse
Agent calls the tool the wrong way: DELETE instead of SELECT, sending email to the wrong person.
1.3 Data Leakage
Agent accidentally exposed sensitive data in response.
2. Defense Layers
class GuardedAgent:
def run(self, user_input):
# Layer 1: Input validation
if self.detect_injection(user_input):
return "Suspicious input detected"
# Layer 2: Tool permission check
# Only allow approved tools
# Layer 3: Output filtering
output = self.agent.run(user_input)
output = self.filter_pii(output)
# Layer 4: Human approval for risky actions
if self.is_risky_action(output):
return self.request_human_approval(output)
return output
Summary
- Agent safety = input validation + tool permissions + output filtering + human approval
- Prompt injection is threat #1
- Never give agent DELETE/UPDATE permissions without approval
- Cost controls prevent runaway spending
- Human-in-the-loop for high-stakes decisions
Exercises
- Implement prompt injection detector
- Build permission system for tools (read-only vs read-write)
- Implement PII filter (mask emails, phone numbers)
- Build human-in-the-loop approval flow