Chuyển đến nội dung chính

Lesson 15: Guardrails & Safety — Protect Agents from "rebellion"

Prompt injection defense, output validation, PII filtering. Guardrails frameworks: NeMo Guardrails, Guardrails AI. Human-in-the-loop patterns. Rate limiting and cost controls.

🧠 AI & ML — Lesson 14 Lesson 15: Guardrails & Safety — Protecting Agents from "rebellion"

Build AI Agents: From Zero to Production

Part 6: Production & Actual Deployment

xdev.asia

Introduction

Agents have the right to act — meaning the wrong agent can cause real damage: deleting files, sending wrong emails, leaking data. Safety is not a nice-to-have — it is a mandatory requirement.


1. Threats

1.1 Prompt Injection

Users intentionally provide hidden instructions to hijack agent behavior.

1.2 Tool Misuse

Agent calls the tool the wrong way: DELETE instead of SELECT, sending email to the wrong person.

1.3 Data Leakage

Agent accidentally exposed sensitive data in response.

2. Defense Layers

class GuardedAgent:
    def run(self, user_input):
        # Layer 1: Input validation
        if self.detect_injection(user_input):
            return "Suspicious input detected"
        
        # Layer 2: Tool permission check
        # Only allow approved tools
        
        # Layer 3: Output filtering
        output = self.agent.run(user_input)
        output = self.filter_pii(output)
        
        # Layer 4: Human approval for risky actions
        if self.is_risky_action(output):
            return self.request_human_approval(output)
        
        return output

Summary

  • Agent safety = input validation + tool permissions + output filtering + human approval
  • Prompt injection is threat #1
  • Never give agent DELETE/UPDATE permissions without approval
  • Cost controls prevent runaway spending
  • Human-in-the-loop for high-stakes decisions

Exercises

  1. Implement prompt injection detector
  2. Build permission system for tools (read-only vs read-write)
  3. Implement PII filter (mask emails, phone numbers)
  4. Build human-in-the-loop approval flow