Chuyển đến nội dung chính

Lesson 14: Multi-Agent RL — Cooperation & Competition

Multi-agent settings: cooperative, competitive, mixed. MAPPO, QMIX algorithms. PettingZoo framework. Game theory basics. Emergent behaviors. Self-playing.

🧠 AI & ML — Lesson 13 Lesson 14: Multi-Agent RL — Cooperation & Competition

Reinforcement Learning: From Basics to Advanced

Part 4: RLHF, LLM Alignment & Production

xdev.asia

Introduction

Multi-Agent RL (MARL) — when multiple agents interact in the same environment. From cooperative (team) to competitive (adversarial) to mixed-motive scenarios.


1. MARL Settings

SettingReward StructureExampleChallenge
CooperativeShared rewardsRobot teamCredit assignment
CompetitiveZero-sumChess, GoNon-stationary
MixedIndividual + sharedTraffic, MarketsEquilibrium

2. PettingZoo Framework

from pettingzoo.mpe import simple_spread_v3

# Parallel API — all agents act simultaneously
env = simple_spread_v3.parallel_env(N=3, max_cycles=25)
observations, infos = env.reset()

while env.agents:
    actions = {
        agent: env.action_space(agent).sample()
        for agent in env.agents
    }
    observations, rewards, terminations, truncations, infos = env.step(actions)

env.close()

3. MAPPO — Multi-Agent PPO

class MAPPOAgent:
    def __init__(self, obs_dim, act_dim, global_state_dim):
        # Decentralized actor
        self.actor = PolicyNetwork(obs_dim, act_dim)
        # Centralized critic (sees global state)
        self.critic = ValueNetwork(global_state_dim)
    
    def act(self, local_obs):
        return self.actor(local_obs)  # Only local observation
    
    def evaluate(self, global_state):
        return self.critic(global_state)  # Full state info

CTDE: Centralized Training, Decentralized Execution

  • Training: Critic sees everything
  • Execution: Actor only sees local observation

4. Self-Play

Train agent by playing against copies of itself:

def self_play_training(env, agent, num_games):
    for game in range(num_games):
        obs = env.reset()
        opponent = agent.clone()  # Create copy
        
        while not done:
            action_agent = agent.act(obs["player_1"])
            action_opponent = opponent.act(obs["player_2"])
            obs, rewards, done, _ = env.step({
                "player_1": action_agent,
                "player_2": action_opponent,
            })
        
        agent.update(trajectory)
        # Periodically update opponent pool
        if game % 100 == 0:
            opponent_pool.append(agent.clone())

5. Emergent Behaviors

Multi-agent training often produces surprising emergent behaviors:

  • Communication: Agents develop protocols
  • Specialization: Role differentiation
  • Deception: Strategic hiding in competitive settings

Summary

AlgorithmSettingKey Ideas
MAPPOCooperativeCentralized critics, decentralized actors
QMIXCooperativeMonotonic value decomposition
Self-playingCompetitivePlay against yourself
MADDPGMixedMulti-agent DDPG with centralized critics