Introduction
Multi-Agent RL (MARL) — when multiple agents interact in the same environment. From cooperative (team) to competitive (adversarial) to mixed-motive scenarios.
1. MARL Settings
| Setting | Reward Structure | Example | Challenge |
|---|---|---|---|
| Cooperative | Shared rewards | Robot team | Credit assignment |
| Competitive | Zero-sum | Chess, Go | Non-stationary |
| Mixed | Individual + shared | Traffic, Markets | Equilibrium |
2. PettingZoo Framework
from pettingzoo.mpe import simple_spread_v3
# Parallel API — all agents act simultaneously
env = simple_spread_v3.parallel_env(N=3, max_cycles=25)
observations, infos = env.reset()
while env.agents:
actions = {
agent: env.action_space(agent).sample()
for agent in env.agents
}
observations, rewards, terminations, truncations, infos = env.step(actions)
env.close()
3. MAPPO — Multi-Agent PPO
class MAPPOAgent:
def __init__(self, obs_dim, act_dim, global_state_dim):
# Decentralized actor
self.actor = PolicyNetwork(obs_dim, act_dim)
# Centralized critic (sees global state)
self.critic = ValueNetwork(global_state_dim)
def act(self, local_obs):
return self.actor(local_obs) # Only local observation
def evaluate(self, global_state):
return self.critic(global_state) # Full state info
CTDE: Centralized Training, Decentralized Execution
- Training: Critic sees everything
- Execution: Actor only sees local observation
4. Self-Play
Train agent by playing against copies of itself:
def self_play_training(env, agent, num_games):
for game in range(num_games):
obs = env.reset()
opponent = agent.clone() # Create copy
while not done:
action_agent = agent.act(obs["player_1"])
action_opponent = opponent.act(obs["player_2"])
obs, rewards, done, _ = env.step({
"player_1": action_agent,
"player_2": action_opponent,
})
agent.update(trajectory)
# Periodically update opponent pool
if game % 100 == 0:
opponent_pool.append(agent.clone())
5. Emergent Behaviors
Multi-agent training often produces surprising emergent behaviors:
- Communication: Agents develop protocols
- Specialization: Role differentiation
- Deception: Strategic hiding in competitive settings
Summary
| Algorithm | Setting | Key Ideas |
|---|---|---|
| MAPPO | Cooperative | Centralized critics, decentralized actors |
| QMIX | Cooperative | Monotonic value decomposition |
| Self-playing | Competitive | Play against yourself |
| MADDPG | Mixed | Multi-agent DDPG with centralized critics |