簡介
多智能體強化學習 (MARL) — 當多個智能體在同一環境中互動。從合作(團隊)到競爭(對抗)再到混合動機場景。
1.MARL 設定
| 設定 | 獎勵結構 | 範例 | 挑戰 |
|---|---|---|---|
| 合作 | 共享獎勵 | 機器人團隊 | 學分分配 |
| 競賽 | 零和 | 西洋棋、圍棋 | 非平穩 |
| 混合 | 個人+共享 | 交通、市場 | 平衡 |
2.PettingZoo框架
from pettingzoo.mpe import simple_spread_v3
# Parallel API — all agents act simultaneously
env = simple_spread_v3.parallel_env(N=3, max_cycles=25)
observations, infos = env.reset()
while env.agents:
actions = {
agent: env.action_space(agent).sample()
for agent in env.agents
}
observations, rewards, terminations, truncations, infos = env.step(actions)
env.close()
3. MAPPO — 多代理 PPO
class MAPPOAgent:
def __init__(self, obs_dim, act_dim, global_state_dim):
# Decentralized actor
self.actor = PolicyNetwork(obs_dim, act_dim)
# Centralized critic (sees global state)
self.critic = ValueNetwork(global_state_dim)
def act(self, local_obs):
return self.actor(local_obs) # Only local observation
def evaluate(self, global_state):
return self.critic(global_state) # Full state info
CTDE:集中訓練,分散執行
- 訓練:批評者看見一切
- 執行:演員只能看到局部觀察
4. 自玩
透過與自身的副本進行比賽來訓練代理:
def self_play_training(env, agent, num_games):
for game in range(num_games):
obs = env.reset()
opponent = agent.clone() # Create copy
while not done:
action_agent = agent.act(obs["player_1"])
action_opponent = opponent.act(obs["player_2"])
obs, rewards, done, _ = env.step({
"player_1": action_agent,
"player_2": action_opponent,
})
agent.update(trajectory)
# Periodically update opponent pool
if game % 100 == 0:
opponent_pool.append(agent.clone())
5. 突發行為
多智能體訓練通常會產生令人驚訝的突發行為為:
- 通訊:代理開發協議
- 專業化:角色分化
- 欺騙:競爭環境中的策略隱藏
總結
| 演算法 | 設定 | 關鍵想法 |
|---|---|---|
| MAPPO | 合作 | 集中的評論家,分散的參與者 |
| QMIX | 合作 | 單調值分解 |
| 自玩 | 競爭 | 與自己對戰 |
| MADDPG | 混合 | 具有集中批評者的多代理 DDPG |