はじめに
マルチエージェント RL (MARL) — 複数のエージェントが同じ環境で対話する場合。協力的 (チーム) から競争的 (敵対的)、そして動機が混在するシナリオまで。
1. MARLの設定
| 設定 | 報酬体系 | 例 | チャレンジ |
|---|---|---|---|
| 協同組合 | 共有特典 | ロボットチーム | 単位の割り当て |
| 競争力 | ゼロサム | チェス、囲碁 | 非定常 |
| 混合 | 個人 + 共有 | 交通、市場 | 平衡 |
2. PettingZoo フレームワーク
from pettingzoo.mpe import simple_spread_v3
# Parallel API — all agents act simultaneously
env = simple_spread_v3.parallel_env(N=3, max_cycles=25)
observations, infos = env.reset()
while env.agents:
actions = {
agent: env.action_space(agent).sample()
for agent in env.agents
}
observations, rewards, terminations, truncations, infos = env.step(actions)
env.close()
3. MAPPO — マルチエージェント PPO
class MAPPOAgent:
def __init__(self, obs_dim, act_dim, global_state_dim):
# Decentralized actor
self.actor = PolicyNetwork(obs_dim, act_dim)
# Centralized critic (sees global state)
self.critic = ValueNetwork(global_state_dim)
def act(self, local_obs):
return self.actor(local_obs) # Only local observation
def evaluate(self, global_state):
return self.critic(global_state) # Full state info
CTDE: 集中トレーニング、分散実行
- トレーニング: 批評家はすべてを見ています
- 実行: アクターはローカルの観察のみを参照します。
4. セルフプレイ
エージェント自身のコピーと対戦してエージェントをトレーニングします。
def self_play_training(env, agent, num_games):
for game in range(num_games):
obs = env.reset()
opponent = agent.clone() # Create copy
while not done:
action_agent = agent.act(obs["player_1"])
action_opponent = opponent.act(obs["player_2"])
obs, rewards, done, _ = env.step({
"player_1": action_agent,
"player_2": action_opponent,
})
agent.update(trajectory)
# Periodically update opponent pool
if game % 100 == 0:
opponent_pool.append(agent.clone())
5. 緊急の行動
マルチエージェントのトレーニングでは、多くの場合、驚くべき創発的な行動が生成されます。
- コミュニケーション: エージェントがプロトコルを開発
- 専門分野: 役割の差別化
- 欺瞞: 競争環境における戦略的な隠蔽
概要
| アルゴリズム | 設定 | 主要なアイデア |
|---|---|---|
| マッポ | 協同組合 | 集中化された批評家と分散化されたアクター |
| QMIX | 協同組合 | 単調値分解 |
| 自動演奏 | 競争力 | 自分自身と対戦する |
| マッドペグ | 混合 | 中央集中型の批評家を備えたマルチエージェント DDPG |