Chuyển đến nội dung chính

第 14 課:多智能體強化學習 — 合作與競爭

多智能體設定:合作、競爭、混合。 MAPPO、QMIX 演算法。 PettingZoo 框架。博弈論基礎。突發行為為。自玩。

🧠 人工智慧與機器學習 — 第 13 課 第 14 課:多智能體強化學習 — 合作與 競爭

強化學習:從基礎到高級

第 4 部分:RLHF、LLM 調整和製作

亞洲開發網

簡介

多智能體強化學習 (MARL) — 當多個智能體在同一環境中互動。從合作(團隊)到競爭(對抗)再到混合動機場景。


1.MARL 設定

設定獎勵結構範例挑戰
合作共享獎勵機器人團隊學分分配
競賽零和西洋棋、圍棋非平穩
混合個人+共享交通、市場平衡

2.PettingZoo框架

from pettingzoo.mpe import simple_spread_v3

# Parallel API — all agents act simultaneously
env = simple_spread_v3.parallel_env(N=3, max_cycles=25)
observations, infos = env.reset()

while env.agents:
    actions = {
        agent: env.action_space(agent).sample()
        for agent in env.agents
    }
    observations, rewards, terminations, truncations, infos = env.step(actions)

env.close()

3. MAPPO — 多代理 PPO

class MAPPOAgent:
    def __init__(self, obs_dim, act_dim, global_state_dim):
        # Decentralized actor
        self.actor = PolicyNetwork(obs_dim, act_dim)
        # Centralized critic (sees global state)
        self.critic = ValueNetwork(global_state_dim)
    
    def act(self, local_obs):
        return self.actor(local_obs)  # Only local observation
    
    def evaluate(self, global_state):
        return self.critic(global_state)  # Full state info

CTDE:集中訓練,分散執行

  • 訓練:批評者看見一切
  • 執行:演員只能看到局部觀察

4. 自玩

透過與自身的副本進行比賽來訓練代理:

def self_play_training(env, agent, num_games):
    for game in range(num_games):
        obs = env.reset()
        opponent = agent.clone()  # Create copy
        
        while not done:
            action_agent = agent.act(obs["player_1"])
            action_opponent = opponent.act(obs["player_2"])
            obs, rewards, done, _ = env.step({
                "player_1": action_agent,
                "player_2": action_opponent,
            })
        
        agent.update(trajectory)
        # Periodically update opponent pool
        if game % 100 == 0:
            opponent_pool.append(agent.clone())

5. 突發行為

多智能體訓練通常會產生令人驚訝的突發行為為:

  • 通訊:代理開發協議
  • 專業化:角色分化
  • 欺騙:競爭環境中的策略隱藏

總結

演算法設定關鍵想法
MAPPO合作集中的評論家,分散的參與者
QMIX合作單調值分解
自玩競爭與自己對戰
MADDPG混合具有集中批評者的多代理 DDPG