Chuyển đến nội dung chính

レッスン 4: Q ラーニングの詳細と多腕の盗賊

Q ラーニング アルゴリズムの詳細。 ε-貪欲な探索。多腕バンディット問題。 UCB、トンプソン・サンプリング。体験タクシー環境。

🧠 AI と ML — レッスン 3 レッスン 4: Q ラーニングの詳細とマルチアーム 山賊

強化学習: 基礎から高度まで

パート 1: RL の基礎 — マルコフの意思決定プロセスと表形式の手法

xdev.asia

はじめに

Q-Learning はオフポリシー TD 制御アルゴリズムであり、DQN およびすべての値ベースのディープ RL の基盤です。 Multi-Armed Bandits は、探索と搾取に焦点を当てた簡素化された RL です。


1. Q ラーニング アルゴリズム

def q_learning(env, num_episodes, alpha=0.1, gamma=0.99, epsilon=0.1):
    Q = np.zeros((env.observation_space.n, env.action_space.n))
    
    for episode in range(num_episodes):
        state, _ = env.reset()
        done = False
        
        while not done:
            # ε-greedy action selection
            if np.random.random() < epsilon:
                action = env.action_space.sample()
            else:
                action = np.argmax(Q[state])
            
            next_state, reward, terminated, truncated, _ = env.step(action)
            done = terminated or truncated
            
            # Q-Learning update (off-policy: max over next actions)
            Q[state, action] += alpha * (
                reward + gamma * np.max(Q[next_state]) * (1 - done) - Q[state, action]
            )
            state = next_state
    return Q

SARSA 対 Q ラーニング

サルサQラーニング
タイプオンポリシーポリシー外
更新Q(s,a) += α[r + γQ(s',a') - Q(s,a)]Q(s,a) += α[r + γ max Q(s',·) - Q(s,a)]
行動より安全、ε-greed に従う最適なポリシーを学習

2. 探索戦略

ε-greedy with Decay

def epsilon_greedy_decay(Q, state, episode, min_epsilon=0.01, decay=0.995):
    epsilon = max(min_epsilon, 1.0 * (decay ** episode))
    if np.random.random() < epsilon:
        return env.action_space.sample()
    return np.argmax(Q[state])

ボルツマン (ソフトマックス) の探査

def boltzmann_action(Q, state, temperature=1.0):
    q_values = Q[state] / temperature
    probs = np.exp(q_values - np.max(q_values))
    probs /= probs.sum()
    return np.random.choice(len(probs), p=probs)

3. 多腕の盗賊

簡略化された RL: 1 つの状態、K のアクション (武器)、報酬を最大化します。

class MultiArmedBandit:
    def __init__(self, k=10):
        self.k = k
        self.true_values = np.random.randn(k)
    
    def pull(self, arm):
        return np.random.randn() + self.true_values[arm]

class UCBAgent:
    def __init__(self, k, c=2.0):
        self.counts = np.zeros(k)
        self.values = np.zeros(k)
        self.c = c
        self.t = 0
    
    def select_arm(self):
        self.t += 1
        if 0 in self.counts:
            return np.argmin(self.counts)
        ucb = self.values + self.c * np.sqrt(np.log(self.t) / self.counts)
        return np.argmax(ucb)
    
    def update(self, arm, reward):
        self.counts[arm] += 1
        self.values[arm] += (reward - self.values[arm]) / self.counts[arm]

4. 体験: タクシー環境

env = gym.make("Taxi-v3")
Q = q_learning(env, num_episodes=10000, alpha=0.1, gamma=0.99, epsilon=0.1)

# Test learned policy
state, _ = env.reset()
total_reward = 0
for _ in range(200):
    action = np.argmax(Q[state])
    state, reward, done, _, _ = env.step(action)
    total_reward += reward
    if done:
        break
print(f"Total reward: {total_reward}")

概要

戦略利点デメリット
ε-貪欲シンプルで効果的均一ランダム探索
ε崩壊残高の探索/活用減衰率を調整する必要があります
UCB原則として、ε は使用しません。決定論的
トンプソンベイズ最適計算コスト
ボルツマンスムーズな温度制御スケールに敏感