はじめに
Q-Learning はオフポリシー TD 制御アルゴリズムであり、DQN およびすべての値ベースのディープ RL の基盤です。 Multi-Armed Bandits は、探索と搾取に焦点を当てた簡素化された RL です。
1. Q ラーニング アルゴリズム
def q_learning(env, num_episodes, alpha=0.1, gamma=0.99, epsilon=0.1):
Q = np.zeros((env.observation_space.n, env.action_space.n))
for episode in range(num_episodes):
state, _ = env.reset()
done = False
while not done:
# ε-greedy action selection
if np.random.random() < epsilon:
action = env.action_space.sample()
else:
action = np.argmax(Q[state])
next_state, reward, terminated, truncated, _ = env.step(action)
done = terminated or truncated
# Q-Learning update (off-policy: max over next actions)
Q[state, action] += alpha * (
reward + gamma * np.max(Q[next_state]) * (1 - done) - Q[state, action]
)
state = next_state
return Q
SARSA 対 Q ラーニング
| サルサ | Qラーニング | |
|---|---|---|
| タイプ | オンポリシー | ポリシー外 |
| 更新 | Q(s,a) += α[r + γQ(s',a') - Q(s,a)] | Q(s,a) += α[r + γ max Q(s',·) - Q(s,a)] |
| 行動 | より安全、ε-greed に従う | 最適なポリシーを学習 |
2. 探索戦略
ε-greedy with Decay
def epsilon_greedy_decay(Q, state, episode, min_epsilon=0.01, decay=0.995):
epsilon = max(min_epsilon, 1.0 * (decay ** episode))
if np.random.random() < epsilon:
return env.action_space.sample()
return np.argmax(Q[state])
ボルツマン (ソフトマックス) の探査
def boltzmann_action(Q, state, temperature=1.0):
q_values = Q[state] / temperature
probs = np.exp(q_values - np.max(q_values))
probs /= probs.sum()
return np.random.choice(len(probs), p=probs)
3. 多腕の盗賊
簡略化された RL: 1 つの状態、K のアクション (武器)、報酬を最大化します。
class MultiArmedBandit:
def __init__(self, k=10):
self.k = k
self.true_values = np.random.randn(k)
def pull(self, arm):
return np.random.randn() + self.true_values[arm]
class UCBAgent:
def __init__(self, k, c=2.0):
self.counts = np.zeros(k)
self.values = np.zeros(k)
self.c = c
self.t = 0
def select_arm(self):
self.t += 1
if 0 in self.counts:
return np.argmin(self.counts)
ucb = self.values + self.c * np.sqrt(np.log(self.t) / self.counts)
return np.argmax(ucb)
def update(self, arm, reward):
self.counts[arm] += 1
self.values[arm] += (reward - self.values[arm]) / self.counts[arm]
4. 体験: タクシー環境
env = gym.make("Taxi-v3")
Q = q_learning(env, num_episodes=10000, alpha=0.1, gamma=0.99, epsilon=0.1)
# Test learned policy
state, _ = env.reset()
total_reward = 0
for _ in range(200):
action = np.argmax(Q[state])
state, reward, done, _, _ = env.step(action)
total_reward += reward
if done:
break
print(f"Total reward: {total_reward}")
概要
| 戦略 | 利点 | デメリット |
|---|---|---|
| ε-貪欲 | シンプルで効果的 | 均一ランダム探索 |
| ε崩壊 | 残高の探索/活用 | 減衰率を調整する必要があります |
| UCB | 原則として、ε は使用しません。決定論的 | |
| トンプソン | ベイズ最適 | 計算コスト |
| ボルツマン | スムーズな温度制御 | スケールに敏感 |