Introduction
Reinforcement Learning (RL) is the third paradigm of Machine Learning — the agent learns how to act in the environment to maximize cumulative reward.
1. RL vs Supervised vs Unsupervised
| Paradigm | Data | Feedback | Example |
|---|---|---|---|
| Supervised | (x, y) pairs | Labels | Image classification |
| Unsupervised | x only | None | Clustering |
| RL | States, actions | Rewards (delayed) | Game playing |
Unique characteristics of RL
- Sequential decision making: Decisions that affect the future
- Delayed reward: Not knowing immediately whether the action is good or bad
- Exploration vs Exploitation: Try new things vs exploit known knowledge
- No supervisor: Only reward signal
2. Agent-Environment Interaction Loop
Every RL problem follows a loop:
Agent quan sát state s_t
→ Chọn action a_t theo policy π
→ Environment trả về reward r_{t+1} và state mới s_{t+1}
→ Lặp lại
Example: Robot goes through a maze
import gymnasium as gym
env = gym.make("FrozenLake-v1", render_mode="human")
state, info = env.reset()
for step in range(100):
action = env.action_space.sample() # Random policy
next_state, reward, terminated, truncated, info = env.step(action)
print(f"State: {state}, Action: {action}, Reward: {reward}")
if terminated or truncated:
state, info = env.reset()
else:
state = next_state
3. Core concepts
State(s)
Fully describes the current state of the environment.
Action (a)
Decide which agent to execute — discrete (left/right) or continuous (rotation angle).
Rewards (r)
Scalar response from environment — agent wants to maximize total reward.
Policy (π)
State → action mapping strategy:
- Deterministic: π(s) = a
- Stochastic: π(a|s) = P(a|s)
Value Function V(s)
Expected cumulative reward when starting from state s:
$$V^\pi(s) = \mathbb{E}\pi\left[\sum{t=0}^{\infty} \gamma^t r_{t+1} | s_0 = s\right]$$
Q-function Q(s,a)
Expected cumulative reward when choosing action a at state s:
$$Q^\pi(s,a) = \mathbb{E}\pi\left[\sum{t=0}^{\infty} \gamma^t r_{t+1} | s_0 = s, a_0 = a\right]$$
4. Markov Decision Process (MDP)
MDP is the standard mathematical framework for RL: (S, A, P, R, γ)
| Ingredients | Symbol | Description |
|---|---|---|
| States | S | Collection of all states |
| Actions | A | Collection of all actions |
| Transition | P(s' | s,a) |
| Rewards | R(s,a,s') | Reward function |
| Discount | γ ∈ [0,1] | Discount factor |
Markov Property: The future depends only on the present, not the past.
5. Exploration vs Exploitation
| Strategy | Description | Trade-off |
|---|---|---|
| Exploration | Try new action | Find a better strategy |
| Exploitation | Use the best known action | Maximize short-term rewards |
Balancing method:
- ε-greedy: Random with probability ε
- UCB: Upper Confidence Bound
- Thompson Sampling: Bayesian approach
Summary
| Concept | Description |
|---|---|
| RL | Agent learns from interaction with environment |
| MDP | Math framework: states, actions, rewards |
| Policy | Action selection strategy |
| Value | Expected cumulative reward |
| Exploration | Balance between trying new and exploiting |