Chuyển đến nội dung chính

Lesson 1: What is Reinforcement Learning? — Agent, Environment & Rewards

Define RL, compare supervised/unsupervised/RL. Agent-Environment interaction loop. State, Action, Reward, Policy, Value function. MDP. Exploration vs Exploitation.

🧠 AI & ML — Lesson 0 Lesson 1: What is Reinforcement Learning? — Agent, Environment & Reward

Reinforcement Learning: From Basics to Advanced

Part 1: RL Foundation — Markov Decision Process & Tabular Methods

xdev.asia

Introduction

Reinforcement Learning (RL) is the third paradigm of Machine Learning — the agent learns how to act in the environment to maximize cumulative reward.


1. RL vs Supervised vs Unsupervised

ParadigmDataFeedbackExample
Supervised(x, y) pairsLabelsImage classification
Unsupervisedx onlyNoneClustering
RLStates, actionsRewards (delayed)Game playing

Unique characteristics of RL

  • Sequential decision making: Decisions that affect the future
  • Delayed reward: Not knowing immediately whether the action is good or bad
  • Exploration vs Exploitation: Try new things vs exploit known knowledge
  • No supervisor: Only reward signal

2. Agent-Environment Interaction Loop

Every RL problem follows a loop:

Agent quan sát state s_t
  → Chọn action a_t theo policy π
  → Environment trả về reward r_{t+1} và state mới s_{t+1}
  → Lặp lại

Example: Robot goes through a maze

import gymnasium as gym

env = gym.make("FrozenLake-v1", render_mode="human")
state, info = env.reset()

for step in range(100):
    action = env.action_space.sample()  # Random policy
    next_state, reward, terminated, truncated, info = env.step(action)
    print(f"State: {state}, Action: {action}, Reward: {reward}")
    
    if terminated or truncated:
        state, info = env.reset()
    else:
        state = next_state

3. Core concepts

State(s)

Fully describes the current state of the environment.

Action (a)

Decide which agent to execute — discrete (left/right) or continuous (rotation angle).

Rewards (r)

Scalar response from environment — agent wants to maximize total reward.

Policy (π)

State → action mapping strategy:

  • Deterministic: π(s) = a
  • Stochastic: π(a|s) = P(a|s)

Value Function V(s)

Expected cumulative reward when starting from state s:

$$V^\pi(s) = \mathbb{E}\pi\left[\sum{t=0}^{\infty} \gamma^t r_{t+1} | s_0 = s\right]$$

Q-function Q(s,a)

Expected cumulative reward when choosing action a at state s:

$$Q^\pi(s,a) = \mathbb{E}\pi\left[\sum{t=0}^{\infty} \gamma^t r_{t+1} | s_0 = s, a_0 = a\right]$$


4. Markov Decision Process (MDP)

MDP is the standard mathematical framework for RL: (S, A, P, R, γ)

IngredientsSymbolDescription
StatesSCollection of all states
ActionsACollection of all actions
TransitionP(s's,a)
RewardsR(s,a,s')Reward function
Discountγ ∈ [0,1]Discount factor

Markov Property: The future depends only on the present, not the past.


5. Exploration vs Exploitation

StrategyDescriptionTrade-off
ExplorationTry new actionFind a better strategy
ExploitationUse the best known actionMaximize short-term rewards

Balancing method:

  • ε-greedy: Random with probability ε
  • UCB: Upper Confidence Bound
  • Thompson Sampling: Bayesian approach

Summary

ConceptDescription
RLAgent learns from interaction with environment
MDPMath framework: states, actions, rewards
PolicyAction selection strategy
ValueExpected cumulative reward
ExplorationBalance between trying new and exploiting