Chuyển đến nội dung chính

Reinforcement Learning: From Basics to Advanced

Comprehensive Reinforcement Learning course — from MDP, Q-Learning to Deep RL platforms with DQN, Policy Gradient, PPO, SAC. Practice game AI, robotics simulation, RLHF for LLM, and deploy RL agents production-ready with Python, Gymnasium, Stable-Baselines3.

Introducing the Series

Reinforcement Learning: From Basics to Advanced is a course that helps you master the entire field of RL — from MDP, Q-Learning platforms to modern Deep RL with PPO, SAC, and RLHF for LLM Alignment.

🎯 After completing the course, you will:

  • Deep understanding of RL theory: MDP, Bellman, Policy Gradient
  • Implement DQN, PPO, SAC from scratch
  • Proficient in using Gymnasium & Stable-Baselines3
  • Understand RLHF/DPO for LLM alignment
  • Build RL agents for games, robotics, and real-world problems

Study path

Part 1: RL Platform

MDP, Dynamic Programming, Q-Learning — your first step into the world of RL.

Part 2: Deep Reinforcement Learning

DQN, Policy Gradient, PPO, SAC — when neural networks encounter RL.

Part 3: Frameworks & Practices

Gymnasium, Stable-Baselines3, robotics simulation.

Part 4: RLHF & Production

RL for LLM alignment, multi-agent, and production deployment.