Chuyển đến nội dung chính

Lesson 16: Capstone — Building RL Agent for Real-world Problem

Summary project: Choose 1 of 3 projects: Game AI Agent, Robot Control, or RLHF for Chatbot. End-to-end pipeline from design to deployment.

🧠 AI & ML — Lesson 15 Lesson 16: Capstone — Building RL Agent for Real-world Problem

Reinforcement Learning: From Basics to Advanced

Part 4: RLHF, LLM Alignment & Production

xdev.asia

Introduction

Capstone project applies all RL knowledge to a real end-to-end problem. Choose 1 of 3 projects below.


Project 1: Game AI Agent

Description

Build AI gaming agents — from custom environments to trained agents that can demo.

Technical Stack

  • Gymnasium custom environment (Snake, Flappy Bird, Tetris)
  • DQN or PPO training
  • Hyperparameter optimization with Optuna
  • Web demo with Gradio

Steps

# 1. Build custom environment
class GameEnv(gym.Env):
    # Implement reset(), step(), render()
    pass

# 2. Train agent
from stable_baselines3 import PPO
model = PPO("MlpPolicy", GameEnv(), verbose=1)
model.learn(total_timesteps=1_000_000)

# 3. Evaluate
mean_reward, std = evaluate_policy(model, GameEnv(), n_eval_episodes=100)
print(f"Score: {mean_reward:.1f} +/- {std:.1f}")

# 4. Demo
import gradio as gr
def play_game(seed):
    env = GameEnv(render_mode="rgb_array")
    frames = record_episode(model, env, seed)
    return frames

Project 2: Robot Control

Description

Train robot locomotion agent in MuJoCo — walking, running, or manipulation.

Technical Stack

  • MuJoCo (Ant, Humanoid, or custom robot)
  • SAC training with domain randomization
  • TensorBoard analysis
  • Sim-to-real transfer analysis

Evaluation

# Compare algorithms
algorithms = {
    "PPO": PPO("MlpPolicy", env),
    "SAC": SAC("MlpPolicy", env),
    "TD3": TD3("MlpPolicy", env),
}

results = {}
for name, model in algorithms.items():
    model.learn(total_timesteps=1_000_000)
    mean_reward, _ = evaluate_policy(model, env, n_eval_episodes=50)
    results[name] = mean_reward

Project 3: RLHF / DPO Chatbot

Description

Align a small LLM with human preferences using DPO or RLHF.

Technical Stack

  • Base model: SmolLM or Qwen2.5 (0.5B-1.5B)
  • TRL library for SFT + DPO
  • Evaluation: MT-Bench, AlpacaEval
  • Gradio chat interface

Pipelines

# 1. SFT
sft_trainer = SFTTrainer(model, train_dataset=sft_data)
sft_trainer.train()

# 2. DPO
dpo_trainer = DPOTrainer(model, ref_model, train_dataset=pref_data)
dpo_trainer.train()

# 3. Evaluate
# - Perplexity
# - Win rate vs base model
# - Human evaluation

# 4. Deploy
import gradio as gr
demo = gr.ChatInterface(fn=generate_response)
demo.launch()

Deliverables

ItemDescriptionWeight
CodeClean, documented GitHub repository30%
Training logsTensorBoard visualizations, learning curves20%
ReportArchitecture decisions, results analysis, ablations30%
DemoInteractive demo (web app or video)20%

Summary

Congratulations on completing the Reinforcement Learning: From Basic to Advanced series!

Learned knowledge

PartMain content
1. PlatformMDP, DP, MC, TD, Q-Learning
2. Deep RLDQN, Policy Gradient, PPO, SAC
3. FrameworksGymnasium, SB3, MuJoCo
4. ProductionRLHF, DPO, Multi-agent, Deploy

Further development direction

  • Research: Read papers on arXiv, reproduce results
  • Competition: Kaggle RL, NeurIPS challenges
  • Open-source: Contribute to SB3, TRL, PettingZoo
  • Career: RL Engineer, AI Safety Researcher, Robotics Engineer