Introduction
Capstone project applies all RL knowledge to a real end-to-end problem. Choose 1 of 3 projects below.
Project 1: Game AI Agent
Description
Build AI gaming agents — from custom environments to trained agents that can demo.
Technical Stack
- Gymnasium custom environment (Snake, Flappy Bird, Tetris)
- DQN or PPO training
- Hyperparameter optimization with Optuna
- Web demo with Gradio
Steps
# 1. Build custom environment
class GameEnv(gym.Env):
# Implement reset(), step(), render()
pass
# 2. Train agent
from stable_baselines3 import PPO
model = PPO("MlpPolicy", GameEnv(), verbose=1)
model.learn(total_timesteps=1_000_000)
# 3. Evaluate
mean_reward, std = evaluate_policy(model, GameEnv(), n_eval_episodes=100)
print(f"Score: {mean_reward:.1f} +/- {std:.1f}")
# 4. Demo
import gradio as gr
def play_game(seed):
env = GameEnv(render_mode="rgb_array")
frames = record_episode(model, env, seed)
return frames
Project 2: Robot Control
Description
Train robot locomotion agent in MuJoCo — walking, running, or manipulation.
Technical Stack
- MuJoCo (Ant, Humanoid, or custom robot)
- SAC training with domain randomization
- TensorBoard analysis
- Sim-to-real transfer analysis
Evaluation
# Compare algorithms
algorithms = {
"PPO": PPO("MlpPolicy", env),
"SAC": SAC("MlpPolicy", env),
"TD3": TD3("MlpPolicy", env),
}
results = {}
for name, model in algorithms.items():
model.learn(total_timesteps=1_000_000)
mean_reward, _ = evaluate_policy(model, env, n_eval_episodes=50)
results[name] = mean_reward
Project 3: RLHF / DPO Chatbot
Description
Align a small LLM with human preferences using DPO or RLHF.
Technical Stack
- Base model: SmolLM or Qwen2.5 (0.5B-1.5B)
- TRL library for SFT + DPO
- Evaluation: MT-Bench, AlpacaEval
- Gradio chat interface
Pipelines
# 1. SFT
sft_trainer = SFTTrainer(model, train_dataset=sft_data)
sft_trainer.train()
# 2. DPO
dpo_trainer = DPOTrainer(model, ref_model, train_dataset=pref_data)
dpo_trainer.train()
# 3. Evaluate
# - Perplexity
# - Win rate vs base model
# - Human evaluation
# 4. Deploy
import gradio as gr
demo = gr.ChatInterface(fn=generate_response)
demo.launch()
Deliverables
| Item | Description | Weight |
|---|---|---|
| Code | Clean, documented GitHub repository | 30% |
| Training logs | TensorBoard visualizations, learning curves | 20% |
| Report | Architecture decisions, results analysis, ablations | 30% |
| Demo | Interactive demo (web app or video) | 20% |
Summary
Congratulations on completing the Reinforcement Learning: From Basic to Advanced series!
Learned knowledge
| Part | Main content |
|---|---|
| 1. Platform | MDP, DP, MC, TD, Q-Learning |
| 2. Deep RL | DQN, Policy Gradient, PPO, SAC |
| 3. Frameworks | Gymnasium, SB3, MuJoCo |
| 4. Production | RLHF, DPO, Multi-agent, Deploy |
Further development direction
- Research: Read papers on arXiv, reproduce results
- Competition: Kaggle RL, NeurIPS challenges
- Open-source: Contribute to SB3, TRL, PettingZoo
- Career: RL Engineer, AI Safety Researcher, Robotics Engineer