About the project
A study of how two training choices, curriculum learning and reward design, each shape a Unitree Go2 quadruped learning to follow velocity commands in Isaac Lab.
Velocity commands are grouped into speed bins, and a curriculum decides which bins the robot trains on at each stage. Uniform sampling trains on every speed from the start. Task-specific sampling unlocks faster bins once the robot tracks the current ones well enough. Teacher-guided sampling puts more training where learning progress is highest.
The policy is trained with PPO, and each curriculum is paired with two reward designs. This separates what the curriculum changes from what the reward changes.
Process
Separate the two questions
Which commands to train on and when is a curriculum question. What behavior to encourage is a reward question. The study tests them independently.
Train every combination
Each of the three curricula, uniform sampling, task-specific, and teacher-guided, is paired with each of the two reward designs and trained with PPO.
Compare speed range and gait
Each policy is judged on how wide a range of speeds it tracks and how stable its gait stays.
Results
- Curriculum learning widens the range of speeds the robot can track.
- A better-shaped reward keeps the gait stable and prevents falls.
- The best result needs both, because they fix different problems.
Figures



System
| Robot | Unitree Go2 quadruped |
|---|---|
| Simulator | Isaac Lab |
| Experiment | 3 curricula × 2 reward designs |

