pboon09
Institute of Field Robotics (FIBO), KMUTT

Curriculum Learning for Velocity Tracking on Unitree Go2

April 2026 – May 2026

About the project

A study of how two training choices, curriculum learning and reward design, each shape a Unitree Go2 quadruped learning to follow velocity commands in Isaac Lab.

Velocity commands are grouped into speed bins, and a curriculum decides which bins the robot trains on at each stage. Uniform sampling trains on every speed from the start. Task-specific sampling unlocks faster bins once the robot tracks the current ones well enough. Teacher-guided sampling puts more training where learning progress is highest.

The policy is trained with PPO, and each curriculum is paired with two reward designs. This separates what the curriculum changes from what the reward changes.

3Curriculum strategies
2Reward designs
6Trained combinations

Process

  1. Separate the two questions

    Which commands to train on and when is a curriculum question. What behavior to encourage is a reward question. The study tests them independently.

  2. Train every combination

    Each of the three curricula, uniform sampling, task-specific, and teacher-guided, is paired with each of the two reward designs and trained with PPO.

  3. Compare speed range and gait

    Each policy is judged on how wide a range of speeds it tracks and how stable its gait stays.

Results

  • Curriculum learning widens the range of speeds the robot can track.
  • A better-shaped reward keeps the gait stable and prevents falls.
  • The best result needs both, because they fix different problems.

Figures

Training pipeline.
Training pipeline.
Achieved versus commanded forward velocity.
Achieved versus commanded forward velocity.
Task-sampling distribution over training.
Task-sampling distribution over training.

System

RobotUnitree Go2 quadruped
SimulatorIsaac Lab
Experiment3 curricula × 2 reward designs