pboon09
Institute of Field Robotics (FIBO), KMUTT

Reinforcement Learning for Wheeled Bipedal Locomotion

June 2026 – July 2026

About the project

A single reinforcement-learning policy for OTTO that balances, rejects pushes, follows velocity commands, and changes body height.

It is the learning-based counterpart to OTTO's model-based controllers, trained in Isaac Lab with the gap to real hardware in mind.

Early velocity policies that were forced to stay upright could not produce thrust and reached only about half the commanded speed, because a wheeled biped has to lean to accelerate. The final policy may lean into acceleration and turns, while a reward keeps its legs close to the robot's analytical balance posture.

The policy only uses signals the real robot can measure, so it never relies on the body's linear velocity. This keeps the same network deployable on hardware without extra estimation.

0% to 0.8%Fall rate on training terrain
70% to 75%Of commanded speed tracked
±10 cmBody-height range
50 HzPolicy rate

Process

  1. Define the interface

    The policy reads a 28-dimensional observation and outputs 6 actions at 50 Hz, position targets for the legs and torques for the wheels.

  2. Train skills in stages

    Balance, velocity, height and terrain are trained in order, each warm-started from the last, then fused into one policy.

  3. Randomize for transfer

    Friction, mass, actuator gains and random pushes vary during training so the policy does not overfit the simulator.

  4. Evaluate across commands

    The policy is tested over the full command range, including the hardest case of reverse drive with a full spin.

Results

  • The fall rate is 0% to 0.8% on training terrain and 8.6% in the reverse-drive plus full-spin case.
  • Velocity tracking reaches about 70% to 75% of the command.
  • Body height can be controlled over ±10 cm, limited by the knee range.

Figures

Course navigation.
Height control.
Spinning.

Where it stands

Everything runs in simulation so far. The blind terrain policy still falls on drops and cliffs outside its training data. Sim-to-sim checks and hardware deployment come next.

System

SimulatorIsaac Lab on Isaac Sim 5.1
AlgorithmPPO (rsl_rl)
Policy28 observations, 6 actions, 50 Hz
Training GPURTX 4070 Ti SUPER, 16 GB