About the project
A single reinforcement-learning policy for OTTO that balances, rejects pushes, follows velocity commands, and changes body height.
It is the learning-based counterpart to OTTO's model-based controllers, trained in Isaac Lab with the gap to real hardware in mind.
Early velocity policies that were forced to stay upright could not produce thrust and reached only about half the commanded speed, because a wheeled biped has to lean to accelerate. The final policy may lean into acceleration and turns, while a reward keeps its legs close to the robot's analytical balance posture.
The policy only uses signals the real robot can measure, so it never relies on the body's linear velocity. This keeps the same network deployable on hardware without extra estimation.
Process
Define the interface
The policy reads a 28-dimensional observation and outputs 6 actions at 50 Hz, position targets for the legs and torques for the wheels.
Train skills in stages
Balance, velocity, height and terrain are trained in order, each warm-started from the last, then fused into one policy.
Randomize for transfer
Friction, mass, actuator gains and random pushes vary during training so the policy does not overfit the simulator.
Evaluate across commands
The policy is tested over the full command range, including the hardest case of reverse drive with a full spin.
Results
- The fall rate is 0% to 0.8% on training terrain and 8.6% in the reverse-drive plus full-spin case.
- Velocity tracking reaches about 70% to 75% of the command.
- Body height can be controlled over ±10 cm, limited by the knee range.
Figures
Where it stands
Everything runs in simulation so far. The blind terrain policy still falls on drops and cliffs outside its training data. Sim-to-sim checks and hardware deployment come next.
System
| Simulator | Isaac Lab on Isaac Sim 5.1 |
|---|---|
| Algorithm | PPO (rsl_rl) |
| Policy | 28 observations, 6 actions, 50 Hz |
| Training GPU | RTX 4070 Ti SUPER, 16 GB |

