Robot Arm

A simulated Franka Panda arm, trained with reinforcement learning, picks up a box and throws it as far as it can. Grab the box and drop it anywhere, and the arm goes and gets it.

Loading the physics engine and the trained policy…

Grab the box and drop it anywhere · drag elsewhere to turn the view

0Throws
–Last
–Best

How it works

The arm is the Franka Emika Panda model from MuJoCo Menagerie, set up as in the PandaPickCube task of MuJoCo Playground. Your browser simulates it with MuJoCo compiled to WebAssembly, in 5 ms physics steps. Every 20 ms a policy, a small neural network, reads the joint angles and speeds and where the box is, and moves each joint's position target by up to 0.04 radians.

Two policies take turns. The pick policy lifts the box. When the box has stayed within 5 cm of the gripper and 10 cm up for 0.2 s, the throw policy takes over and throws it toward the arrow, which points a new random way each time. A throw ends where the box first touches the floor. Its length is measured from the robot's base along the arrow.

Training

Both policies were trained with PPO, using Brax and MJX (the GPU version of MuJoCo), on a laptop with an RTX 5050. It simulated 2,048 copies of the scene at once, about 120,000 steps per second.

Measured

The policies were trained on MJX, but this page runs them on MuJoCo's regular CPU engine, which computes contacts differently. So I measured them with the same WebAssembly build and the same code as this page: 1,200 rounds, with a new box after each one.

Limits

Panda model: MuJoCo Menagerie (Apache 2.0). Task based on MuJoCo Playground's PandaPickCube (Apache 2.0). Colors sampled from Mary Mark Ockerbloom's photo Alphabet Blocks.