How it works
The arm is the Franka Emika Panda model from MuJoCo Menagerie, set up as in the PandaPickCube task of MuJoCo Playground. Your browser simulates it with MuJoCo compiled to WebAssembly, in 5 ms physics steps. Every 20 ms a policy, a small neural network, reads the joint angles and speeds and where the box is, and moves each joint's position target by up to 0.04 radians.
Two policies take turns. The pick policy lifts the box. When the box has stayed within 5 cm of the gripper and 10 cm up for 0.2 s, the throw policy takes over and throws it toward the arrow, which points a new random way each time. A throw ends where the box first touches the floor. Its length is measured from the robot's base along the arrow.
Training
Both policies were trained with PPO, using Brax and MJX (the GPU version of MuJoCo), on a laptop with an RTX 5050. It simulated 2,048 copies of the scene at once, about 120,000 steps per second.
- Picking: 124 million steps, about 20 minutes. This is Playground's task, changed so that the box can land on any side, is sometimes dropped from up to 25 cm, and can reappear somewhere else in the middle of an attempt. That is why the arm goes after a box you drop.
- Throwing: 213 million steps, about 30 minutes. Each attempt starts from one of 22,096 moments where the pick policy had just lifted the box. The reward is how far the box gets along a random direction before it first touches the floor, minus how far it lands to the side of that line. There were no demonstrations; the throwing motion came from the reward alone.
Measured
The policies were trained on MJX, but this page runs them on MuJoCo's regular CPU engine, which computes contacts differently. So I measured them with the same WebAssembly build and the same code as this page: 1,200 rounds, with a new box after each one.
- 1,182 boxes were picked up and thrown. 15 were not picked up within 12 s, and 3 were picked up but never thrown.
- Throws went 1.90 m on average (median 2.03 m, best 2.64 m). 59 of them (5%) were fumbles that dropped the box within 0.5 m of the base.
- The median landing spot was 5° off the arrow.
- In MJX, the same throw policy averages 2.05 m and almost never fumbles. Most of the difference is boxes slipping out of the gripper early.
Limits
- Playground's Panda model has no collisions between the robot's own parts. A penalty keeps the hand away from the base, and I have not seen the arm pass through itself, but nothing physically stops it.
- Position targets move at most 2 radians per second, close to the real Panda's joint speed limits of 2.2 to 2.6. Throws behind the robot go farther (median about 2.2 m) than throws in front of it (about 1.8 m).
- Sometimes the arm pushes the box out of reach instead of picking it up. After 1.5 s a new box drops.
Panda model: MuJoCo Menagerie (Apache 2.0). Task based on MuJoCo Playground's PandaPickCube (Apache 2.0). Colors sampled from Mary Mark Ockerbloom's photo Alphabet Blocks.