Air Hockey RL Agent
Imitation, then league self-play on a GPU
Move your paddle with the mouse or a finger.
The game
The puck slides with little friction, bounces off the walls, and reflects off the paddles the way it does off a real mallet: the faster the paddle moves into it, the harder it comes back. Both paddles have the same top speed, and your paddle follows your mouse or finger through the same smoothing as the AI's. The physics runs at a fixed 60 steps per second.
How it was trained
The browser physics was ported to JAX and checked against this page frame by frame, then trained on one laptop GPU at about 600,000 game steps per second.
- Copy an expert. A scripted player predicts the puck's path, blocks, and lines up shots. The network first learns to imitate it (DAgger). Starting from scratch, the agent learned to hide in a corner, because random touches score more own goals than goals.
- League self-play. PPO then trains the network against its current self, earlier snapshots (picked more often when they beat it), an exploiter trained only to beat it, and scripted players with human reaction delays.
What it sees (12 numbers)
- Its own paddle: position and velocity
- The puck: position and velocity
- Your paddle: position and velocity
Results
Measured in 5-minute matches with the same rules as this page.
- Against human-limited networks: 37 to 1. These were trained the same way, specifically to beat this policy, but they see the puck 200 ms late (a typical human reaction time) and their aim wobbles. They score 0.2 goals a minute to its 7.4. With a quicker 133 ms reaction and steadier aim it is still 9 to 1.
- Against a scripted expert with human reaction time: 15.4 goals a minute for, 0.15 against. Against the same expert with no delay, 26 to 0.5.
- It is not unbeatable. A network trained only to beat it, with no handicap, wins 49% of points from random positions and loses 32%.
Code
Physics and scripted players | Imitation | League training | Evaluation | Web inference