Air Hockey RL Agent

Imitation, then league self-play on a GPU

Move your paddle with the mouse or a finger.

The game

The puck slides with little friction, bounces off the walls, and reflects off the paddles the way it does off a real mallet: the faster the paddle moves into it, the harder it comes back. Both paddles have the same top speed, and your paddle follows your mouse or finger through the same smoothing as the AI's. The physics runs at a fixed 60 steps per second.

How it was trained

The browser physics was ported to JAX and checked against this page frame by frame, then trained on one laptop GPU at about 600,000 game steps per second.

  1. Copy an expert. A scripted player predicts the puck's path, blocks, and lines up shots. The network first learns to imitate it (DAgger). Starting from scratch, the agent learned to hide in a corner, because random touches score more own goals than goals.
  2. League self-play. PPO then trains the network against its current self, earlier snapshots (picked more often when they beat it), an exploiter trained only to beat it, and scripted players with human reaction delays.

What it sees (12 numbers)

Results

Measured in 5-minute matches with the same rules as this page.

Code

Physics and scripted players | Imitation | League training | Evaluation | Web inference