Air Hockey RL Agent
Payoff-matrix PSRO + sampled league self-play
Move your paddle with the pointer.
Training Methodology
The agent was initialized from the deterministic incumbent, then trained from randomized sides with stochastic PPO best responses against a payoff-matrix league. Prioritized fictitious self-play emphasizes both the approximate Nash mixture and difficult opponents, while dedicated exploiters expose weaknesses for later rounds. The web opponent samples from three frozen league policies after each goal, and self-play assigns a different sampled policy to each side.
Observation Space (12 features)
- Own paddle: position (x, y), velocity (dx, dy)
- Puck: position (x, y), velocity (dx, dy)
- Opponent paddle: position (x, y), velocity (dx, dy)
Results
- 72–28 with 100 draws against the incumbent over 200 side-balanced games
- 61.0% match score, with a 95% interval of 56.3–65.7%
- 66.3% against random and 40.6% against the scripted baseline, versus the incumbent's 63.1% and 25.0%
- 16/24 bank shots saved, matching the incumbent while passing all six adversarial promotion gates
- Wall-clipped paddle velocity, side-wall wedges, and scoreless self-play rounds now have explicit recovery behavior
Code
Training Script | Environment | Adversarial Evaluation | Web Inference