Learning on Ice

A simulated humanoid robot that learns continuously. Change its environment (ice, an injured leg, a 15 kg backpack) and it adapts while you watch.

Normal floor
Loading the physics engine and the pretrained agent…
World Speed
Return per episode each episode20-episode average
Steps learned
0
Speed
—
Average return
—
Optimizer push
—

Published

October 5, 2026

Robot

Unitree G1 in MuJoCo

Learner

Stream-AC [1], running in your browser

This Unitree G1 robot was trained to walk, and it keeps learning while it runs. At every control step (20 ms of simulated time) it chooses an action, observes the result, updates its two neural networks once, and discards the data.

Switch the floor to ice . The copied gait slips and the robot falls within a second. With learning at Max speed it walks again after 400,000 to over a million steps, depending on the run (about 1.5 to 4 minutes). The agent is not told that the floor changed. It detects the change only through larger prediction errors. With the injured leg or the backpack, it falls within a few seconds at first and walks again after 100,000 to 200,000 steps (under a minute).

Continual, streaming learning

Most deep reinforcement learning trains a policy on stored experience (a replay buffer, sampled in batches) and then fixes the weights. A fixed policy cannot adapt when the environment changes.

This agent learns continually, in streaming mode: each transition (state, action, reward, next state) is used for one update and then discarded. There is no replay buffer and no batching.

Continual learning has two known problems. New learning can overwrite old skills (catastrophic forgetting), and networks can lose the ability to learn over time (loss of plasticity) [2]. Streaming updates are also noisy, which made streaming deep RL unstable. Elsayed, Vasan and Mahmood describe a method that makes it stable [1]. This page implements that method.

The method

The agent is an actor-critic. The critic \(\hat v\) estimates future reward. The actor \(\pi\) outputs target angles for the robot's 12 leg joints, which its motors then track. The waist and arms hold still. After each step, the TD error \(\delta\) measures the critic's error:

$$\delta_t = R_{t+1} + \gamma\,\hat v(S_{t+1}) - \hat v(S_t)$$

An eligibility trace \(z\) is a decaying sum of recent gradients. It lets each TD error update the weights for the last few dozen steps without storing those steps:

$$z \leftarrow \gamma\lambda\, z + \nabla_{w}\hat v(S_t)$$

The optimizer divides each weight's update by the largest recent value of \(|\delta z|\) for that weight. This limits every weight change to at most \(\alpha\) per step:

$$v \leftarrow \max\!\big(\beta v,\ |\delta z|\big), \qquad w \leftarrow w + \alpha\,\frac{\delta z}{v}$$

The optimizer push readout shows \(|\delta z| / v\) averaged over all weights. It rises when the environment changes. The method also normalizes observations and rewards with running statistics, applies LayerNorm in both networks, and starts each network with 95% of its weights at zero.

Where the walk comes from

Learning to walk from scratch with this method takes millions of steps and produces an odd gait, so this actor did not start from random weights. We copied Unitree's published walking controller for the G1 [3] into the actor network. The copy drove the robot, Unitree's controller labeled each state it visited with its own action, and the copy was refit on all the labels so far, eight times over (DAgger [4]). To track the copied gait, the actor also observes a gait clock: the phase of a 0.8 s stepping cycle.

Then streaming learning took over: 300,000 steps that trained only the critic, so its first TD errors were meaningful, then 3 million steps of the full method on a normal floor. From there on, every change in the gait comes from streaming learning.

The experiment

Dohare et al. tested continual learning by switching a simulated ant's floor between high and low friction [2]. We ran a similar test with this agent: ice (friction 0.15), normal floor, injured right leg (hip, knee and ankle motors limited to 20% of their torque), normal floor, then a 15 kg backpack. Each condition lasted 800,000 steps, with learning on throughout.

Return per episode during the tour (one run; we fixed its random seed before recording it). A full 20 s episode scores about 2,900. On ice the return starts near 100 and passes 2,500 after about 380,000 steps. By the end of the ice phase the robot finishes 84% of its episodes. Back on the normal floor it walks as before, so learning on ice did not overwrite the old gait. The injured leg causes a short dip, and the backpack a deep one that recovers in about 120,000 steps. Every condition ends with the robot walking at 0.77 to 0.80 m/s.

With learning turned off, the same agent finishes 99% of its episodes on the normal floor. In each changed environment it falls within a few seconds: after 0.7 s on ice, 3.4 s with the injured leg, and 5.7 s with the backpack (averages over at least 175 episodes).