This Unitree G1 robot was trained to walk, and it keeps learning while it runs. At every control step (20 ms of simulated time) it chooses an action, observes the result, updates its two neural networks once, and discards the data.
Switch the floor to ice . The copied gait slips and the robot falls within a second. With learning at Max speed it walks again after 400,000 to over a million steps, depending on the run (about 1.5 to 4 minutes). The agent is not told that the floor changed. It detects the change only through larger prediction errors. With the injured leg or the backpack, it falls within a few seconds at first and walks again after 100,000 to 200,000 steps (under a minute).
Continual, streaming learning
Most deep reinforcement learning trains a policy on stored experience (a replay buffer, sampled in batches) and then fixes the weights. A fixed policy cannot adapt when the environment changes.
This agent learns continually, in streaming mode: each transition (state, action, reward, next state) is used for one update and then discarded. There is no replay buffer and no batching.
Continual learning has two known problems. New learning can overwrite old skills (catastrophic forgetting), and networks can lose the ability to learn over time (loss of plasticity) [2]. Streaming updates are also noisy, which made streaming deep RL unstable. Elsayed, Vasan and Mahmood describe a method that makes it stable [1]. This page implements that method.
The method
The agent is an actor-critic. The critic \(\hat v\) estimates future reward. The actor \(\pi\) outputs target angles for the robot's 12 leg joints, which its motors then track. The waist and arms hold still. After each step, the TD error \(\delta\) measures the critic's error:
$$\delta_t = R_{t+1} + \gamma\,\hat v(S_{t+1}) - \hat v(S_t)$$An eligibility trace \(z\) is a decaying sum of recent gradients. It lets each TD error update the weights for the last few dozen steps without storing those steps:
$$z \leftarrow \gamma\lambda\, z + \nabla_{w}\hat v(S_t)$$The optimizer divides each weight's update by the largest recent value of \(|\delta z|\) for that weight. This limits every weight change to at most \(\alpha\) per step:
$$v \leftarrow \max\!\big(\beta v,\ |\delta z|\big), \qquad w \leftarrow w + \alpha\,\frac{\delta z}{v}$$The optimizer push readout shows \(|\delta z| / v\) averaged over all weights. It rises when the environment changes. The method also normalizes observations and rewards with running statistics, applies LayerNorm in both networks, and starts each network with 95% of its weights at zero.
Where the walk comes from
Learning to walk from scratch with this method takes millions of steps and produces an odd gait, so this actor did not start from random weights. We copied Unitree's published walking controller for the G1 [3] into the actor network. The copy drove the robot, Unitree's controller labeled each state it visited with its own action, and the copy was refit on all the labels so far, eight times over (DAgger [4]). To track the copied gait, the actor also observes a gait clock: the phase of a 0.8 s stepping cycle.
Then streaming learning took over: 300,000 steps that trained only the critic, so its first TD errors were meaningful, then 3 million steps of the full method on a normal floor. From there on, every change in the gait comes from streaming learning.
The experiment
Dohare et al. tested continual learning by switching a simulated ant's floor between high and low friction [2]. We ran a similar test with this agent: ice (friction 0.15), normal floor, injured right leg (hip, knee and ankle motors limited to 20% of their torque), normal floor, then a 15 kg backpack. Each condition lasted 800,000 steps, with learning on throughout.
With learning turned off, the same agent finishes 99% of its episodes on the normal floor. In each changed environment it falls within a few seconds: after 0.7 s on ice, 3.4 s with the injured leg, and 5.7 s with the backpack (averages over at least 175 episodes).