Learning to Walk: Reinforcement Learning Gait Controllers
A musculoskeletal model can show what forces act inside the body, but it cannot move on its own. It needs a controller that decides, moment by moment, how strongly each muscle should contract. We train these controllers with deep reinforcement learning. Neural-network controllers learn by trial and error in physics simulation to drive full-body models with up to 150 muscles, first imitating measured human motion and then walking and running on their own. Once trained, a controller can be placed in new conditions it was never shown, such as a different surface, body or task. This lets us ask "what if" questions that are difficult or unsafe to test in people.
How we use it
Walking on slippery ground. Controllers that used higher coactivation of opposing leg muscles walked more stably on slippery surfaces. Stable walking appears to require muscle coactivation beyond the minimum needed to save energy.
Treadmill versus overground walking. By removing psychological factors and keeping only physical ones, we showed that an ideal treadmill produces the same leg motion as overground walking. Differences appear when the treadmill belt cannot hold its speed against the forces of each step.
Rethinking standard gait analysis. By comparing forward dynamics simulation with conventional inverse dynamics on the same subjects, we found that inverse dynamics underestimates mechanical power. The main reason is the residual forces it needs to balance the equations.
Benchmarks. Our team placed second in the locomotion track of the NeurIPS 2023 MyoChallenge and won first place in 2024 among 54 teams from 15 countries, and we contributed to the MyoChallenge 2024 benchmark paper.
Neural Circuits of Locomotion: Supraspinal and Reflex Control
Walking is controlled at several levels of the nervous system. The brain sends predictive, long-latency commands that shape each step. Spinal reflexes react within tens of milliseconds to what the muscles and joints sense. In people, these pathways always work together, so it is hard to tell what each one contributes. In simulation we can separate them. We build neural controllers with distinct pathways, train them to walk a musculoskeletal model, and remove one pathway at a time in a "virtual ablation." This shows what each pathway does and why the nervous system needs it.
How we use it
What spinal reflexes are for. We compared four controller architectures: supraspinal only, reflex only, reflex with dynamic gain modulation, and combined supraspinal and reflex control. On level ground without disturbances, supraspinal control alone reproduced normal leg motion. When the model was pushed at the pelvis or the floor speed changed, controllers with reflexes stayed upright more often and used less muscle activation. Reflexes therefore act as a fast, effort-saving way to reject disturbances.
How reflex gains change across the gait cycle. When the combined controller was trained with disturbances, it learned to separate local reflex feedback from baseline drive. It then adjusted reflex gains across the gait cycle in a way that matches patterns measured in human experiments. This pattern appeared only after training with disturbances, which suggests that both reflex pathways and their modulation are needed for robust walking.
Muscle coactivation as a control strategy. Simulations on slippery ground showed that the nervous system may deliberately co-activate opposing muscles, spending extra energy to stiffen the joints and stabilize walking.
Muscle activation patterns. We are studying how lower-limb muscle activation patterns are organized during human movement, combining EMG measurements with simulation.
Understanding Human Movement Physiology
Why do we move the way we do? A widely accepted idea is that humans walk and run in ways that minimize metabolic energy. Yet people show behaviors that this idea alone cannot explain. We tense opposing leg muscles at the same time, which costs extra energy, and we swing our arms actively when running. These questions are hard to answer through experiments alone, because people cannot simply switch off a strategy or freeze a body part without changing everything else. In forward dynamics simulation we can. Musculoskeletal models driven by neural-network controllers let us turn individual strategies on and off and measure the consequences for stability and energy.
How we use it
Why we co-activate opposing muscles. A gait controller that used only the minimal muscle activation needed to walk often fell on slippery ground and uneven terrain. When controllers were trained on slippery ground, coactivation of lower-limb muscles emerged by itself. Controllers with physiological levels of coactivation fell much less often. Stable walking therefore appears to require muscle activity beyond the energy-optimal minimum, and both stability and energy efficiency shape how we control our muscles.
Why we swing our arms when running. We simulated running with active arm swing, passive arm swing and fixed arms using full-body models with 150 muscles. Active arm swing gave the least torso rotation and the lowest total metabolic cost: 5.52 J/kg-m, compared with 5.73 for passive swing and 5.82 for fixed arms. This holds even though the arm muscles themselves used more energy. Arm swing is an active, energy-saving strategy, not a passive side effect of leg motion.
Treadmills are widely used in rehabilitation and gait analysis. However, previous studies have reported differences in terms of kinematics and kinetics between treadmill and overground walking due to physical and psychological factors. The aim of this study was to analyze gait differences due to only the physical factors of treadmill walking. The study has been published in Frontiers in Bioengineering and Biotechnology. - by Mingi Jung