Skip to content

About

Continuous-control PPO Lunar Lander with a verified 98.3% strict landing rate

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

PPO Lunar Lander

A continuous-control PPO agent for Gymnasium's LunarLander-v3. The final policy landed both legs safely inside the pad on 98.3% of 20,000 unseen episodes.

PPO agent landing the LunarLander between both flags

Deterministic policy on an unseen evaluation seed.

Results

Evaluation Strict landings Rate Crashes Timeouts
Promotion test 4,919 / 5,000 98.38% 0 0
Audit A 9,838 / 10,000 98.38% 2 0
Audit B 9,823 / 10,000 98.23% 0 0
Combined audits 19,661 / 20,000 98.31% 2 0

On the same 10,000-episode audit, the final contact-aware policy improved the previous model as follows:

Outcome Previous model Final model
Strict landing 9,632 9,838
Incomplete contact 347 157
Outside the pad 17 3
Crash 4 2

What counts as a landing?

A strict landing requires all three conditions:

  1. The lander stops safely without crashing.
  2. Both legs maintain ground contact.
  3. Both legs are between the two pad flags.

This is stricter than relying on Gymnasium's reward alone. A one-leg landing or a safe stop outside the flags is counted as a failure.

Quick Start

Use Python 3.12 from the repository directory:

python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

Watch the included trained model:

python watch.py
python watch.py --episodes 3 --seed 900000

Evaluate it without rendering:

python evaluate.py
python evaluate.py --episodes 10000 --seed 900000

The scripts load models/ppo_continuous.zip by default.

How It Works

The policy receives the standard eight-value LunarLander observation: position, velocity, angle, angular velocity, and both leg-contact states. It outputs two continuous controls for the main engine and lateral thrusters.

PPO improves the policy by collecting trajectories, estimating which actions performed better than expected, and making limited updates so one training batch cannot change the policy too aggressively.

Beyond basic PPO

Problem Addition
Reward did not guarantee a valid pad landing Strict terminal classifier requiring both legs inside the flags
The lander settled near or outside the pad Phase-aware body and leg-placement shaping
Fast or tilted touchdowns Near-ground speed, angle, and rotation penalties
Abrupt engine changes Action-change smoothing
Rare failures disappeared in random training Failure-seed mining and focused replay
The policy balanced indefinitely on one leg Second-leg contact shaping and a small one-leg hold penalty
PPO varied between runs Eight candidates across four reward presets
A lucky validation score could win Rotating validation blocks and a separate promotion test

Training

train.py repeats the final fine-tuning experiment from the included checkpoint:

python train.py

The full run:

  1. Protects the current checkpoint.
  2. Audits it on 10,000 episodes.
  3. Mines failures from 5,000 training-only seeds.
  4. Assigns 70% of focused replay to incomplete-contact failures.
  5. Trains eight candidates for up to 400,000 steps each.
  6. Rotates across four disjoint 500-episode validation blocks.
  7. Tests the top four checkpoints on 5,000 separate promotion episodes.
  8. Promotes a checkpoint only when it beats the protected model.

All candidates keep the same PPO and base reward settings. Only contact-shaping strength changes:

Preset Contact shaping
baseline 0.00
contact-light 0.50
contact 1.00
contact-strong 1.50

contact-light won. Stronger shaping often made the policy slower and less stable, showing that more reward pressure was not automatically better.

Use python train.py --help to reduce the seeds, timesteps, or evaluation sizes for a shorter experiment.

Training Progress

PPO fine-tuning validation success and episode duration

Each point is a deterministic evaluation on one rotating 500-episode validation block. The dashed line marks the 98% target. Colors identify contact-shaping strength; solid and dashed lines are separate training seeds.

Validation is intentionally noisy because adjacent checkpoints use different seed blocks. The promoted seed 5 contact-light checkpoint was selected by the separate 5,000-episode promotion test, not by this chart alone. The duration panel shows the tradeoff: careful settling took longer, while excessive contact shaping produced long episodes without consistently improving success.

Regenerate the chart:

python plot_training.py

Regenerate the animated demo:

python record.py --seed 900000

Repository Layout

Path Purpose
models/ppo_continuous.zip Verified PPO checkpoint
results/ Training history, failure telemetry, metadata, and audits
assets/ Animated demo and training chart
landing_env.py Strict success rules and reward shaping
train.py Failure mining, fine-tuning, validation, and promotion
evaluate.py Deterministic strict evaluation
watch.py Interactive rendering
record.py Reproducible GIF recording
plot_training.py Training chart generation

The exact winning configuration and promotion results are stored in results/training_metadata.json. Raw checkpoint evaluations are in results/training_evaluations.csv.

License

MIT

About

Continuous-control PPO Lunar Lander with a verified 98.3% strict landing rate

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages