A continuous-control PPO agent for Gymnasium's LunarLander-v3. The final
policy landed both legs safely inside the pad on 98.3% of 20,000 unseen
episodes.
Deterministic policy on an unseen evaluation seed.
| Evaluation | Strict landings | Rate | Crashes | Timeouts |
|---|---|---|---|---|
| Promotion test | 4,919 / 5,000 | 98.38% | 0 | 0 |
| Audit A | 9,838 / 10,000 | 98.38% | 2 | 0 |
| Audit B | 9,823 / 10,000 | 98.23% | 0 | 0 |
| Combined audits | 19,661 / 20,000 | 98.31% | 2 | 0 |
On the same 10,000-episode audit, the final contact-aware policy improved the previous model as follows:
| Outcome | Previous model | Final model |
|---|---|---|
| Strict landing | 9,632 | 9,838 |
| Incomplete contact | 347 | 157 |
| Outside the pad | 17 | 3 |
| Crash | 4 | 2 |
A strict landing requires all three conditions:
- The lander stops safely without crashing.
- Both legs maintain ground contact.
- Both legs are between the two pad flags.
This is stricter than relying on Gymnasium's reward alone. A one-leg landing or a safe stop outside the flags is counted as a failure.
Use Python 3.12 from the repository directory:
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txtWatch the included trained model:
python watch.py
python watch.py --episodes 3 --seed 900000Evaluate it without rendering:
python evaluate.py
python evaluate.py --episodes 10000 --seed 900000The scripts load models/ppo_continuous.zip by default.
The policy receives the standard eight-value LunarLander observation: position, velocity, angle, angular velocity, and both leg-contact states. It outputs two continuous controls for the main engine and lateral thrusters.
PPO improves the policy by collecting trajectories, estimating which actions performed better than expected, and making limited updates so one training batch cannot change the policy too aggressively.
| Problem | Addition |
|---|---|
| Reward did not guarantee a valid pad landing | Strict terminal classifier requiring both legs inside the flags |
| The lander settled near or outside the pad | Phase-aware body and leg-placement shaping |
| Fast or tilted touchdowns | Near-ground speed, angle, and rotation penalties |
| Abrupt engine changes | Action-change smoothing |
| Rare failures disappeared in random training | Failure-seed mining and focused replay |
| The policy balanced indefinitely on one leg | Second-leg contact shaping and a small one-leg hold penalty |
| PPO varied between runs | Eight candidates across four reward presets |
| A lucky validation score could win | Rotating validation blocks and a separate promotion test |
train.py repeats the final fine-tuning experiment from the included checkpoint:
python train.pyThe full run:
- Protects the current checkpoint.
- Audits it on 10,000 episodes.
- Mines failures from 5,000 training-only seeds.
- Assigns 70% of focused replay to incomplete-contact failures.
- Trains eight candidates for up to 400,000 steps each.
- Rotates across four disjoint 500-episode validation blocks.
- Tests the top four checkpoints on 5,000 separate promotion episodes.
- Promotes a checkpoint only when it beats the protected model.
All candidates keep the same PPO and base reward settings. Only contact-shaping strength changes:
| Preset | Contact shaping |
|---|---|
baseline |
0.00 |
contact-light |
0.50 |
contact |
1.00 |
contact-strong |
1.50 |
contact-light won. Stronger shaping often made the policy slower and less
stable, showing that more reward pressure was not automatically better.
Use python train.py --help to reduce the seeds, timesteps, or evaluation sizes
for a shorter experiment.
Each point is a deterministic evaluation on one rotating 500-episode validation block. The dashed line marks the 98% target. Colors identify contact-shaping strength; solid and dashed lines are separate training seeds.
Validation is intentionally noisy because adjacent checkpoints use different
seed blocks. The promoted seed 5 contact-light checkpoint was selected by the
separate 5,000-episode promotion test, not by this chart alone. The duration
panel shows the tradeoff: careful settling took longer, while excessive contact
shaping produced long episodes without consistently improving success.
Regenerate the chart:
python plot_training.pyRegenerate the animated demo:
python record.py --seed 900000| Path | Purpose |
|---|---|
models/ppo_continuous.zip |
Verified PPO checkpoint |
results/ |
Training history, failure telemetry, metadata, and audits |
assets/ |
Animated demo and training chart |
landing_env.py |
Strict success rules and reward shaping |
train.py |
Failure mining, fine-tuning, validation, and promotion |
evaluate.py |
Deterministic strict evaluation |
watch.py |
Interactive rendering |
record.py |
Reproducible GIF recording |
plot_training.py |
Training chart generation |
The exact winning configuration and promotion results are stored in
results/training_metadata.json. Raw checkpoint evaluations are in
results/training_evaluations.csv.

