Skip to content

About

A reinforcement-learning agent that plays Atari 2600 Space Invaders.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Space Invaders RL

Python 3.12 PyTorch Gymnasium License: MIT

An observation-dependent reinforcement-learning agent that plays Atari 2600 Space Invaders directly from pixels. The five-million-step champion averaged 7,407.6 points across 100 unseen seeds—8.4× its 2.5-million-step parent.

The trained agent playing Space Invaders

Results

Both models were evaluated on the same 100 unseen seeds using native Atari score.

Metric 5M champion 2.5M parent
Mean 7,407.6 886.7
Median 5,860.0 690.0
Interquartile mean 6,921.1 790.3
10th percentile 2,060.0 570.0

The paired improvement was +6,520.9 points with a bootstrapped 95% confidence interval of [5,604.2, 7,436.3]. See the validation reports and model card for full metrics and limitations.

Does it actually look at the screen?

The final audit corrupts the policy's observations to detect memorized action loops.

Observation Mean score Retained
Normal 6,504.5 100%
Shuffled frame history 593.3 9.1%
Frozen first frame 270.0 4.2%
Blank screen 0.0 0%

The policy uses all six actions, no action exceeds 31.2% of decisions, and the audit produced no behavioral warnings. Its collapse under corrupted inputs is evidence that it reacts to gameplay instead of replaying a fixed sequence.

How it works

Atari RGB frames
      │  max-pool, grayscale, resize, stack 4
      ▼
  4 × 84 × 84 pixels
      ▼
IMPALA residual CNN + IQN dueling/noisy heads
      ▼
NOOP / FIRE / LEFT / RIGHT / LEFTFIRE / RIGHTFIRE
      ▼
prioritized 3-step replay + Munchausen IQN targets

The learner combines an IMPALA-style visual encoder, spectral normalization, Implicit Quantile Networks, dueling and noisy layers, prioritized replay, three-step returns, and Munchausen targets. It receives no emulator RAM, object coordinates, or privileged game state.

Quick start

Requires Python 3.12 and a PyTorch/ALE-supported platform.

git clone https://github.com/Cp557/space-invaders-rl.git
cd space-invaders-rl
./setup.sh
./run.sh

The default 5,000-step run tests the complete pipeline but is too short to produce a strong player.

Command Purpose
./run.sh End-to-end smoke test
./run.sh pilot One-million-step compact run
./run.sh full Train a 2.5-million-step candidate
./run.sh continue Add 2.5 million steps to a champion
./run.sh validate Behavioral audit and 100-seed comparison
./run.sh watch artifacts/model_champion.pt Record a checkpoint

Later stages are gated by score, lower-tail performance, action diversity, and observation dependence. The full strategy is documented in docs/DESIGN.md.

Champion checkpoint

Download the 99 MB model_champion.pt checkpoint from the v1.0.0 GitHub release. Place it at artifacts/model_champion.pt, then run:

./run.sh watch artifacts/model_champion.pt

Project structure

assets/                  showcase GIF
docs/                    design and experiment plan
results/                 final reports and model card
scripts/cloud/           detached jobs with auto-shutdown
src/space_invaders_rl/   model, replay, training, and evaluation
tests/                    unit and integration tests

Generated experiments, ROMs, logs, virtual environments, and checkpoints are excluded by .gitignore.

Evaluation protocol

  • ALE/SpaceInvaders-v5
  • Four 84×84 grayscale frames
  • Four-frame action repeat and 25% sticky actions
  • Random 0–30 no-op starts
  • Five million training steps
  • 100 paired validation seeds and 20 episodes per corruption condition

Atari ROMs are not included. Scores should only be compared under the same ALE configuration.

References

This independent Gymnasium/PyTorch adaptation is released under the MIT License. Atari and Space Invaders belong to their respective owners.

About

A reinforcement-learning agent that plays Atari 2600 Space Invaders.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages