An observation-dependent reinforcement-learning agent that plays Atari 2600 Space Invaders directly from pixels. The five-million-step champion averaged 7,407.6 points across 100 unseen seeds—8.4× its 2.5-million-step parent.
Both models were evaluated on the same 100 unseen seeds using native Atari score.
| Metric | 5M champion | 2.5M parent |
|---|---|---|
| Mean | 7,407.6 | 886.7 |
| Median | 5,860.0 | 690.0 |
| Interquartile mean | 6,921.1 | 790.3 |
| 10th percentile | 2,060.0 | 570.0 |
The paired improvement was +6,520.9 points with a bootstrapped 95% confidence interval of [5,604.2, 7,436.3]. See the validation reports and model card for full metrics and limitations.
The final audit corrupts the policy's observations to detect memorized action loops.
| Observation | Mean score | Retained |
|---|---|---|
| Normal | 6,504.5 | 100% |
| Shuffled frame history | 593.3 | 9.1% |
| Frozen first frame | 270.0 | 4.2% |
| Blank screen | 0.0 | 0% |
The policy uses all six actions, no action exceeds 31.2% of decisions, and the audit produced no behavioral warnings. Its collapse under corrupted inputs is evidence that it reacts to gameplay instead of replaying a fixed sequence.
Atari RGB frames
│ max-pool, grayscale, resize, stack 4
▼
4 × 84 × 84 pixels
▼
IMPALA residual CNN + IQN dueling/noisy heads
▼
NOOP / FIRE / LEFT / RIGHT / LEFTFIRE / RIGHTFIRE
▼
prioritized 3-step replay + Munchausen IQN targets
The learner combines an IMPALA-style visual encoder, spectral normalization, Implicit Quantile Networks, dueling and noisy layers, prioritized replay, three-step returns, and Munchausen targets. It receives no emulator RAM, object coordinates, or privileged game state.
Requires Python 3.12 and a PyTorch/ALE-supported platform.
git clone https://github.com/Cp557/space-invaders-rl.git
cd space-invaders-rl
./setup.sh
./run.shThe default 5,000-step run tests the complete pipeline but is too short to produce a strong player.
| Command | Purpose |
|---|---|
./run.sh |
End-to-end smoke test |
./run.sh pilot |
One-million-step compact run |
./run.sh full |
Train a 2.5-million-step candidate |
./run.sh continue |
Add 2.5 million steps to a champion |
./run.sh validate |
Behavioral audit and 100-seed comparison |
./run.sh watch artifacts/model_champion.pt |
Record a checkpoint |
Later stages are gated by score, lower-tail performance, action diversity, and
observation dependence. The full strategy is documented in
docs/DESIGN.md.
Download the 99 MB model_champion.pt
checkpoint from the v1.0.0 GitHub release. Place it at
artifacts/model_champion.pt, then run:
./run.sh watch artifacts/model_champion.ptassets/ showcase GIF
docs/ design and experiment plan
results/ final reports and model card
scripts/cloud/ detached jobs with auto-shutdown
src/space_invaders_rl/ model, replay, training, and evaluation
tests/ unit and integration tests
Generated experiments, ROMs, logs, virtual environments, and checkpoints are
excluded by .gitignore.
ALE/SpaceInvaders-v5- Four 84×84 grayscale frames
- Four-frame action repeat and 25% sticky actions
- Random 0–30 no-op starts
- Five million training steps
- 100 paired validation seeds and 20 episodes per corruption condition
Atari ROMs are not included. Scores should only be compared under the same ALE configuration.
- Hessel et al., Rainbow
- Clark et al., Beyond The Rainbow
- MIT-licensed BTR reference implementation
- Gymnasium and ALE
This independent Gymnasium/PyTorch adaptation is released under the MIT License. Atari and Space Invaders belong to their respective owners.
