Back to Projects
Catan RL Training System

Catan Multi-Agent RL Training System

A comprehensive deep reinforcement learning framework for training AI agents to play Settlers of Catan. Features DQN, PPO, and SAC algorithms with a custom game engine, interactive web interface, and multi-agent training infrastructure.

Completed: December 2025 Reinforcement Learning & Game AI
Python PyTorch Ray RLlib FastAPI
1,017
Discrete Actions
90-dim
State Encoding
928 LOC
Rules Engine
3 Algorithms
DQN, PPO, SAC
Multi-Agent RL Architecture

Technical Overview

Core Architecture

The system implements a complete Settlers of Catan game engine with support for all standard rules including resource production, trading, development cards, the robber mechanic, and victory point tracking. The engine is designed for high-throughput simulation to enable efficient RL training.

Key Components:

  • Immutable game state representation for safe multi-agent rollouts
  • Efficient action masking for legal move enforcement
  • Vectorized state encoding for neural network input
  • Gymnasium-compatible environment interface

RL Algorithms

Multiple reinforcement learning algorithms are implemented for comparative analysis and to handle Catan's unique challenges: large action spaces, partial observability, and strategic long-term planning.

Implemented Algorithms:

  • DQN: Deep Q-Network with experience replay and target networks
  • PPO: Proximal Policy Optimization via Ray RLlib
  • SAC: Soft Actor-Critic for continuous action refinement
  • LLM Agent: Experimental GPT-based strategic reasoning
Training Metrics - Learning Curves, Win Rates, Loss Convergence

Challenges & Solutions

Challenge: Massive Action Space

Catan has 1,017 possible discrete actions: 54 settlement locations, 54 city upgrades, 72 road edges, 19 robber positions, plus trading combinations and development cards. Most actions are illegal at any given time, causing severe exploration inefficiency.

Solution: Dynamic Action Masking

Implemented a legal action mask that filters the policy output before sampling. The game engine computes valid actions each turn, and the mask is applied to Q-values (DQN) or logits (PPO/SAC) to prevent illegal moves. This reduced wasted exploration by ~95%.

# Action masking in forward pass
q_values = self.network(state)
q_values[~legal_mask] = float('-inf')  # Mask illegal actions
action = q_values.argmax()

Challenge: Settlement Distance Rule

Settlements must be at least 2 edges apart (the "distance rule"). Early agents would repeatedly attempt to build adjacent to existing settlements, triggering rule violations. Debug scripts revealed the connectivity check wasn't accounting for opponent structures.

Solution: Graph-Based Validation

Rewrote the settlement validation using proper graph traversal. The board topology is represented as a vertex-edge graph, and BFS checks that no settlement exists within 2 hops of the target vertex, including opponent buildings.

Challenge: State Representation

The full game state includes hexagonal board topology, per-player resources, building positions, development cards, and ports—far too complex for naive flattening. Initial encodings caused the network to overfit on irrelevant features.

Solution: 90-Dimensional Feature Engineering

Designed a compact 90-dimensional state vector with semantic groupings: resources (5), buildings (8), strategic position (10), building readiness (12), opponent comparison (35), and game state (10). Features capture relative advantage rather than absolute positions.

Challenge: Sparse Rewards & Long Horizons

Games last 50-100+ turns with victory (+1) or loss (-1) only at the end. Agents struggled to learn meaningful policies with such delayed feedback—random play achieved similar early performance.

Solution: Reward Shaping

Added intermediate rewards for strategic milestones: +0.1 for each victory point gained, +0.05 for completing a road segment, +0.02 for resource diversity. Negative shaping (-0.01) for hoarding resources without building. This provided gradient signal throughout the game.

State Encoder: 90-Dimensional Feature Space

Action Space Design

1,017 Discrete Actions

Catan presents a challenging discrete action space with variable legality. The system uses action masking to ensure only legal moves are considered during training and inference.


# Action space breakdown
Pass turn:           1 action
Roll dice:           1 action
Build settlement:   54 actions (one per vertex)
Build city:         54 actions (upgrade existing)
Build road:         72 actions (one per edge)
Buy dev card:        1 action
Play dev card:       4 actions (knight, monopoly, year of plenty, road building)
Move robber:        19 actions (one per tile)
Steal resource:     ~12 actions (varies by adjacent players)
Bank trades:       ~200 actions (resource combinations)
Discard:           ~600 actions (7+ card combinations)
─────────────────────────────
Total:            1,017 discrete actions
           

Game Engine Features

  • Complete Catan rules (928 lines of game logic)
  • Immutable state for safe parallel rollouts
  • Efficient legal action computation
  • Support for 2-4 players
  • All development card types
  • Port trading (3:1 and 2:1 ratios)

Training Infrastructure

  • Self-play with population-based training
  • Curriculum learning from random to skilled opponents
  • Reward shaping for intermediate goals
  • Checkpoint saving and tournament evaluation
  • TensorBoard integration for metrics
  • Reproducible experiments via config files

Interactive Web Interface

A FastAPI-powered web interface allows humans to play against trained agents or watch agent vs agent matches. The interface renders the game board using SVG with real-time state updates and supports all game actions through an intuitive UI.

Web Interface Screenshot

Human vs AI

Play against trained agents

AI vs AI

Watch agents compete

Live Stats

Real-time game metrics

Results & Performance

Training Performance

After training against random and self-play opponents, agents achieve competitive performance. SAC slightly outperforms PPO due to better exploration in the large action space.

  • SAC win rate vs random: ~28% (4-player game baseline: 25%)
  • PPO win rate vs random: ~23% after 5K episodes
  • Policy loss converges within 2K episodes
  • Learns trading and development card strategies

Technical Achievements

  • Complete Catan rules implementation with all edge cases
  • Modular design allowing easy algorithm swapping
  • 90-dimensional feature encoder with semantic groupings
  • 1,017-action space with efficient masking
  • Interactive web interface for human-AI play
  • Comprehensive test suite (game logic verification)

Future Development

Algorithm Improvements

  • Implement AlphaZero-style MCTS + neural network
  • Add transformer-based policy networks
  • Explore multi-agent communication protocols
  • Implement opponent modeling

Feature Additions

  • Support for Catan expansions (Seafarers, Cities & Knights)
  • LLM agent for natural language negotiation
  • Online multiplayer with ELO rating system
  • Mobile-friendly web interface