A comprehensive deep reinforcement learning framework for training AI agents to play Settlers of Catan. Features DQN, PPO, and SAC algorithms with a custom game engine, interactive web interface, and multi-agent training infrastructure.
The system implements a complete Settlers of Catan game engine with support for all standard rules including resource production, trading, development cards, the robber mechanic, and victory point tracking. The engine is designed for high-throughput simulation to enable efficient RL training.
Multiple reinforcement learning algorithms are implemented for comparative analysis and to handle Catan's unique challenges: large action spaces, partial observability, and strategic long-term planning.
Catan has 1,017 possible discrete actions: 54 settlement locations, 54 city upgrades, 72 road edges, 19 robber positions, plus trading combinations and development cards. Most actions are illegal at any given time, causing severe exploration inefficiency.
Implemented a legal action mask that filters the policy output before sampling. The game engine computes valid actions each turn, and the mask is applied to Q-values (DQN) or logits (PPO/SAC) to prevent illegal moves. This reduced wasted exploration by ~95%.
# Action masking in forward pass
q_values = self.network(state)
q_values[~legal_mask] = float('-inf') # Mask illegal actions
action = q_values.argmax()
Settlements must be at least 2 edges apart (the "distance rule"). Early agents would repeatedly attempt to build adjacent to existing settlements, triggering rule violations. Debug scripts revealed the connectivity check wasn't accounting for opponent structures.
Rewrote the settlement validation using proper graph traversal. The board topology is represented as a vertex-edge graph, and BFS checks that no settlement exists within 2 hops of the target vertex, including opponent buildings.
The full game state includes hexagonal board topology, per-player resources, building positions, development cards, and ports—far too complex for naive flattening. Initial encodings caused the network to overfit on irrelevant features.
Designed a compact 90-dimensional state vector with semantic groupings: resources (5), buildings (8), strategic position (10), building readiness (12), opponent comparison (35), and game state (10). Features capture relative advantage rather than absolute positions.
Games last 50-100+ turns with victory (+1) or loss (-1) only at the end. Agents struggled to learn meaningful policies with such delayed feedback—random play achieved similar early performance.
Added intermediate rewards for strategic milestones: +0.1 for each victory point gained, +0.05 for completing a road segment, +0.02 for resource diversity. Negative shaping (-0.01) for hoarding resources without building. This provided gradient signal throughout the game.
Catan presents a challenging discrete action space with variable legality. The system uses action masking to ensure only legal moves are considered during training and inference.
# Action space breakdown
Pass turn: 1 action
Roll dice: 1 action
Build settlement: 54 actions (one per vertex)
Build city: 54 actions (upgrade existing)
Build road: 72 actions (one per edge)
Buy dev card: 1 action
Play dev card: 4 actions (knight, monopoly, year of plenty, road building)
Move robber: 19 actions (one per tile)
Steal resource: ~12 actions (varies by adjacent players)
Bank trades: ~200 actions (resource combinations)
Discard: ~600 actions (7+ card combinations)
─────────────────────────────
Total: 1,017 discrete actions
A FastAPI-powered web interface allows humans to play against trained agents or watch agent vs agent matches. The interface renders the game board using SVG with real-time state updates and supports all game actions through an intuitive UI.
Play against trained agents
Watch agents compete
Real-time game metrics
After training against random and self-play opponents, agents achieve competitive performance. SAC slightly outperforms PPO due to better exploration in the large action space.