The problem
My bachelor's thesis asks how reinforcement-learning agents behave when they cannot observe the full state, and which training strategies best recover the missing information. The work extends that question beyond architecture alone to measure the effects of memory, intrinsic motivation, and curriculum learning.
Approach & tradeoffs
A modular, reproducible, multi-seed MiniGrid benchmark compares four algorithms:
- PPO and A2C as on-policy actor-critic baselines.
- DQN as a value-based point of comparison.
- RecurrentPPO to test explicit memory under partial observability.
The evaluation varies observability and training strategy while keeping runs config-driven and repeatable across seeds. This separates algorithm choice from the contribution of recurrent memory, intrinsic-reward signals, and a staged curriculum.
Results
Success rate at evaluation, mean over seeds with a bootstrap 95% confidence interval. The agent sees a 7×7 egocentric view unless noted; 5 seeds per configuration.
| Question | Setting | Success | 95% CI |
|---|---|---|---|
| Curriculum | DoorKey-8x8, staged from 5×5 | 0.948 | [0.90, 0.98] |
DoorKey-8x8, trained directly | 0.372 | [0.00, 0.76] | |
| Exploration | MultiRoom-N6, NovelD bonus | 0.856 | [0.80, 0.92] |
MultiRoom-N6, RND bonus | 0.028 | [0.01, 0.04] | |
MultiRoom-N6, no bonus, ICM or RIDE | 0.000 | [0.00, 0.00] | |
| Memory | MemoryS13Random, PPO, single frame | 0.540 | [0.47, 0.61] |
MemoryS13Random, RecurrentPPO (LSTM) | 0.476 | [0.40, 0.55] | |
MemoryS13Random, PPO, 4-frame stack | 0.440 | [0.38, 0.49] |
What the numbers say:
- Curriculum was the most reliable lever. Staging DoorKey from 5×5 to 8×8 reached 0.948 success; training on 8×8 directly averaged 0.372, with an interval spanning 0.00 to 0.76 — whether it learns depends on the seed.
- Exploration bonuses are not interchangeable. On MultiRoom-N6 only NovelD got off the ground; RND, ICM and RIDE stayed at or near zero, like plain PPO. A staged curriculum over room counts reached 0.892 with no bonus at all.
- Explicit memory did not pay off. Neither an LSTM nor frame stacking beat single-frame PPO on the memory task at this training budget.
- More of the grid is not more signal. Giving PPO the full allocentric grid on FourRooms dropped success from 0.653 to 0.007 (3 seeds): the state representation mattered more than how much of it the agent could see.