← NMRIL Labs
🐿 SIRL — Interactive GridWorld Simulator
DQN
PPO
SIRL
All
▶ Play
›Step
↺
Speed
6×
🌱 Seed (Real Experiment Data)
Seed 42
Seed 43
Seed 44
partial
🖌 Paint Tool
🏁
Goal
+50 reward, ends ep.
💀
Fatal Zone
−10 reward, ends ep.
⚠️
Risky Cell
50%: +0.5 or −5.0
🤖
Start
Agent spawn point
⬜
Erase to Safe
+0.15 per step
⚙️ Grid Size
5×5
10×10 ✓
15×15
🗃 Presets
Default (paper setup)
⚡ Force Risky Path (key demo!)
Corner goal challenge
Maze-like layout
Dense fatal zones
Easy: No obstacles
🔧 Actions
Clear all cells
🎲 Randomize layout
Tip:
Click or drag on the grid to paint. Switch tools above. Try moving the 🏁 goal to a corner blocked by 💀 zones!
DQN — click ▶ Play to run | click grid to edit
0
Step
0.00
Reward
0.00
Penalty
—
Status
Risky cells visited per algorithm
📊 Algorithm Results
DQN
idle
Deep Q-Network with experience replay. No risk awareness — learns from reward signal alone. May pass through risky cells.
Steps:
—
Rew:
—
Path:
—
PPO
idle
On-policy actor-critic. Uses short rollouts — struggles on large grids with sparse rewards. Often wanders without reaching goal.
Steps:
—
Rew:
—
Path:
—
SIRL
idle
Squirrel-Inspired RL with RiskNet threat predictor. Penalises dangerous cells and routes around them. Safest path to goal.
Steps:
—
Rew:
—
Path:
—
📋 Step Log
Edit the grid, then press ▶ Play...
🧠 Algorithm Behaviour
DQN
— Cost: fatal=∞,
risky=1.5
, safe=1. Takes the shortest path even through risky zones.
PPO
— Random walk with weak goal bias. Fails on large grids.
SIRL
— Cost: fatal=∞,
risky=20
, safe=1. Pays the extra steps to stay safe.
🔑 Why DQN ≈ SIRL on Simple Grids?
⚠ This is expected!
On simple grids,
both reach the goal
— the real difference is:
Penalty
— DQN walks through risky cells; SIRL avoids them → 5× less penalty
Failures during training
— DQN dies many times before learning; SIRL avoids early
Try preset:
"Force Risky Path" to make the difference obvious!
📊 Penalty Comparison
🧠 Algorithm Behaviour