← NMRIL Labs

🐿 SIRL — Interactive GridWorld Simulator

Speed
🌱 Seed (Real Experiment Data)
🖌 Paint Tool
🏁
Goal
+50 reward, ends ep.
💀
Fatal Zone
−10 reward, ends ep.
⚠️
Risky Cell
50%: +0.5 or −5.0
🤖
Start
Agent spawn point
Erase to Safe
+0.15 per step
⚙️ Grid Size
🗃 Presets
🔧 Actions
Tip: Click or drag on the grid to paint. Switch tools above. Try moving the 🏁 goal to a corner blocked by 💀 zones!
DQN — click ▶ Play to run | click grid to edit
0
Step
0.00
Reward
0.00
Penalty
Status
Risky cells visited per algorithm
📊 Algorithm Results
DQN idle
Deep Q-Network with experience replay. No risk awareness — learns from reward signal alone. May pass through risky cells.
Steps: Rew: Path:
PPO idle
On-policy actor-critic. Uses short rollouts — struggles on large grids with sparse rewards. Often wanders without reaching goal.
Steps: Rew: Path:
SIRL idle
Squirrel-Inspired RL with RiskNet threat predictor. Penalises dangerous cells and routes around them. Safest path to goal.
Steps: Rew: Path:
📋 Step Log
Edit the grid, then press ▶ Play...
🧠 Algorithm Behaviour
DQN — Cost: fatal=∞, risky=1.5, safe=1. Takes the shortest path even through risky zones.
PPO — Random walk with weak goal bias. Fails on large grids.
SIRL — Cost: fatal=∞, risky=20, safe=1. Pays the extra steps to stay safe.
🔑 Why DQN ≈ SIRL on Simple Grids?
⚠ This is expected!
On simple grids, both reach the goal — the real difference is:
  • Penalty — DQN walks through risky cells; SIRL avoids them → 5× less penalty
  • Failures during training — DQN dies many times before learning; SIRL avoids early
  • Try preset: "Force Risky Path" to make the difference obvious!
📊 Penalty Comparison
🧠 Algorithm Behaviour