Back to Home

Q-Learning Pro

Advanced reinforcement learning with comprehensive features

Quick Load Environment

Learning Parameters

Controls how much future rewards matter (0 = only immediate, 1 = future matters most)

How quickly the agent learns from new experiences

Probability of exploring random actions vs exploiting learned policy

Number of training episodes

Environment Configuration

Reward Matrix

Define rewards for each state-action pair. Positive values encourage actions, negative values discourage them.

📄 CSV Import Format:

Download Sample

Your CSV should have this structure:

State,Action0,Action1,Action2
0,0,0,0
1,0,10,0
2,0,0,50
3,0,0,100
  • • First row: Headers (State, Action0, Action1, ...)
  • • Other rows: State index followed by reward values for each action
  • • Number of actions determined by number of columns after State column
  • • Supports any number of states and actions (minimum 2 each)
State \ ActionA0A1
S0
S1
S2
S3