Q-Learning Pro
Advanced reinforcement learning with comprehensive features
Quick Load Environment
Learning Parameters
Controls how much future rewards matter (0 = only immediate, 1 = future matters most)
How quickly the agent learns from new experiences
Probability of exploring random actions vs exploiting learned policy
Number of training episodes
Environment Configuration
Reward Matrix
Define rewards for each state-action pair. Positive values encourage actions, negative values discourage them.
📄 CSV Import Format:
Download SampleYour CSV should have this structure:
- • First row: Headers (State, Action0, Action1, ...)
- • Other rows: State index followed by reward values for each action
- • Number of actions determined by number of columns after State column
- • Supports any number of states and actions (minimum 2 each)
| State \ Action | A0 | A1 |
|---|---|---|
| S0 | ||
| S1 | ||
| S2 | ||
| S3 |
Related Topics & Algorithms
Explore these related algorithms and concepts to deepen your understanding and discover complementary techniques.
Neural Networks
Deep Q-networks and policy networks
Gradient Descent
Policy gradient optimization methods
Time Series
Sequential decision-making over time
Statistics
Probability and reward distributions
Hyperparameter Tuning
Tune exploration-exploitation tradeoff
GAN
Adversarial training concepts