Downloads · 30 days
0
Athmapriyan/RLT
RLT is a machine learning model from Athmapriyan. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
Just open index.html in any modern browser. No build step, no server needed.
Downloads · 30 days
0
Access
Public
Updated Apr 3, 2026
Repo size
—
Likes
0
Public
Click a slice to open those files.
.js14.6 KB · 35%
From the Hugging Face model README
bandit-project/
├── index.html ← Main page (hero, theory, simulator, insights)
├── style.css ← All styles (dark theme, responsive layout)
├── simulation.js ← Core bandit engine (BanditEnv, EpsilonGreedyAgent, runSimulation)
└── main.js ← UI logic (Plotly charts, sliders, hero canvas)
Just open index.html in any modern browser. No build step, no server needed.
# Option 1: Direct
open index.html
# Option 2: Local server (recommended)
python3 -m http.server 8080
# then visit http://localhost:8080
ε-Greedy with three initialization strategies:
| Strategy | Initial Q₀ | Exploration drive |
|---|---|---|
| Optimistic | +5 to +20 | High (natural) |
| Normal | 0 | Medium (ε only) |
| Pessimistic | −5 to −20 | Low |
Update rule: Q(a) ← Q(a) + [R − Q(a)] / N(a) (incremental sample mean)
| Parameter | Description | Range |
|---|---|---|
| K (arms) | Number of bandit arms | 2 – 10 |
| Iterations | Steps per simulation run | 100 – 3000 |
| ε (epsilon) | Exploration probability | 0.00 – 0.50 |
| Runs | Repetitions averaged together | 1 – 50 |
| σ (noise) | Reward standard deviation | 0.1 – 3.0 |
| Opt Q₀ | Optimistic initial value | +1 to +20 |
| Pess Q₀ | Pessimistic initial value | −1 to −20 |
sanjay