AdaGrad vs RMSProp — Three Honest Scenarios

Each scenario is numerically verified to show a genuine algorithmic difference — not a tuning artifact

0.90
GDGradient Descent
Effective lr over time
AGAdaGrad
Effective lr over time
RMRMSProp
Effective lr over time
Effective lr: GD (orange) AdaGrad (blue) RMSProp (purple)
GD Steps
0
GD Loss
GD theta
AG Steps
0
AG Loss
AG theta
RM Steps
0
RM Loss
RM theta
Gradient Descent θt = θt−1 − η · gt

Fixed effective lr = η every step
AdaGrad Gt = Gt−1 + gt²
θt = θt−1 η Gt + ε · gt

Gt grows monotonically → η/(√Gt + ε) → 0
RMSProp Gt = ρ · Gt−1 + (1−ρ) · gt²
θt = θt−1 η Gt + ε · gt

EMA keeps Gt bounded → η/√Gt stays healthy
Loading...