Adadelta
Adagrad
Adam
AdamW
AMSGrad
LAMB Optimizer
Lion Optimizer
Lookahead Optimizer
Nadam
RMSProp
-
-
-
N
AMSGrad Optimizer — annotated PyTorch implementation with the max-of-past-second-moments fix Art… 1 term
-
-
O
Incorporating Nesterov Momentum into Adam (Dozat, ICLR 2016 workshop) — the Nadam paper Pap… 1 term
-
-
-
A
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes (the LAMB paper, arXiv) Pap… 1 term
-
-
G
michaelrzhang/lookahead — reference Lookahead optimizer implementation by the paper's first author Code 1 term
-
-
J
Original Adagrad paper: Adaptive Subgradient Methods for Online Learning and Stochastic Optimization Pap… 1 term
-
-
-
-
-
-
-
-
No terms or links match your search.