Batch Size
Epoch
Mini-Batch
Training Loop
Training Step
Adadelta
Adagrad
Adam
AdamW
AMSGrad
LAMB Optimizer
Lion Optimizer
Lookahead Optimizer
Nadam
RMSProp
Cosine Annealing Schedule
Learning Rate Warmup
OneCycle Learning Rate
Reduce LR on Plateau
Gradient Accumulation
Gradient Clipping
Loss Scaling
Mixed Precision Training
-
A
Revisiting Small Batch Training for Deep Neural Networks — best performance at batch sizes 2–32 Pap… 2 terms
-
A
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour — the gradual learning-rate warmup scheme Pap… 1 term
-
-
-
N
AMSGrad Optimizer — annotated PyTorch implementation with the max-of-past-second-moments fix Art… 1 term
-
-
-
-
-
-
-
O
Incorporating Nesterov Momentum into Adam (Dozat, ICLR 2016 workshop) — the Nadam paper Pap… 1 term
-
-
-
-
A
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes (the LAMB paper, arXiv) Pap… 1 term
-
-
G
michaelrzhang/lookahead — reference Lookahead optimizer implementation by the paper's first author Code 1 term
-
-
A
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima (Keskar et al., 2016) Pap… 1 term
-
-
-
D
Optimizing Model Parameters — PyTorch official training-loop tutorial (train_loop / test_loop) Docs 1 term
-
J
Original Adagrad paper: Adaptive Subgradient Methods for Online Learning and Stochastic Optimization Pap… 1 term
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
A
Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates (the one-cycle policy) Pap… 1 term
-
-
-
-
-
-
No terms or links match your search.