Artificial Neuron
Bias
Perceptron
Pre-Activation
Feedforward Network
Forward Pass
Fully Connected Layer
Hidden Layer
Input Layer
Multilayer Perceptron (MLP)
Output Layer
Network Depth
Network Width
Parameter Count
Hard Sigmoid
Sigmoid
Softmax
ELU
Leaky ReLU
Maxout
PReLU
ReLU
SELU
GELU
Mish
Swish
He Initialization
Identity Initialization
LeCun Initialization
Orthogonal Initialization
Random Initialization
Xavier (Glorot) Initialization
Backpropagation
Backpropagation Through Time (BPTT)
Exploding Gradients
Gradient Checking
Gradient Flow
Local Gradient
Upstream Gradient
Batch Size
Epoch
Mini-Batch
Training Loop
Training Step
Adadelta
Adagrad
Adam
AdamW
AMSGrad
LAMB Optimizer
Lion Optimizer
Lookahead Optimizer
Nadam
RMSProp
Cosine Annealing Schedule
Learning Rate Warmup
OneCycle Learning Rate
Reduce LR on Plateau
Gradient Accumulation
Gradient Clipping
Loss Scaling
Mixed Precision Training
DropConnect
Dropout
Label Smoothing
Spectral Normalization
Stochastic Depth
Batch Normalization
Group Normalization
Instance Normalization
Layer Normalization
CutMix
Data Augmentation
Channel
Convolution Operation
Convolutional Layer
Feature Map
Kernel (Filter)
Padding (Same, Valid)
Receptive Field
1x1 Convolution
Depthwise Convolution
Dilated (Atrous) Convolution
Separable Convolution
Transposed Convolution
Average Pooling
Global Average Pooling
Max Pooling
Bidirectional RNN
Cell State
Forget Gate
Gated Recurrent Unit (GRU)
Hidden State
Input Gate
Long Short-Term Memory (LSTM)
Output Gate
Recurrent Neural Network (RNN)
Reset Gate (GRU)
Stacked RNN
Teacher Forcing
Truncated BPTT
Attention Mechanism
Attention Score
Causal (Masked) Attention
Cross-Attention
Key Vector
Multi-Head Attention
Padding Mask
Query Vector
Scaled Dot-Product Attention
Self-Attention Mechanism
Value Vector
Decoder Block
Encoder Block
Position-wise Feed-Forward Network
Positional Embeddings
Pre-LN vs Post-LN
Transformer
Autoregressive
Sequence-to-Sequence
-
+8 more
-
-
C
Stanford CS231n — Neural Networks Part 1: Modeling one neuron, activation functions, architecture Art… 5 terms
-
-
D
Dive into Deep Learning 5.4 — Xavier initialization and variance stability across layers Art… 4 terms
-
-
-
-
-
C
CS231n — Backpropagation, Intuitions (computational graphs, chain rule, local and upstream gradients) Art… 3 terms
-
D
Dive into Deep Learning 5.3 — Forward Propagation, Backward Propagation, and Computational Graphs (intermediate pre-activation z) Art… 3 terms
-
-
-
-
-
-
P
Understanding the difficulty of training deep feedforward neural networks (Xavier initialization) Pap… 3 terms
-
-
D
7.2. Convolutions for Images — defines the feature map as a convolutional layer's output (Dive into Deep Learning) Art… 2 terms
-
-
-
A
A guide to convolution arithmetic for deep learning (Dumoulin & Visin) — how kernel, stride and padding set feature-map size Pap… 2 terms
-
D
Deep Learning (Goodfellow et al.) — Ch. 10 Sequence Modeling: Recurrent and Recursive Nets Art… 2 terms
-
D
Dive into Deep Learning 5.1 — Multilayer Perceptrons (feedforward architecture and hidden layers) Art… 2 terms
-
-
-
-
-
-
A
Revisiting Small Batch Training for Deep Neural Networks — best performance at batch sizes 2–32 Pap… 2 terms
-
-
-
D
10.1. Long Short-Term Memory (LSTM) — Dive into Deep Learning (forget gate and internal state) Art… 1 term
-
D
10.3. Deep Recurrent Neural Networks — Dive into Deep Learning (stacking recurrent layers) Art… 1 term
-
-
-
-
L
A survey on Image Data Augmentation for Deep Learning (Shorten & Khoshgoftaar, J. Big Data 2019) Pap… 1 term
-
A
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour — the gradual learning-rate warmup scheme Pap… 1 term
-
-
-
-
N
AMSGrad Optimizer — annotated PyTorch implementation with the max-of-past-second-moments fix Art… 1 term
-
-
-
-
C
Backpropagation Through Time: What It Does and How to Do It — Werbos, Proc. IEEE 78(10), 1990 (PDF) PDF 1 term
-
D
Bidirectional Recurrent Neural Networks — Schuster & Paliwal, IEEE Trans. Signal Processing 45(11), 1997 (PDF) PDF 1 term
-
-
-
-
-
-
-
-
T
Data augmentation — official TensorFlow Core tutorial (Keras preprocessing layers, tf.image) Docs 1 term
-
-
A
DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs Pap… 1 term
-
-
-
-
-
D
Dive into Deep Learning 8.6 — Residual Networks (ResNet) and ResNeXt: residual blocks and skip connections Art… 1 term
-
L
Efficient BackProp (LeCun, Bottou, Orr & Müller, 1998) — the 1/√fan-in initialization rule Art… 1 term
-
-
G
facebookresearch/mixup-cifar10 — official mixup implementation from the paper's authors Code 1 term
-
A
Fixup Initialization: Residual Learning Without Normalization (Zhang, Dauphin & Ma, ICLR 2019) Pap… 1 term
-
-
-
-
-
A
How to Construct Deep Recurrent Neural Networks — Pascanu, Gulcehre, Cho & Bengio (ICLR 2014) Pap… 1 term
-
-
O
Incorporating Nesterov Momentum into Adam (Dozat, ICLR 2016 workshop) — the Nadam paper Pap… 1 term
-
-
-
K
Keras Dense layer — API documentation (output = activation(dot(input, kernel) + bias)) Docs 1 term
-
-
-
-
-
L
Label Smoothing — Lei Mao's Log Book (derivation against cross-entropy with soft targets) Art… 1 term
-
-
A
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes (the LAMB paper, arXiv) Pap… 1 term
-
-
-
A
Maxout Networks (Goodfellow, Warde-Farley, Mirza, Courville, Bengio, ICML 2013) — arXiv full text Pap… 1 term
-
-
G
michaelrzhang/lookahead — reference Lookahead optimizer implementation by the paper's first author Code 1 term
-
-
-
-
-
A
Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau, Cho, Bengio) Pap… 1 term
-
D
Neural networks: Multi-class classification (softmax output layer) — Google Machine Learning Crash Course Docs 1 term
-
-
N
Nielsen, Neural Networks and Deep Learning — Ch. 4: A visual proof that neural nets can compute any function Art… 1 term
-
-
A
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima (Keskar et al., 2016) Pap… 1 term
-
-
-
-
D
Optimizing Model Parameters — PyTorch official training-loop tutorial (train_loop / test_loop) Docs 1 term
-
J
Original Adagrad paper: Adaptive Subgradient Methods for Online Learning and Stochastic Optimization Pap… 1 term
-
-
A
Original ELU paper: Fast and Accurate Deep Network Learning by Exponential Linear Units Pap… 1 term
-
-
-
-
-
A
Professor Forcing: A New Algorithm for Training Recurrent Networks — Lamb et al. (NeurIPS 2016) Pap… 1 term
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
C
Rectified Linear Units Improve Restricted Boltzmann Machines (Nair & Hinton, ICML 2010) PDF 1 term
-
A
Rectifier Nonlinearities Improve Neural Network Acoustic Models (Maas, Hannun & Ng, ICML 2013) PDF 1 term
-
A
ReZero is All You Need: Fast Convergence at Large Depth — zero-initialized residual gate = identity at init Pap… 1 term
-
-
-
A
Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks (Bengio et al., 2015) Pap… 1 term
-
-
-
-
-
E
Softmax function — Wikipedia (logits to probability distribution, numerical stability, attention) Ref… 1 term
-
-
C
Stanford CS231n — Setting up the data and the model: “Pitfall: all zero initialization” Art… 1 term
-
A
Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates (the one-cycle policy) Pap… 1 term
-
-
-
-
-
-
H
Text generation — autoregressive generation with LLMs (Hugging Face Transformers docs) Docs 1 term
-
T
tf.keras.activations.tanh — TensorFlow API documentation (formula, range [-1,1], zero-centred) Docs 1 term
-
-
-
-
-
-
-
-
-
-
-
-
D
torch.autograd.gradcheck.gradcheck — PyTorch documentation (analytical vs finite-difference gradients) Docs 1 term
-
-
-
P
torch.nn.SELU — PyTorch documentation (scale/alpha constants, self-normalizing networks note) Docs 1 term
-
-
-
-
-
-
A
Understanding the Difficulty of Training Transformers (Liu et al., 2020) — Post-LN instability and Admin Pap… 1 term
-
A
Understanding the Effective Receptive Field in Deep Convolutional Neural Networks (Luo et al., NIPS 2016) Pap… 1 term
-
-
A
Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks Pap… 1 term
-
-
-
No terms or links match your search.