Computing Atlas

How Computing Was Built
Sign In
Text size
100%
Theme
Algorithm

Adam Optimizer Algorithm

Optimization Algorithm

Adam, short for Adaptive Moment Estimation, is an optimization algorithm used to train machine learning models by adjusting each parameter's own learning rate based on the recent history of its gradients. It keeps a running exponential average of each parameter's gradient, its first moment, and of that gradient's square, its second moment, then corrects both averages for the bias introduced by starting them at zero and uses the corrected values to scale how much each parameter is updated on every training step. Diederik Kingma and Jimmy Ba published Adam in 2014, describing it as an update to the earlier RMSProp optimizer that adds the momentum idea from an older method, and later researchers found that Adam's initial convergence proof does not strictly hold for every case, which produced later variants such as AMSGrad and AdamW meant to correct that gap. Despite that theoretical shortfall, Adam's strong practical performance made it the default optimizer built into major deep learning frameworks such as TensorFlow and PyTorch.

Comments (0)
No comments yet. Be the first to share a thought.
Reader Challenges (0)
No disputes yet. Spotted an error or a better source? Open the first one.