cd/entity/SGD· home entities SGD
grep -l @sgd /news/*.json | wc -l → 6

SGD

mentions 6 type Organization feed RSS

// recent coverage 6 mentions

19:34
2026-07-31
idlemachines.co.uk
machine-learning

Adam and AdamW: adaptive optimisation and weight decay

In a technical explainer, the author demonstrates that L2 regularization and weight decay, equivalent in plain SGD, diverge under the Adam optimizer because Adam's adaptive step rescales the L2 penalt…

14:46
2026-07-23
promptcube3.com
machine-learning

Deep Learning: A Complete Guide from Scratch

A guide explains that a neuron is a weighted sum followed by a non-linear activation, without which a deep network would be a linear regression. The training workflow consists of forward pass, loss ca…

04:55
2026-07-16
thegustafson.com
machine-learning

Optimizers: Momentum, Adam, and Learning Rate Schedules

A technical blog post explains that AdamW with warmup and cosine decay is the optimizer almost everyone uses today to train large language models, tracing the evolution from vanilla SGD through moment…

08:23
2026-07-01
machinebrief.com
large-language-models

Reining in AI Hallucinations: A New Approach to Dialogue Safety

A new study introduces the Guided-Retry strategy to reduce AI hallucinations in task-oriented dialogues without retraining models. Tested on models like DeepSeek-R1 and Llama-3, the method cut halluci…

// co-occurs with top 8 entities
// topics top 6 topics