cd /news/machine-learning/continuous-delayed-memory-stochastic… · home topics machine-learning article
[ARTICLE · art-135542] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies

A new arXiv paper (2609.20906v1) introduces Continuous-Delayed-Memory Stochastic Gradient Descent, an optimizer that depends on the past state of the discrete iteration process, and reports that simulations on 2-dimensional landscapes showed wider exploration and more precise convergence than Vanilla SGD when hyperparameters were adjusted. The paper also proposes a continuous-time policy-gradient reinforcement learning structure that avoids solving the Hamilton-Jacobi-Bellman PDE, with optimality conditions that recover the Gibbs policy of prior work. The work reviews how stochastic differential equations with neural network parameterizations have been adapted to model quasar light curves from ground-based survey time series.

by read1 min views1 publishedSep 21, 2026

arXiv:2609.20906v1 Announce Type: new Abstract: Quasars are luminous objects in the universe that exhibit stochastic brightness variations encoding information about the supermassive black holes powering them, and modeling these variations from ground-based survey data time series, known as light curves, is a statistical challenge. This paper reviews how stochastic differential equations (SDEs) have been adapted with neural network parameterizations to overcome this challenge in history. We create the Continuous-Delayed-Memory Stochastic Gradient Descent which depend on the past state of the discrete iteration process. We performed the simulation on some 2-dimensional landscape and observed some wider-exploration and more precise convergent behavior compared to Vanilla SGD by adjusting hyperparameters. Besides, we proposed a reinforcement learning structure with continuous time policy gradients for exploratory policies without solving HJB PDE, and we show that its optimality conditions recover the Gibbs policy of previous works.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/continuous-delayed-m…] indexed:0 read:1min 2026-09-21 ·