Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
Abstract
Stochastic gradient descent dynamics are modeled as a percolation process where architectural symmetries cause subnetworks to merge in discrete blocks, producing variance spikes resembling phase transitions, with similar trapping behavior extending to Adam and AdamW under heavy-tailed noise.
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which architectural symmetries force subnetworks to merge in discrete simultaneous blocks rather than one at a time. These structural transitions register as variance spikes in a macroscopic order parameter, echoing physical phase transitions. We further show this trapping mechanism and its associated scaling cascade extend to Adam and AdamW under an explicit heavy-tailed noise model.
Community
SGD collapses deep neural networks toward sparse, low-rank representations generated by architectural symmetry. We show how this collapse progresses over time, and that the mechanism extends to Adam and AdamW.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Persistent Gaussian Perturbations Prevent Oversmoothing in Recurrent Graph Neural Networks (2026)
- Learning as a Geometric Phase Transition: Renormalization Group Flow and Anisotropic Symmetry Breaking in Deep Networks (2026)
- Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory (2026)
- Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries (2026)
- Teacher Knows It Best: Spontaneous Symmetry Breaking and Tipping Points in Networked Langevin Dynamics AI Sycophancy (2026)
- DiPhon: Diffusion on Graphons for Scalable Graph Generation (2026)
- Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.02373 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper