Statistical Mechanics for Deep Learning & CS
A rigorous, modern curriculum bridging statistical physics, information theory, and generative AI / LLMs (2020–2026).
Statistical Mechanics for Deep Learning & Computer Science
Why does Statistical Mechanics sit at the core of Modern AI?
Statistical mechanics was created to answer: How do we predict the emergent behavior of microscopic particles without solving individual trajectories?
Deep learning asks the exact same question: How do we understand the emergent behavior of parameters and millions of tokens without tracking individual weights?
🗺️ The CS Stat-Mech Rosetta Stone
| Statistical Mechanics | Computer Science & Machine Learning | Primary Application / Literature |
|---|---|---|
| Microstate | Parameter vector / Token sequence | State spaces, Latent configurations |
| Energy | Loss function / Negative Logit | Energy-Based Models (EBMs), Softmax |
| Boltzmann / Gibbs Distribution | Softmax with Temperature / Gibbs measure | Autoregressive LLM Sampling, Policy |
| Partition Function | Normalization factor / Log-Sum-Exp | EBMs, Evidence, DPO cancellation trick |
| Entropy | Shannon Entropy / Exploration Regularizer | Policy diversity, Attention entropy collapse |
| Free Energy | Variational Free Energy / ELBO | Variational Autoencoders (VAEs), KL Optimization |
| Spin Glasses / Ising Model | Hopfield Networks / Transformer Attention | Associative Memory, In-Context Learning theory |
| Thermodynamic Limit () | Infinite-Width Limit / High-Dimensional Regime | NNGP, Neural Tangent Kernel (NTK), Mean Field |
| Order Parameter / Criticality | Generalization error / Sharp loss transitions | Grokking, Phase transitions in learning |
| Langevin Dynamics & Noise | Diffusion Models & SGLD | DDPM (2020), Score-SDE (2021), LLaDA (2025) |
| Liouville Continuity Equation | Continuous Flow Matching | Flow Matching (ICLR 2023), Normalizing Flows |
📚 Curriculum Structure
1. Foundations: Gibbs, Softmax & EBMs
Microstates, Jaynes MaxEnt, derivation of Boltzmann distribution, softmax with temperature, and Energy-Based Models.
2. Free Energy, Variational Inference & DPO
Helmholtz free energy, ELBO in VAEs, and the mathematical derivation of DPO with partition function cancellation.
3. Large Numbers & Infinite-Width Networks
Stirling approximation, Central Limit Theorem, Gaussian Processes (NNGP/NTK), and mini-batch SGD variance.
4. Spin Glasses, Hopfield & Attention
The Ising model, Modern Hopfield Networks, attention as energy minimization, and spin-glass models of In-Context Learning.
5. Phase Transitions & Grokking
Order parameters, mean field theory, interpolation thresholds, grokking as a 1st-order phase transition, and REM.
6. Langevin Dynamics & Diffusion Models
Brownian motion, Fluctuation-Dissipation, DDPM, Score-SDEs, and discrete language diffusion (SEDD, LLaDA 8B).
7. Liouville & Continuous Flow Matching
Phase space density, Liouville theorem, probability continuity equation, and Flow Matching for generative modeling.
8. The Rosetta Stone & 2020–2026 Research Map
Full conceptual dictionary, annotated 16-paper reading list, and open theoretical frontiers in AI.
2024 Final Examination (27th Batch)
Questions and complete step-by-step mathematical solutions for the CS-4125 Machine Learning 2024 final examination.
1. Foundations: Gibbs Distributions, Softmax & Energy-Based Models
Microstates vs macrostates, Jaynes' Maximum Entropy derivation of the Boltzmann/Gibbs measure, exact equivalence with temperature-scaled softmax, and training Energy-Based Models in modern generative AI.