WP1 – Asymptotics and Turnpike for Deep Neural Networks
Turnpike theory, originally an economic concept, describes how optimal trajectories remain close to a stable, efficient regime — the “turnpike” — for most of their evolution, departing from it primarily to accommodate initial and terminal conditions. In Machine Learning, this principle provides a powerful framework for understanding optimization dynamics, particularly in the training of Deep Neural Networks. By identifying and exploiting turnpike structures, we develop efficient training strategies, accelerating convergence and reducing computational costs while preserving model performance.
Task 1.1 – Neural networks with varying width
Neural networks do not need to have the same size at every layer. We explore how varying the width across layers can lead to simpler and more effective architectures, using switching dynamical systems and turnpike theory.
Task 1.2 – Sparse neural network architectures
Large neural networks often contain many connections that are not essential. Sparsity provides a way to keep only the relevant ones, producing smaller networks whose architecture adapts naturally to the learning problem.
Task 1.3 – Architectures that work across different data
A good neural network architecture should not depend excessively on one particular dataset. We look at optimal architectures that remain effective across different data, connecting neural network design with ideas from optimal sensor and actuator placement.
Task 1.4 – Training with multiple optimal solutions
A learning problem may have several different solutions that are equally good. Turnpike theory can explain which of these solutions the training dynamics approach and how this choice depends on the loss function and the network dynamics.
Task 1.5 – Turnpike theory for discrete neural networks
Turnpike theory provides a new perspective on why the parameters of very deep networks may display simple patterns during training. This perspective is extended from continuous-time models to ResNets and other discrete neural network architectures.
Publications
Wu, Y. & Ji Z. (2026). Mean-Covariance Turnpikes in Wasserstein Distributionally Robust Linear-Quadratic Control. arXiv preprint. arXiv: 2503.20342
Qian, M., Faria, J. R. D., Santos, A. J., Sokołowski, J., & Wyse, A. P. (2026). Topological Derivative Method for Design and Control of Timoshenko Beam Networks. Appl. Math. Optim., 93(1), 15.
Trélat, E. & Zuazua, E. (2025). Turnpike in Optimal Control and Beyond: A Survey. arXiv preprint. arXiv: 2503.20342