Speaker
Description
Estimating the latent capacity required to model high-dimensional data is important for representation learning, compression, model design, and downstream control. We formulate latent capacity estimation as a sequential measurement problem over an ordered family of nested representations induced by prefix-masked autoencoders. This turns latent capacity selection from a costly architectural search into a one-dimensional decision variable, allowing effective representational dimensionality to emerge during training within a single shared model.
Beyond reconstruction, the learned latent space provides a compact control space for inverse problems and adaptive optimization. We show that latent representations learned from simulated point spread function data and beamline datasets can be used to control the underlying system efficiently, enabling fast adaptation toward target outputs without expensive search in the original high-dimensional parameter space.
Empirically, adaptive allocation focuses training on the most informative capacity regimes, giving lower reconstruction error, reduced variance, and more stable dimension estimates than fixed schedules under matched budgets. Experiments on synthetic manifolds, structured image datasets, simulated PSF data, and beamline control benchmarks show that effective representational capacity can be identified reliably within a single training run while supporting fast latent-space adaptation for scientific instrument control.