aDarXivDesk
ExploreDocs

Statistics Theory

28,659 papers in this slice of arXiv.

All fieldsArtificial IntelligenceMachine LearningComputation and LanguageComputer Vision and Pattern RecognitionNeural and Evolutionary ComputingRoboticsInformation RetrievalHuman-Computer InteractionCryptography and SecurityData Structures and AlgorithmsSoftware EngineeringDistributed, Parallel, and Cluster ComputingProgramming LanguagesSystems and Control
2608.13520
2 days ago

The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity

Martin J. Wainwright

We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the unmasking growth complexity (UGC). Its local increments directly control Kullback--Leibler (KL) discretization error, yielding a unified analysis of Bernoulli-subset and fixed-cardinality unmasking schemes. In log-reveal-odds coordinates, this structure yields optimized single-block and multi-block schedules, and quantifies the gains from adapting computational effort to data geometry. Crucially, we show how UGC increments can be estimated from samples via KL increments along coupled reveal trajectories. This leads to certified-optimal samplers that achieve a prescribed KL error with high probability and iteration complexity within a constant factor of the corresponding oracle procedure. Collapsing the path yields the aggregate UGC mass, which connects to classical multivariate dependence measures and complexity measures from previous analyses of discrete diffusion. In the fine-partition limit, the squared integral of the square-root UGC density determines the sharp leading-order optimal Euler discretization error. Examples exhibit substantial dimension-dependent gains over coarse schedules, including Ω~(d)\widetildeΩ(\sqrt{d})Ω(d​)

PreviousNext
improvements achievable with a constant number of adaptively placed blocks.
Machine LearningArtificial IntelligenceInformation Theory
2608.13510
2 days ago

On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective

Nestor R. Barraza, Gabriel Pena

Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. However, their achievable performance is fundamentally constrained by structural properties of the underlying data-generating process, which are formalized in terms of informational bounds. In this work we examine intrinsic limits of data-driven decision systems from an information-theoretic and interaction-based perspective. We analyze minimal achievable error in classification through Fano-type bounds and precision limits in parametric estimation via the Cramér-Rao inequality, emphasizing that such limits depend on the underlying model rather than on algorithmic sophistication alone. We further discuss how implicit assumptions, such as independence, ergodicity, and distributional stability, affect the validity of inferential procedures. Building on interaction-based modeling principles, we review typical frameworks such as Markov Random Fields and potential based representations for encoding dependence mechanisms. We also describe decision systems, including LLM-integrated agent architectures, as feedback-driven stochastic processes where state-dependent dynamics may induce emergent macroscopic behavior. This perspective highlights the importance of having adequate models for the data as a prerequi- site for expanding predictive capability, and situates algorithmic learning within the informational limits imposed by the models.

Statistics TheoryMachine Learning
2608.13442
2 days ago

Theoretical Properties of Covariate-Adaptive Randomization with a Diverging Number of Covariates

Yuhang Tao, Li-Xin Zhang

Covariate-adaptive randomization procedures are widely used in clinical trials to improve covariate balance. In modern applications, experimenters often have access to many covariates, motivating the need for a theory of covariate-adaptive randomization procedures with a diverging number of covariates. In this paper, we study the theoretical properties of two unified families of covariate-adaptive randomization procedures under high-dimensional settings. For the two procedures, we establish the convergence rate of the imbalance measure corresponding to the specified covariates. In addition, for one of them, we study the asymptotic properties of the imbalance of unspecified covariates under and apply these results to derive the asymptotic properties of the difference-in-means estimator for the average treatment effect and construct asymptotic 95% confidence intervals. Furthermore, we provide extensive numerical and empirical studies to illustrate the practical relevance of our theoretical results.

MethodologyStatistics Theory
2608.13383
2 days ago

On Bridging Mixture Distributions

Pierre Del Moral, Ajay Jasra, Ke Zhao

In this article we consider bridging between two mixture probability measures. In particular, given access to a Markov kernel between two component distributions, we provide a general mechanism to generate samples from one mixture to the other. Associated to a given reference and extended state space, we prove entropic optimality of this approach. In order to use this idea one needs to know the underlying mixtures and the Markov kernel, which is seldom available, and so we consider the case of Gaussian mixtures and Schrödinger Bridges. We prove a general 2−2-2−Wasserstein continuity bound between the exact bridge and one that is approximated, based on ε−ε-ε−covariance inflation, and these rely on a novel continuity analysis of perturbed Riccati maps. We apply our results in the context of bridging mixtures of Gaussians, single Gaussians and empirical estimators of the Gaussian parameters and the Monge map. For mixtures of Gaussians, when the parameters are estimated using the Expectation-Maxmization algorithm, the upper-bound on the 2−2-2−Wasserstein distance between the true and approximated bridges is, under assumptions and with probability at least 1−10N−11-10N^{-1}1−10N−1, O([(dlog⁡NN)1/2{1+(dlog⁡NN)1/2(ε−2+1)}+ε2])\mathcal{O}\big(\big[\big(\tfrac{d\log N}{N}\big)^{1/2}\left\{1+\big(\tfrac{d\log N}{N}\right)^{1/2}(ε^{-2}+1)\big\}+ε^2\big]\big)O([(NdlogN​)1/2{1+(NdlogN​)1/2(ε−2+1)}+ε2]) and for the other two cases, in expectation, O(d{1+ε−21+N+ε2})\mathcal{O}\left(d\left\{\tfrac{1+ε^{-2}}{1+N}+ε^2\right\}\right)O(d{1+N1+ε−2​+ε2}), where d,N∈Nd,N\in\mathbb{N}d,N∈N is the dimension of the Gaussian and the number of empirical samples respectively. We also investigate our bounds numerically.

Statistics TheoryProbability
2608.13374
2 days ago

Nearly sharp comparison results for sliced and max-sliced Wasserstein distances

Jonathan Niles-Weed, Jacob Shkrob

We prove new comparison results between the Wasserstein distance and its sliced and max-sliced counterparts. First, we show that the Hölder exponent 2d+2\frac{2}{d+2}d+22​ obtained by Bobkov and Götze for the max-sliced 1-Wasserstein distance on the unit ball is optimal for every d≥2d \geq 2d≥2, settling a question raised in their work. Second, we show that sharper comparisons are possible under stronger structural assumptions: if ννν is a discrete measure and the optimal coupling between μμμ and ννν transports each point to a nearest atom of ννν, then Wp(μ,ν)≤Cd K SWp,1(μ,ν)W_p(μ, ν) \leq C \sqrt{d}\, K \, \mathrm{SW}_{p,1}(μ, ν)Wp​(μ,ν)≤Cd​KSWp,1​(μ,ν) for a universal constant CCC, where the complexity parameter KKK is always at most the number of atoms NNN and can be substantially smaller. This complements a similar bound due to Park and Slepčev. An analogous bound holds for the sliced Wasserstein distance based on kkk-dimensional projections. Finally, using a construction from geometric discrepancy theory due to Chen and Travaglini, we prove that the linear dependence on KKK in this bound cannot be improved, up to polylogarithmic factors.

ProbabilityFunctional AnalysisStatistics Theory
2608.13363
2 days ago

Weighted cumulative past inaccuracy and Kullback-Leibler divergence based on extropy: properties, estimation, and applications

Bighneswar Sahoo, Suchandan Kayal

This study develops a weighted framework for measuring the discrepancy between two nonnegative lifetime distributions through cumulative past extropy. We propose two measures, referred to as the weighted cumulative past extropy inaccuracy (WCPEI) and the weighted cumulative past extropy Kullback-Leibler divergence (WCPED). The generalized weight function is considered in this study. We investigate a number of theoretical properties of these measures. Empirical distribution function-based nonparametric estimator is subsequently constructed for the weighted cumulative past extropy inaccuracy ration (WCPEIR). Its finite-sample behavior is studied through Monte Carlo simulation experiments for different sample sizes. To illustrate the practical relevance of the WCPED, two applications are considered. First, an extropy-based goodness-of-fit procedure for testing uniformity is developed using the proposed divergence measure. Its power is then compared with that of several established uniformity tests under a variety of alternatives. Second, an image analysis application is presented in which the proposed measure is employed to assess changes in the distributions of pixel intensities when the image resolution is altered. The framework is further extended to a dynamic setting by conditioning on the lifetime information available up to a specified time point. This leads to the dynamic weighted cumulative past extropy inaccuracy (DWCPEI) and dynamic weighted cumulative past extropy divergence (DWCPED). Their theoretical properties are derived, and the corresponding nonparametric estimation procedures are proposed. The finite-sample performance of these estimators is evaluated through simulation studies using R software.

MethodologyStatistics Theory
2608.13229
2 days ago

Foundations of Independent Component Analysis

Patrick Forré

We present the mathematical foundations of linear independent component analysis (ICA) models based on standard literature in a self-contained note. It is aimed at readers with a background in measure-theoretic probability theory. We first develop the theory of the characteristic functions of probability measures on Rd\mathbb{R}^dRd, including their analyticity and the way in which they determine and characterise the distributions. We then focus on several identifiability results of ICA models with successively strengthened assumptions on the sources: from merely non-constant, to non-Gaussian, to Gaussian-free independent sources. Under the strictest assumptions, we show that the independent sources are identifiable up to translation, permutation, scales and signs, and this even in the presence of additive Gaussian noise. Furthermore, we present the online equivariant gradient descent ICA algorithm for recovering the independent sources from data, in the standard complete noiseless non-Gaussian ICA setting.

Statistics TheoryMachine LearningProbability
2608.13201
2 days ago

Sinkhorn Linearization and the Spectral Proxy: Unifying the Statistical and Algorithmic Theory of Feature-Parameterized Inverse Optimal Transport via a Single Spectral Sandwich

Han Dong, Jiaming Li, Yongqiang Gong +2

We develop the statistical and algorithmic theory of inverse optimal transport (IOT) under the feature-parameterized cost C_theta(i,j) = -theta^T phi(i,j). The core technical contribution is the Sinkhorn linearization -- the implicit-function sensitivity of the entropic OT plan to the cost -- together with its spectral proxy, a formula that is spectrally exact yet geometrically transparent. The restricted Hessian on the tangent space satisfies the spectral sandwich (pi_min/epsilon) I <= H_T^{-1} <= (pi_max/epsilon) I, yielding the single core bound sigma_min >= (pi_min/(a_max epsilon)) sqrt(lambda_min(Sigma)) that drives the entire theory. On this core we establish four theorems and one observation. T1 (identifiability): theta is globally injective on the quotient of the gauge kernel, with dimension bound F <= (K-1)^2. T2 (sparsistency): the l1-penalized estimator recovers the true support under irrepresentability and score concentration, with exponential failure probability. T3 (well-posedness): the feature-moment map M(theta) = Phi^T x_theta is strongly monotone, and the inverse is Lipschitz with constant L <= epsilon ||Phi^T S_a||_op / (pi_min lambda_min(Sigma)). T4 (convergence): local strong convexity with mu >= pi_min^2 lambda_min(Sigma) / epsilon^2 guarantees monotone gradient descent convergence. O5 (misspecification): the estimator converges to the OT-model projection of the truth; the Holder continuity of the projection map is assessed numerically, yielding setting-dependent empirical exponents alpha_eff in (0,1).

Machine LearningMachine LearningOptimization and Control
2608.13154
2 days ago

Extreme principal minors of Wishart and deformed GOE matrices

Zhanrui Dong, Tiefeng Jiang, Tuan Pham +1

We study the laws of large numbers for the largest eigenvalues among all principal minors of Wishart matrices and deformed GOE matrices. We propose a new method based on identifying the deterministic sets to which the random sets formed by suitably normalized principal minors converge in Hausdorff distance, thereby reducing the original extreme-value problems to finite-dimensional convex optimization problems. We demonstrate the effectiveness of this method in regimes not covered by the existing second-moment arguments in cai2021asymptotic,hu2023extreme. For deformed GOE matrices with fixed minor size kkk, we determine the limit for every diagonal variance a>0a>0a>0 and identify a phase transition at a=2a=2a=2. Above the transition, the limiting constant satisfies an explicit recursion with no close-form expression, and the optimizers exhibit a nested hierarchical structure, thereby resolving the case left open in cai2021asymptotic. For Wishart matrices with general sub-Gaussian entries and fixed kkk, we characterize the limit through an entropy-constrained deterministic convex set. When the entries are standard Gaussian, we solve the resulting optimization problem explicitly and obtain the exact value of the limiting constant.

ProbabilityStatistics Theory
2608.13152
2 days ago

Estimation of distribution functions, their jumps and interval probabilities under measurement error

Kairat Mynbaev, Carlos Martins-Filho, Chad Brown

We consider the classical additive measurement-error model X=Y+ZX=Y+ZX=Y+Z, where the latent random variable YYY has unknown distribution FYF_YFY​ and the error ZZZ has a known distribution. We develop direct estimators for three functionals of FYF_YFY​: (i) FY(x)F_Y(x)FY​(x) at continuity points; (ii) interval probabilities FY(y)−FY(x)F_Y(y)-F_Y(x)FY​(y)−FY​(x) when x<yx<yx<y are continuity points; and (iii) the size of a jump at a prespecified discontinuity. We derive non-asymptotic bias and variance bounds, and establish asymptotic unbiasedness and consistency. Unlike previous work, we do not require FYF_YFY​ to admit a density, have a mixture representation, or satisfy global Sobolev smoothness assumptions. The framework accommodates arbitrary latent distributions, including those with both discrete and continuous components, and distributions with multiple jumps. These results rely on a link between Fourier inversion theorems and the algebraic structure of a class of estimators proposed in Mynbaev, Martins-Filho and Henderson (2022). A simulation study evaluates feasible tuning procedures and, where available, compares the finite-sample performance of the proposed estimators with existing methods.

EconometricsStatistics TheoryMethodology
2608.12949
2 days ago

Bayesian Inference Procedures for A/B Testing: An Overview

Mårten Schultzberg, Mattias Frånberg

Bayesian inference for A/B testing is a family of prior and stopping-rule configurations with fundamentally different statistical properties, but it is often discussed as a single method, and no systematic overview exists. This paper organizes common configurations into a three-tier hierarchy: 1) posterior coherence with no error control, 2) false positive rates bounded under continuous monitoring via Bayes factor stopping, and 3) false discovery rate control and calibrated shrinkage via empirical Bayes. Many commercial platforms operate at the lowest tier by default. We show that Bayes factor stopping is near-optimal for a broad class of cost functions, including most proposed in the A/B testing literature; because the same rule also controls the false positive rate, the choice between a decision-theoretic and a frequentist formulation is largely one of parameterization. Furthermore, the empirical Bayes prior is the only path to the third tier, but winner-selected corpora, pooled programs, and heterogeneous metrics can each prevent calibration regardless of corpus size. Simulations against group-sequential and always-valid frequentist baselines show that flat-prior posterior stopping exactly reproduces naive peeking, that a well-calibrated empirical Bayes prior achieves the lowest estimation error, and that expected-loss stopping minimizes regret only when shipping a null-effect variant is nearly free. Error rates, estimation accuracy, and regret are all different risks, and the appropriate method follows from the risks an experimentation program needs to control, not the other way around.

MethodologyStatistics Theory
2608.12838
2 days ago

Weighted Nuclear Elastic Net Estimation of (Near-) Low-Rank Drift Matrices in Ornstein-Uhlenbeck Processes

Dmytro Marushkevych, Francisco Pina, Mark Podolskij

We study estimation of the drift matrix in a continuously observed high-dimensional Ornstein-Uhlenbeck process when the drift is exactly or approximately low rank. In this setting, exact low rank induces non-stable directions and hence a non-ergodic regime, resulting in a poorly conditioned empirical covariance matrix. To address this difficulty, we introduce a Weighted Nuclear Elastic Net Estimator that combines ridge regularization with a nuclear-norm penalty expressed in the empirical likelihood geometry. Under a general diagonalizable spectral framework, we establish oracle inequalities relative to arbitrary low-rank comparison matrices. For near low-rank drifts, the approximation error is naturally measured through the singular-value decay of the drift after weighting by the regularized empirical covariance. The stochastic term is controlled by self-normalized martingale arguments under appropriate choice of the tuning parameter. For a symmetric positive-semidefinite exact low-rank model, we verify the empirical-curvature condition required to translate the weighted bound into a Frobenius-norm bound. With an appropriate choice of tuning parameters, the resulting estimator satisfies, up to a logarithmic factor, the standard rank-rrr matrix-estimation scaling rd/Tr d/Trd/T: specifically, its squared Frobenius error is of order rdlog⁡(T)/Tr d\log(T)/Trdlog(T)/T with high probability, under an explicit dimension-horizon condition.

Statistics Theory
2608.12740
2 days ago

Stable convergence of partial sum processes towards discontinuous limits

Johannes Brutsche

We develop a stable convergence theorem for partial sum processes on sample-size dependent stochastic bases. The result allows multidimensional semimartingale limits that have conditionally independent increments and both a continuous and discontinuous martingale part. Motivated by infill asymptotics, it complements classical Gaussian stable limit theorems and supports applications to likelihood based statistical inference.

ProbabilityStatistics Theory
2608.12738
2 days ago

Hyper-V uniform ergodicity of Markov chains

Austin Brown, Kshitij Khare

We develop a new uniform drift condition and local minorization that implies a stronger weighted form of uniform ergodicity for Markov chains we call hyper-V uniform ergodicity. The convergence guarantees geometric decay of the bias towards the invariant measure independently of the initialization for all functions controlled by a dominating function V. A key advantage of the approach is that it bypasses the need to establish a global minorization condition, which is often substantially more difficult to verify in practice, while yielding stronger convergence guarantees than global minorization. Optimal convergence bounds in a minimax sense of the framework are established. The utility of the framework is demonstrated through applications to the P'olya-Gamma and Kolmogorov-Gamma Gibbs samplers. We also show qualitative hyper-V uniform ergodicity convergence for two-variable Gibbs samplers can be inferred by the form of the invariant measure, bypassing convergence analysis entirely.

Statistics TheoryComputation
2608.12701
3 days ago

Sharp proper estimation of fixed-component Gaussian location mixtures in polynomial time

Hengzhi He, Guang Cheng

We consider a mixture of at most kkk unit-covariance Gaussians in Rd\mathbb{R}^dRd whose means belong to a fixed-radius ball, with no separation or minimum-weight condition. Doss, Wu, Yang and Zhou (2023) proved that the minimax Hellinger risk is of order d/n∧1\sqrt{d/n}\wedge 1d/n​∧1 and constructed a proper polynomial-time estimator with the slower general bound (d/n)1/4(d/n)^{1/4}(d/n)1/4; obtaining the sharp rate in polynomial time for fixed k≥3k\geq 3k≥3 was left open. We resolve this question. The key device is a moment-fiber range finder. A second-moment subspace controls the energy missed by projection. We then estimate finitely many one-free-index Hermite contractions. These vector-valued contractions recover every tensor component containing exactly one missed direction at the sharp d/n\sqrt{d/n}d/n​ scale. Every remaining term contains at least two missed factors and is therefore controlled by the residual second-moment energy. The resulting subspace has dimension depending only on kkk. Exhaustive moment fitting in this constant-dimensional space produces a proper mixture and, together with the dimension-free moment characterization of Gaussian mixtures, achieves the optimal Hellinger rate in polynomial arithmetic time for every fixed kkk.

Statistics Theory
2608.12660
3 days ago

Universality of e-detectors for ARL control

Aaditya Ramdas

An e-detector for a pre-change class P\mathcal PP is a nonnegative process MMM such that EP[Mτ]≤EP[τ]\mathbb E_P[M_τ] \leq \mathbb E_P[τ]EP​[Mτ​]≤EP​[τ] for all stopping times τττ and all P∈PP \in \mathcal PP∈P. Thresholding e-detectors controls the average run length (ARL): declaring a change at the first time TbT_bTb​ when MMM crosses bbb ensures that inf⁡P∈PEP[T]≥b\inf_{P \in \mathcal P}\mathbb E_P[T] \geq binfP∈P​EP​[T]≥b. But e-detectors do substantially more than control the ARL; they also satisfy a optional-horizon inequality: P(Tb≤σ)≤EP[σ]/bP(T_b\leqσ)\leq \mathbb E_P[σ]/bP(Tb​≤σ)≤EP​[σ]/b for every data-dependent stopping time (monitoring horizon) σσσ and P∈PP\in \mathcal PP∈P. In particular, every e-detector-based procedure obeys P(T≤t)≤t/bP(T\leq t)\leq t/bP(T≤t)≤t/b at each fixed ttt, thus avoiding early false alarms. Remarkably, the converse also holds: every stopping time TTT that satisfies the optional-horizon inequality must in fact arise from thresholding an e-detector. We also derive a universal representation of stopping times that satisfy (only) ARL control. These are represented by weak e-detectors, that only require EP[Mτ]≤EP[τ]\mathbb E_P[M_τ] \leq \mathbb E_P[τ]EP​[Mτ​]≤EP​[τ] to hold at all threshold stopping times TbT_bTb​. Appendices present universal representations for other (less common) change detection metrics.

Statistics TheoryInformation TheorySignal Processing
2608.12637
3 days ago

Density Estimation on Compact Manifolds under Intrinsic Spectral Block Variation

Olga Klopp, Fedor Noskov

We introduce an intrinsic spectral sparsity model for nonparametric density estimation on compact connected Riemannian manifolds. Instead of penalizing coefficients in an arbitrarily chosen Laplace--Beltrami eigenbasis, we group each complete eigenspace and measure the Hilbert norm of its spectral component. The resulting block-variation space is basis independent and isometry invariant. We establish its structural, atomic, and nonlinear approximation properties and clarify its relation to Sobolev, Besov, and coefficientwise spectral ℓ1\ell^1ℓ1 classes. We then construct a coordinate-free block-shrinkage estimator and prove a nonasymptotic signal-dependent L2L^2L2-oracle inequality that adapts to the unknown set of detectable eigenspaces. Under polynomial spectral growth, the risk theory separates the number of spectral blocks from their multiplicities and exhibits two regimes: one driven by a single high-dimensional eigenspace and the other by cumulative spectral complexity. Under matching spectral-growth and nondegeneracy assumptions, corresponding minimax lower bounds show that this multiplicity dependence is intrinsic, with sharp consequences for spheres and the rotation group SO(3)SO(3)SO(3). Finally, we develop a positive, normalized, block-penalized exponential spectral sieve for log-densities and derive likelihood oracle inequalities together with expected Kullback--Leibler, Hellinger, and L2L^2L2 risk bounds. The resulting framework provides a geometry-respecting theory of sparse density estimation that remains invariant under changes of eigenbasis.

Statistics TheoryFunctional Analysis
2608.12309
3 days ago

A contiguity approach to replica symmetric marginals

Ernesto Mordecki, Anas A. Rahman, Manuel Sáenz

We develop a probabilistic cavity-contiguity framework for proving replica-symmetric convergence of local marginals in mean-field Gibbs systems. The approach is based on cavity decompositions of the Hamiltonian together with direct comparison of probability measures through Radon-Nikodym derivatives and Hellinger-type estimates. At a conceptual level, the method separates the concentration of the relevant order parameters from the identification of the asymptotic cavity model and the comparison of the associated Gibbs measures. In contrast with interpolation-based approaches, the argument relies only weakly on the Gaussianity of the disorder and naturally accommodates concentration tools such as Poincaré and log-Sobolev inequalities. Rather than pursuing maximal generality, with the aim of making the method transparent, we implement the framework in a canonical example: the high-temperature Sherrington-Kirkpatrick model. In this setting, we prove that the marginal law of a fixed spin converges in total variation toward the effective one-dimensional cavity measure predicted by the replica method. Beyond the specific result for the Sherrington-Kirkpatrick model, the paper illustrates a broader cavity-contiguity methodology which is expected to extend naturally to other mean-field Gibbs systems, particularly Bayesian inference models with or without mismatch.

Mathematical PhysicsProbabilityStatistics Theory
2608.12242
3 days ago

Sharp Berry-Esseen Bounds for the Log Determinant of a Gaussian Sample Correlation Matrix

Hongru Zhao

Let R^\widehat RR be the Pearson sample correlation matrix formed from nnn independent Gaussian observations in ppp dimensions, and write m=n−1≥pm=n-1\ge pm=n−1≥p. Under the null correlation R=IpR=I_pR=Ip​, the classical independent beta product, exact cumulants, and full Fourier inversion yield, along every sequence p→∞p\to\inftyp→∞ with m≥pm\ge pm≥p, a uniform first Edgeworth expansion for log⁡det⁡R^\log\det\widehat RlogdetR, centered by its exact mean and scaled by its exact standard deviation. The expansion identifies the exact finite dimensional skewness correction and gives the sharp Kolmogorov equivalent Am,p/{62πVm,p3/2}A_{m,p}/\{6\sqrt{2π}V_{m,p}^{3/2}\}Am,p​/{62π​Vm,p3/2​}, where Vm,pV_{m,p}Vm,p​ is the exact variance and Am,pA_{m,p}Am,p​ is the absolute third cumulant. This equivalent unifies the square, fixed gap, growing gap, proportional, and dilute regimes; in the square regime the error has order (log⁡p)−3/2(\log p)^{-3/2}(logp)−3/2 with an exact constant. For every positive definite population correlation matrix RRR, we prove a uniform finite sample Berry-Esseen bound that explicitly tracks population dependence. All theoretical results have exact or proved equivalent Lean 4 formulations whose declarations and dependencies are kernel checked.

ProbabilityStatistics Theory
2608.12183
3 days ago

Spectral phase transitions in Gaussian multi-index models

Florent Krzakala, Pierre Mergny, Vanessa Piccolo

Recovering a low-dimensional latent subspace from nonlinear observations of Gaussian covariates in high dimensions is a fundamental problem in feature learning. Here, we consider Gaussian multi-index models in which the covariates xi∼i.i.d.N(0,Id)\boldsymbol{x}_i \stackrel{\mathrm{i.i.d.}}{\sim} \mathcal{N}(0,\boldsymbol{I}_d)xi​∼i.i.d.N(0,Id​) and the responses yi\boldsymbol{y}_iyi​ depend on xi\boldsymbol{x}_ixi​ only through its projection onto an unknown rrr-dimensional subspace. Earlier work based on approximate message passing (AMP) identified a sharp threshold for weak recovery [Troiani et al., 2025], raising the question of whether it can be attained, without side information, by a spectral method. We answer this affirmatively and develop a general random matrix theory for matrix-valued spectral estimators of the form Dn=1n∑i=1nT(yi)⊗xixi⊤,\boldsymbol{D}_n=\frac{1}{n}\sum_{i=1}^n\boldsymbol{T}(\boldsymbol{y}_i)\otimes\boldsymbol{x}_i\boldsymbol{x}_i^\top,Dn​=n1​i=1∑n​T(yi​)⊗xi​xi⊤​, where T\boldsymbol{T}T is an arbitrary bounded symmetric matrix-valued preprocessing map of fixed dimension. As n,d→∞n,d \to \inftyn,d→∞ with n/d→αn/d\toαn/d→α, we prove that the empirical spectral measure of Dn\boldsymbol{D}_nDn​ converges almost surely to a deterministic compactly supported distribution characterized by a matrix-valued self-consistent equation. We then establish a spectral phase transition for the largest eigenvalue: below threshold it sticks to the bulk edge, while above threshold an outlier emerges. We characterize the outlier location through a finite-dimensional deterministic equation and show that the associated spectral estimator achieves weak recovery of the latent subspace. Finally, we prove that the AMP-derived preprocessing of [Defilippis et al., 2025] is optimal among all bounded matrix-valued preprocessing maps of any fixed dimension. Its transition coincides with the AMP weak-recovery threshold, proving the general spectral conjecture of [Defilippis et al., 2025].

Statistics TheoryProbability