28,659 papers in this slice of arXiv.
Martin J. Wainwright
We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the unmasking growth complexity (UGC). Its local increments directly control Kullback--Leibler (KL) discretization error, yielding a unified analysis of Bernoulli-subset and fixed-cardinality unmasking schemes. In log-reveal-odds coordinates, this structure yields optimized single-block and multi-block schedules, and quantifies the gains from adapting computational effort to data geometry. Crucially, we show how UGC increments can be estimated from samples via KL increments along coupled reveal trajectories. This leads to certified-optimal samplers that achieve a prescribed KL error with high probability and iteration complexity within a constant factor of the corresponding oracle procedure. Collapsing the path yields the aggregate UGC mass, which connects to classical multivariate dependence measures and complexity measures from previous analyses of discrete diffusion. In the fine-partition limit, the squared integral of the square-root UGC density determines the sharp leading-order optimal Euler discretization error. Examples exhibit substantial dimension-dependent gains over coarse schedules, including Ω(d)
Nestor R. Barraza, Gabriel Pena
Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. However, their achievable performance is fundamentally constrained by structural properties of the underlying data-generating process, which are formalized in terms of informational bounds. In this work we examine intrinsic limits of data-driven decision systems from an information-theoretic and interaction-based perspective. We analyze minimal achievable error in classification through Fano-type bounds and precision limits in parametric estimation via the Cramér-Rao inequality, emphasizing that such limits depend on the underlying model rather than on algorithmic sophistication alone. We further discuss how implicit assumptions, such as independence, ergodicity, and distributional stability, affect the validity of inferential procedures. Building on interaction-based modeling principles, we review typical frameworks such as Markov Random Fields and potential based representations for encoding dependence mechanisms. We also describe decision systems, including LLM-integrated agent architectures, as feedback-driven stochastic processes where state-dependent dynamics may induce emergent macroscopic behavior. This perspective highlights the importance of having adequate models for the data as a prerequi- site for expanding predictive capability, and situates algorithmic learning within the informational limits imposed by the models.
Yuhang Tao, Li-Xin Zhang
Covariate-adaptive randomization procedures are widely used in clinical trials to improve covariate balance. In modern applications, experimenters often have access to many covariates, motivating the need for a theory of covariate-adaptive randomization procedures with a diverging number of covariates. In this paper, we study the theoretical properties of two unified families of covariate-adaptive randomization procedures under high-dimensional settings. For the two procedures, we establish the convergence rate of the imbalance measure corresponding to the specified covariates. In addition, for one of them, we study the asymptotic properties of the imbalance of unspecified covariates under and apply these results to derive the asymptotic properties of the difference-in-means estimator for the average treatment effect and construct asymptotic 95% confidence intervals. Furthermore, we provide extensive numerical and empirical studies to illustrate the practical relevance of our theoretical results.
Pierre Del Moral, Ajay Jasra, Ke Zhao
In this article we consider bridging between two mixture probability measures. In particular, given access to a Markov kernel between two component distributions, we provide a general mechanism to generate samples from one mixture to the other. Associated to a given reference and extended state space, we prove entropic optimality of this approach. In order to use this idea one needs to know the underlying mixtures and the Markov kernel, which is seldom available, and so we consider the case of Gaussian mixtures and Schrödinger Bridges. We prove a general 2−Wasserstein continuity bound between the exact bridge and one that is approximated, based on ε−covariance inflation, and these rely on a novel continuity analysis of perturbed Riccati maps. We apply our results in the context of bridging mixtures of Gaussians, single Gaussians and empirical estimators of the Gaussian parameters and the Monge map. For mixtures of Gaussians, when the parameters are estimated using the Expectation-Maxmization algorithm, the upper-bound on the 2−Wasserstein distance between the true and approximated bridges is, under assumptions and with probability at least 1−10N−1, O([(NdlogN)1/2{1+(NdlogN)1/2(ε−2+1)}+ε2]) and for the other two cases, in expectation, O(d{1+N1+ε−2+ε2}), where d,N∈N is the dimension of the Gaussian and the number of empirical samples respectively. We also investigate our bounds numerically.
Jonathan Niles-Weed, Jacob Shkrob
We prove new comparison results between the Wasserstein distance and its sliced and max-sliced counterparts. First, we show that the Hölder exponent d+22 obtained by Bobkov and Götze for the max-sliced 1-Wasserstein distance on the unit ball is optimal for every d≥2, settling a question raised in their work. Second, we show that sharper comparisons are possible under stronger structural assumptions: if ν is a discrete measure and the optimal coupling between μ and ν transports each point to a nearest atom of ν, then Wp(μ,ν)≤CdKSWp,1(μ,ν) for a universal constant C, where the complexity parameter K is always at most the number of atoms N and can be substantially smaller. This complements a similar bound due to Park and Slepčev. An analogous bound holds for the sliced Wasserstein distance based on k-dimensional projections. Finally, using a construction from geometric discrepancy theory due to Chen and Travaglini, we prove that the linear dependence on K in this bound cannot be improved, up to polylogarithmic factors.
Bighneswar Sahoo, Suchandan Kayal
This study develops a weighted framework for measuring the discrepancy between two nonnegative lifetime distributions through cumulative past extropy. We propose two measures, referred to as the weighted cumulative past extropy inaccuracy (WCPEI) and the weighted cumulative past extropy Kullback-Leibler divergence (WCPED). The generalized weight function is considered in this study. We investigate a number of theoretical properties of these measures. Empirical distribution function-based nonparametric estimator is subsequently constructed for the weighted cumulative past extropy inaccuracy ration (WCPEIR). Its finite-sample behavior is studied through Monte Carlo simulation experiments for different sample sizes. To illustrate the practical relevance of the WCPED, two applications are considered. First, an extropy-based goodness-of-fit procedure for testing uniformity is developed using the proposed divergence measure. Its power is then compared with that of several established uniformity tests under a variety of alternatives. Second, an image analysis application is presented in which the proposed measure is employed to assess changes in the distributions of pixel intensities when the image resolution is altered. The framework is further extended to a dynamic setting by conditioning on the lifetime information available up to a specified time point. This leads to the dynamic weighted cumulative past extropy inaccuracy (DWCPEI) and dynamic weighted cumulative past extropy divergence (DWCPED). Their theoretical properties are derived, and the corresponding nonparametric estimation procedures are proposed. The finite-sample performance of these estimators is evaluated through simulation studies using R software.
Patrick Forré
We present the mathematical foundations of linear independent component analysis (ICA) models based on standard literature in a self-contained note. It is aimed at readers with a background in measure-theoretic probability theory. We first develop the theory of the characteristic functions of probability measures on Rd, including their analyticity and the way in which they determine and characterise the distributions. We then focus on several identifiability results of ICA models with successively strengthened assumptions on the sources: from merely non-constant, to non-Gaussian, to Gaussian-free independent sources. Under the strictest assumptions, we show that the independent sources are identifiable up to translation, permutation, scales and signs, and this even in the presence of additive Gaussian noise. Furthermore, we present the online equivariant gradient descent ICA algorithm for recovering the independent sources from data, in the standard complete noiseless non-Gaussian ICA setting.
Han Dong, Jiaming Li, Yongqiang Gong +2
We develop the statistical and algorithmic theory of inverse optimal transport (IOT) under the feature-parameterized cost C_theta(i,j) = -theta^T phi(i,j). The core technical contribution is the Sinkhorn linearization -- the implicit-function sensitivity of the entropic OT plan to the cost -- together with its spectral proxy, a formula that is spectrally exact yet geometrically transparent. The restricted Hessian on the tangent space satisfies the spectral sandwich (pi_min/epsilon) I <= H_T^{-1} <= (pi_max/epsilon) I, yielding the single core bound sigma_min >= (pi_min/(a_max epsilon)) sqrt(lambda_min(Sigma)) that drives the entire theory. On this core we establish four theorems and one observation. T1 (identifiability): theta is globally injective on the quotient of the gauge kernel, with dimension bound F <= (K-1)^2. T2 (sparsistency): the l1-penalized estimator recovers the true support under irrepresentability and score concentration, with exponential failure probability. T3 (well-posedness): the feature-moment map M(theta) = Phi^T x_theta is strongly monotone, and the inverse is Lipschitz with constant L <= epsilon ||Phi^T S_a||_op / (pi_min lambda_min(Sigma)). T4 (convergence): local strong convexity with mu >= pi_min^2 lambda_min(Sigma) / epsilon^2 guarantees monotone gradient descent convergence. O5 (misspecification): the estimator converges to the OT-model projection of the truth; the Holder continuity of the projection map is assessed numerically, yielding setting-dependent empirical exponents alpha_eff in (0,1).
Zhanrui Dong, Tiefeng Jiang, Tuan Pham +1
We study the laws of large numbers for the largest eigenvalues among all principal minors of Wishart matrices and deformed GOE matrices. We propose a new method based on identifying the deterministic sets to which the random sets formed by suitably normalized principal minors converge in Hausdorff distance, thereby reducing the original extreme-value problems to finite-dimensional convex optimization problems. We demonstrate the effectiveness of this method in regimes not covered by the existing second-moment arguments in cai2021asymptotic,hu2023extreme. For deformed GOE matrices with fixed minor size k, we determine the limit for every diagonal variance a>0 and identify a phase transition at a=2. Above the transition, the limiting constant satisfies an explicit recursion with no close-form expression, and the optimizers exhibit a nested hierarchical structure, thereby resolving the case left open in cai2021asymptotic. For Wishart matrices with general sub-Gaussian entries and fixed k, we characterize the limit through an entropy-constrained deterministic convex set. When the entries are standard Gaussian, we solve the resulting optimization problem explicitly and obtain the exact value of the limiting constant.
Kairat Mynbaev, Carlos Martins-Filho, Chad Brown
We consider the classical additive measurement-error model X=Y+Z, where the latent random variable Y has unknown distribution FY and the error Z has a known distribution. We develop direct estimators for three functionals of FY: (i) FY(x) at continuity points; (ii) interval probabilities FY(y)−FY(x) when x<y are continuity points; and (iii) the size of a jump at a prespecified discontinuity. We derive non-asymptotic bias and variance bounds, and establish asymptotic unbiasedness and consistency. Unlike previous work, we do not require FY to admit a density, have a mixture representation, or satisfy global Sobolev smoothness assumptions. The framework accommodates arbitrary latent distributions, including those with both discrete and continuous components, and distributions with multiple jumps. These results rely on a link between Fourier inversion theorems and the algebraic structure of a class of estimators proposed in Mynbaev, Martins-Filho and Henderson (2022). A simulation study evaluates feasible tuning procedures and, where available, compares the finite-sample performance of the proposed estimators with existing methods.
Mårten Schultzberg, Mattias Frånberg
Bayesian inference for A/B testing is a family of prior and stopping-rule configurations with fundamentally different statistical properties, but it is often discussed as a single method, and no systematic overview exists. This paper organizes common configurations into a three-tier hierarchy: 1) posterior coherence with no error control, 2) false positive rates bounded under continuous monitoring via Bayes factor stopping, and 3) false discovery rate control and calibrated shrinkage via empirical Bayes. Many commercial platforms operate at the lowest tier by default. We show that Bayes factor stopping is near-optimal for a broad class of cost functions, including most proposed in the A/B testing literature; because the same rule also controls the false positive rate, the choice between a decision-theoretic and a frequentist formulation is largely one of parameterization. Furthermore, the empirical Bayes prior is the only path to the third tier, but winner-selected corpora, pooled programs, and heterogeneous metrics can each prevent calibration regardless of corpus size. Simulations against group-sequential and always-valid frequentist baselines show that flat-prior posterior stopping exactly reproduces naive peeking, that a well-calibrated empirical Bayes prior achieves the lowest estimation error, and that expected-loss stopping minimizes regret only when shipping a null-effect variant is nearly free. Error rates, estimation accuracy, and regret are all different risks, and the appropriate method follows from the risks an experimentation program needs to control, not the other way around.
Dmytro Marushkevych, Francisco Pina, Mark Podolskij
We study estimation of the drift matrix in a continuously observed high-dimensional Ornstein-Uhlenbeck process when the drift is exactly or approximately low rank. In this setting, exact low rank induces non-stable directions and hence a non-ergodic regime, resulting in a poorly conditioned empirical covariance matrix. To address this difficulty, we introduce a Weighted Nuclear Elastic Net Estimator that combines ridge regularization with a nuclear-norm penalty expressed in the empirical likelihood geometry. Under a general diagonalizable spectral framework, we establish oracle inequalities relative to arbitrary low-rank comparison matrices. For near low-rank drifts, the approximation error is naturally measured through the singular-value decay of the drift after weighting by the regularized empirical covariance. The stochastic term is controlled by self-normalized martingale arguments under appropriate choice of the tuning parameter. For a symmetric positive-semidefinite exact low-rank model, we verify the empirical-curvature condition required to translate the weighted bound into a Frobenius-norm bound. With an appropriate choice of tuning parameters, the resulting estimator satisfies, up to a logarithmic factor, the standard rank-r matrix-estimation scaling rd/T: specifically, its squared Frobenius error is of order rdlog(T)/T with high probability, under an explicit dimension-horizon condition.
Johannes Brutsche
We develop a stable convergence theorem for partial sum processes on sample-size dependent stochastic bases. The result allows multidimensional semimartingale limits that have conditionally independent increments and both a continuous and discontinuous martingale part. Motivated by infill asymptotics, it complements classical Gaussian stable limit theorems and supports applications to likelihood based statistical inference.
Austin Brown, Kshitij Khare
We develop a new uniform drift condition and local minorization that implies a stronger weighted form of uniform ergodicity for Markov chains we call hyper-V uniform ergodicity. The convergence guarantees geometric decay of the bias towards the invariant measure independently of the initialization for all functions controlled by a dominating function V. A key advantage of the approach is that it bypasses the need to establish a global minorization condition, which is often substantially more difficult to verify in practice, while yielding stronger convergence guarantees than global minorization. Optimal convergence bounds in a minimax sense of the framework are established. The utility of the framework is demonstrated through applications to the P'olya-Gamma and Kolmogorov-Gamma Gibbs samplers. We also show qualitative hyper-V uniform ergodicity convergence for two-variable Gibbs samplers can be inferred by the form of the invariant measure, bypassing convergence analysis entirely.
Hengzhi He, Guang Cheng
We consider a mixture of at most k unit-covariance Gaussians in Rd whose means belong to a fixed-radius ball, with no separation or minimum-weight condition. Doss, Wu, Yang and Zhou (2023) proved that the minimax Hellinger risk is of order d/n∧1 and constructed a proper polynomial-time estimator with the slower general bound (d/n)1/4; obtaining the sharp rate in polynomial time for fixed k≥3 was left open. We resolve this question. The key device is a moment-fiber range finder. A second-moment subspace controls the energy missed by projection. We then estimate finitely many one-free-index Hermite contractions. These vector-valued contractions recover every tensor component containing exactly one missed direction at the sharp d/n scale. Every remaining term contains at least two missed factors and is therefore controlled by the residual second-moment energy. The resulting subspace has dimension depending only on k. Exhaustive moment fitting in this constant-dimensional space produces a proper mixture and, together with the dimension-free moment characterization of Gaussian mixtures, achieves the optimal Hellinger rate in polynomial arithmetic time for every fixed k.
Aaditya Ramdas
An e-detector for a pre-change class P is a nonnegative process M such that EP[Mτ]≤EP[τ] for all stopping times τ and all P∈P. Thresholding e-detectors controls the average run length (ARL): declaring a change at the first time Tb when M crosses b ensures that infP∈PEP[T]≥b. But e-detectors do substantially more than control the ARL; they also satisfy a optional-horizon inequality: P(Tb≤σ)≤EP[σ]/b for every data-dependent stopping time (monitoring horizon) σ and P∈P. In particular, every e-detector-based procedure obeys P(T≤t)≤t/b at each fixed t, thus avoiding early false alarms. Remarkably, the converse also holds: every stopping time T that satisfies the optional-horizon inequality must in fact arise from thresholding an e-detector. We also derive a universal representation of stopping times that satisfy (only) ARL control. These are represented by weak e-detectors, that only require EP[Mτ]≤EP[τ] to hold at all threshold stopping times Tb. Appendices present universal representations for other (less common) change detection metrics.
Olga Klopp, Fedor Noskov
We introduce an intrinsic spectral sparsity model for nonparametric density estimation on compact connected Riemannian manifolds. Instead of penalizing coefficients in an arbitrarily chosen Laplace--Beltrami eigenbasis, we group each complete eigenspace and measure the Hilbert norm of its spectral component. The resulting block-variation space is basis independent and isometry invariant. We establish its structural, atomic, and nonlinear approximation properties and clarify its relation to Sobolev, Besov, and coefficientwise spectral ℓ1 classes. We then construct a coordinate-free block-shrinkage estimator and prove a nonasymptotic signal-dependent L2-oracle inequality that adapts to the unknown set of detectable eigenspaces. Under polynomial spectral growth, the risk theory separates the number of spectral blocks from their multiplicities and exhibits two regimes: one driven by a single high-dimensional eigenspace and the other by cumulative spectral complexity. Under matching spectral-growth and nondegeneracy assumptions, corresponding minimax lower bounds show that this multiplicity dependence is intrinsic, with sharp consequences for spheres and the rotation group SO(3). Finally, we develop a positive, normalized, block-penalized exponential spectral sieve for log-densities and derive likelihood oracle inequalities together with expected Kullback--Leibler, Hellinger, and L2 risk bounds. The resulting framework provides a geometry-respecting theory of sparse density estimation that remains invariant under changes of eigenbasis.
Ernesto Mordecki, Anas A. Rahman, Manuel Sáenz
We develop a probabilistic cavity-contiguity framework for proving replica-symmetric convergence of local marginals in mean-field Gibbs systems. The approach is based on cavity decompositions of the Hamiltonian together with direct comparison of probability measures through Radon-Nikodym derivatives and Hellinger-type estimates. At a conceptual level, the method separates the concentration of the relevant order parameters from the identification of the asymptotic cavity model and the comparison of the associated Gibbs measures. In contrast with interpolation-based approaches, the argument relies only weakly on the Gaussianity of the disorder and naturally accommodates concentration tools such as Poincaré and log-Sobolev inequalities. Rather than pursuing maximal generality, with the aim of making the method transparent, we implement the framework in a canonical example: the high-temperature Sherrington-Kirkpatrick model. In this setting, we prove that the marginal law of a fixed spin converges in total variation toward the effective one-dimensional cavity measure predicted by the replica method. Beyond the specific result for the Sherrington-Kirkpatrick model, the paper illustrates a broader cavity-contiguity methodology which is expected to extend naturally to other mean-field Gibbs systems, particularly Bayesian inference models with or without mismatch.
Hongru Zhao
Let R be the Pearson sample correlation matrix formed from n independent Gaussian observations in p dimensions, and write m=n−1≥p. Under the null correlation R=Ip, the classical independent beta product, exact cumulants, and full Fourier inversion yield, along every sequence p→∞ with m≥p, a uniform first Edgeworth expansion for logdetR, centered by its exact mean and scaled by its exact standard deviation. The expansion identifies the exact finite dimensional skewness correction and gives the sharp Kolmogorov equivalent Am,p/{62πVm,p3/2}, where Vm,p is the exact variance and Am,p is the absolute third cumulant. This equivalent unifies the square, fixed gap, growing gap, proportional, and dilute regimes; in the square regime the error has order (logp)−3/2 with an exact constant. For every positive definite population correlation matrix R, we prove a uniform finite sample Berry-Esseen bound that explicitly tracks population dependence. All theoretical results have exact or proved equivalent Lean 4 formulations whose declarations and dependencies are kernel checked.
Florent Krzakala, Pierre Mergny, Vanessa Piccolo
Recovering a low-dimensional latent subspace from nonlinear observations of Gaussian covariates in high dimensions is a fundamental problem in feature learning. Here, we consider Gaussian multi-index models in which the covariates xi∼i.i.d.N(0,Id) and the responses yi depend on xi only through its projection onto an unknown r-dimensional subspace. Earlier work based on approximate message passing (AMP) identified a sharp threshold for weak recovery [Troiani et al., 2025], raising the question of whether it can be attained, without side information, by a spectral method. We answer this affirmatively and develop a general random matrix theory for matrix-valued spectral estimators of the form Dn=n1i=1∑nT(yi)⊗xixi⊤, where T is an arbitrary bounded symmetric matrix-valued preprocessing map of fixed dimension. As n,d→∞ with n/d→α, we prove that the empirical spectral measure of Dn converges almost surely to a deterministic compactly supported distribution characterized by a matrix-valued self-consistent equation. We then establish a spectral phase transition for the largest eigenvalue: below threshold it sticks to the bulk edge, while above threshold an outlier emerges. We characterize the outlier location through a finite-dimensional deterministic equation and show that the associated spectral estimator achieves weak recovery of the latent subspace. Finally, we prove that the AMP-derived preprocessing of [Defilippis et al., 2025] is optimal among all bounded matrix-valued preprocessing maps of any fixed dimension. Its transition coincides with the AMP weak-recovery threshold, proving the general spectral conjecture of [Defilippis et al., 2025].