4,329 papers in this slice of arXiv.
Abdulrahman Qadi, Akash Sharma, Francesca Medda
Shariah-compliant equity screening provides a transparent setting in which institutional rules determine who may own a stock. A binary label identifies current eligibility but not whether the feasible investor base is fragmented across standards or close to changing. We define this instability as classification uncertainty and formalize its investor-base consequence through permitted investor mass. In a 1999-2024 CRSP-Compustat panel of 13,188 securities classified under seven researcher-emulated Shariah rulebooks, screening-rule disagreement and proximity to active boundaries rank next-month screen-implied transitions. U.S. Fama-MacBeth diagnostics do not support an unconditional equal-weighted permission premium, and a September 2023 DJIM/S&P methodology change produces no robust matched repricing. The central event evidence uses 25 official Securities Commission Malaysia lists. The 410 inclusions already trading before the preceding review have positive but imprecise matched returns. Applying the pre-event turnover floor yields 295 inclusions with 1.76 percentage points over [0,10]
Sebastian Frank, Jingrao Lyu, Max Jarmey +5
As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, identifying economically meaningful peer companies for comparison is a fundamental challenge for valuation, due diligence, portfolio construction, and risk management. We propose an ensemble tree-based supervised similarity learning framework that defines company similarity through the lens of market valuation rather than static feature matching or semantic descriptions. Specifically, we train a CatBoost gradient-boosted decision tree model on observed private company valuations and derive a valuation-aware similarity metric from importance-weighted leaf-node co-occurrences across the ensemble. The similarity metric captures shared valuation drivers while accommodating nonlinear relationships, mixed data types, and pervasive missing data common in private markets. Using a global private-market universe of approximately 270,000 companies, including more than 53,000 firms with observed or derivable post-money valuations spanning multiple industries, geographies, and deal stages, we demonstrate that the proposed similarity framework improves upon traditional distance-based and text-embedding-based approaches in downstream k-nearest-neighbor valuation tasks in the evaluated industry groups, while retaining case-based explainability.
Junyi Ye, Ivy Gateri Wanjiku
Financial forecasting models are typically developed in full precision, yet production deployment often requires low-precision inference to reduce memory and computational cost. Post-training quantization (PTQ) enables such deployment without retraining. However, reliable activation quantization requires calibration: activation ranges are estimated from historical data before deployment and then remain fixed during future inference. The importance of this deployment choice for financial forecasting remains poorly understood. We present a systematic study of activation calibration for PTQ in cross-sectional volatility forecasting on the S&P 500. Our evaluation covers seven representative neural architectures, eight walk-forward test years (2018-2025), and 560 trained models. We find that activation calibration has little effect at 8 bits but becomes the primary determinant of predictive performance at 4 bits. Under default absolute-maximum (abs-max) calibration, static 4-bit quantization of both weights and activations removes 11-62% of the full-precision mean information coefficient in affected architectures. Replacing abs-max with percentile calibration recovers 53-94% of this degradation in the four most affected architectures. The preferred activation range also varies across market periods. Narrow ranges improve resolution under typical market conditions but lose part of their advantage when test-period market dispersion exceeds the calibration history. These findings show that activation calibration is a first-class deployment decision for reliable 4-bit PTQ in financial forecasting. When substantial degradation remains, 8-bit activations or weight-only 4-bit quantization provide more robust deployment choices.
Junyi Ye, Gargi Vijay Borde
Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training. This paper asks where such information should enter a neural cross-sectional volatility forecasting model. We study five-day realized-volatility forecasts for 1,027 U.S. equities using a rolling walk-forward evaluation framework in which information, model capacity, hyperparameter tuning, and random seeds are matched across architectures. We propose RG-ResMoE, a regime-gated residual mixture-of-experts architecture in which regime information is used only for expert routing rather than for direct forecasting. The base predictor models volatility from stock features, while a gating network uses regime state variables to route residual corrections. RG-ResMoE consistently outperforms a capacity-matched MLP in both forecasting accuracy and training stability in the main U.S. study. Similar gains are observed on an independent Japanese panel. The integration pathway is decisive: appending the same regime variables directly to the forecasting input degrades both predictive performance and training stability, whereas restricting them to the routing gate improves accuracy and Value-at-Risk calibration. Hard routing consistently underperforms soft routing. The results suggest that, in compact neural volatility forecasting models, the primary value of mixture-of-experts models lies less in increasing model capacity than in controlling how nonstationary regime information influences prediction.
Ekkehardt Bauer, Dirk Holländer, Linus Wolff +3
This study focuses on developing an AI-supported prototype for multiperspective interest rate forecasting that combines classical econometric models with modern artificial intel-ligence methods. Tested in a major European bank, the system enables more precise and flexible prediction of interest rate developments, supporting strategic decision-making in Asset-Liability Management (ALM). It integrates topic modeling, sentiment analysis, econometric forecasting, and market-based analyses within an interactive platform. Leveraging AI to analyze large volumes of financial documents and market data enables the identification of monetary policy trends and sentiment signals at an early stage. The core econometric model is a Bayesian vector autoregression (BVAR) that enables simulation-based scenario analyses to evaluate economic developments from multiple perspectives. The system's innovation lies in its integration of several forecasting approaches that consolidate previously separate information sources and present them transparently and interpretably. Financial analysts and risk managers thus gain a better basis for making decisions, allowing them to assess interest rate risks more accurately and manage market movements more proactively. While the prototype demonstrates how AI can transform interest rate management in banking, further development is required to optimize real-time data integration and regulatory compliance. Even at this stage, the study shows that multi-perspective, AI-driven forecasting provides substantial added value for banks by increasing transparency, strengthening evidence-based decision-making, and improving risk management.
Yannik Pitcan
Studies of association-football forecasting routinely report three-way accuracy in the low fifties and present it as competitive with the betting market. Accuracy against a uniform benchmark answers the wrong question; the question worth asking is whether a model carries information a margin-free closing price has not already absorbed. We formalise that test as the fitted weight in a logarithmic opinion pool and apply it to nineteen complete Serie A seasons (7,220 matches). The answer is negative and stable. A Dixon-Coles model with tuned exponential decay attains 53.4% accuracy and a Ranked Probability Score of 0.1972 against the market's 0.1905; the paired difference is +0.0067 (95% CI [0.0046, 0.0088]) and the market wins in all seven test seasons. The fitted pooling weight on the structural model is 0.000, and the log-loss profile is monotone increasing in that weight on validation and test alike, so this is a boundary solution, not an optimisation artefact. Refitting the same machinery to shots on target yields a variant earning weight 0.35 against the goals model -- it carries information the goals model lacks -- and 0.000 against the market. Two structural signals, each informative about the other, both priced. The structural model is better calibrated than the market on the home-win margin (slope 0.995 versus 1.103) while clearly less sharp: the market's advantage is discrimination rather than honesty, which accuracy alone cannot distinguish. Value lies not in a better forecast but in what is built on a calibrated one. We define match leverage, the change in a club's probability of achieving a season objective between winning and losing a fixture, and compute it for ACF Fiorentina: an away fixture against a relegation rival carried 2.25x the leverage of hosting the eventual champions. The paper also documents and corrects errors in an earlier study of our own.
Jaesung Kim, Changhee Cho, Jae Woo Lee
This study investigates whether the macroscopic statistical maturity of cryptocurrencies implies dynamical equivalence with traditional equity markets. We analyze high-frequency data (2020--2025) using the Complexity--Entropy Causality Plane (CECP) and directed horizontal visibility graphs (directed HVG) to uncover complex temporal patterns and time-directed structures in the return series. While conventional stylized facts show striking convergence across all assets, structural diagnostics reveal a compelling paradox: cryptocurrencies appear more locally random than the equity benchmark during ordinary periods, yet exhibit significantly stronger directional time-irreversibility around high-visibility return events. The absolute-return results show that large cryptocurrency fluctuations tend to begin abruptly and remain elevated afterward. Separate analyses of positive returns and negative-return magnitudes show that this pattern is shared across cryptocurrencies on the upside but varies across assets on the downside. We conclude that statistical maturity is only skin-deep; the underlying dynamical processes of mature cryptocurrencies remain fundamentally distinct from traditional benchmarks.
Alberto Acedo
The Triadic Stress Index (TSI) takes a network index whose four factors were first observed in soil microbiome co-occurrence networks and applies it, without alteration, to the correlation network of financial assets. We test it on five markets spanning 2006-2026 (equities including banking crises and the AI sector, cryptocurrencies, commodities, foreign exchange and sovereign debt), against three independent definitions of a crisis episode, at a fixed alarm budget, out of sample, with block-bootstrap intervals and a Holm correction across the family of tests. The benchmarks are the Absorption Ratio, the industry standard used by MSCI and central banks; the effective rank and the Vendi score, the sharpest spectral measures available; Ollivier-Ricci curvature; and the global and local balance indices of signed correlation networks. Three comparisons favour the index. It carries a per-node decomposition, diag(A^3), naming which asset is carrying the concentration with no parameter to select, and scores 0.97-0.99 against 0.33-0.84 for the only published per-node alternative, whereas spectral attribution must first choose how many components to read and collapses under a standard but wrong choice. Its alarms are the cleanest of anything tested, 4.0% of them with no matching episode against 14.7% for the effective rank and roughly 59% for the Absorption Ratio. And it beats the Absorption Ratio on detection by 0.273 in F1 out of sample, p<0.0005. The remaining comparisons are ties. Against the effective rank and the Vendi score the index ties in every scheme and both samples, and the margin over the Absorption Ratio narrows under the strictest labelling. On real matrices the far simpler node degree reproduces the attribution. A lead-lag analysis puts the peak cross-correlation at zero lag: this is a coincident state index, not a forecast.
Lukasz Adamski, Robert Slepaczuk
Our primary goal is to forecast and empirically examine the evolution of the implied volatility (IV) surface, with particular focus on the dates of scheduled meetings of the Federal Open Market Committee (FOMC). Firstly, we check if IV increases before the announcement and if thes effect is stronger for short-dated, out-the-money (OTM) options in high volatility regimes. In the second part, we turn the focus to verifying if the ML framework can beat the benchmark random walk in forecasting this effect. A feature related to dates of scheduled FOMC meetings augments the model, which allows us to discover if it can learn the effect of elevated pre-announcement uncertainty. Our contribution relies mainly on the quantitative prediction of the pre-announcement effect and the inclusion of exogenous information inside the ML framework used for the IV surface forecasting. It is also on of the first attempts to apply ML models directly on the IV surface without relying on dimensionality reduction. To achieve this, we employ a convolutional two-dimensional LSTM model, which is capable of learning spatio-temporal signals in the surface. Our analysis reveals that the edge of the ML framework can be limited due to the noisy characteristics of the IV surface. Nevertheless, our study reinforces the perspective that ML models can effectively forecast the IV surface also during abnormal days.
Rosanna Grassi, Caterina Pastorino, Pierpaolo Uberti
In this paper we investigate the information content of the lower part of the spectrum of financial correlation matrices, as a source of information on market synchronization. In a financial context, a classical application of Principal Component Analysis and Random Matrix Theory identifies the largest eigenvalues as indicators of dominant market factors and synchronization patterns. We complement this perspective by showing that the smallest eigenvalues also contain relevant information about the effective structure of financial markets. The paper presents the methodological proposal and validates its effectiveness through comprehensive real data experiments in both descriptive and predictive settings.
Kundan Mukhia, Sabat Rai, Vivek Shrivastav +2
Stablecoins have rapidly emerged as an important class of digital assets and a component of the digital financial ecosystem. Despite their growing importance, the statistical properties of stablecoin transaction activity remain largely unexplored. To the best of our knowledge, this is the first study to investigate scaling behavior in stablecoin transaction data, focusing on USDT and USDC. We analyze approximately 370 million USDT and USDC transactions recorded on the Ethereum blockchain across six periods spanning June 2024 to February 2026. Based on interactions between Externally Owned Accounts (EOAs) and Smart Contracts (SCs), we classify transactions into four categories: EOA-EOA, EOA-SC, SC-EOA, and SC-SC. Using maximum-likelihood estimation of power-law exponents, we find that transaction value distributions exhibit heavy-tailed scaling for both stablecoins across all periods and interaction categories. We identify two distinct scaling regimes: EOA-involved categories cluster around 1.45-1.60, whereas SC-SC transactions exhibit higher exponents of approximately 1.72-1.73. Sensitivity analysis confirms that this separation is robust across periods, stablecoins, and fitting sample sizes. Counterfactual analysis shows that changes in category weights alone cannot explain the observed variation in the overall exponent. Across different sample sizes, the counterfactual path accounts for only about 10%-35% of the total temporal range observed in the actual data. Overall, our results indicate two broadly differentiated scaling regimes in the tail of stablecoin transaction values. Power-law tail behavior is observed throughout stablecoin transaction activity, but the exponent depends on whether transactions are driven by EOAs or SCs. These findings provide a basis for further research on scaling behavior and transaction heterogeneity in blockchain-based financial systems.
Kasun Dewage, Suranadi De Silva, Shankhadeep Mondal
Foundation models for time series forecasting demonstrate impressive zero-shot generalization but often underperform on specialized domains such as high-frequency finance. We present a comprehensive study of hybrid neural-classical correction for adapting frozen TimesFM (200M parameters) to stock return prediction during the volatile opening trading hour. We compare two neural correction architectures - AttnCorrect (multi-head self-attention, approximately 471K parameters) and GatedLinear (low-rank bilinear projection with gating, approximately 49K parameters) - each augmented with Random Forest residual learning. Through systematic ablation across 10 major technology stocks (NVDA, MSFT, AAPL, GOOG, GOOGL, AMZN, META, AVGO, TSLA, NFLX) spanning 2 million data points, we reveal critical insights: (1) The hybrid neural-classical approach achieves 0.597 pooled correlation and 6.4x mean per-day correlation improvement over frozen TimesFM; (2) Classical residual learning (Random Forest) provides the largest single-component contribution, matching or exceeding the neural correction component; (3) Simpler neural architectures surprisingly outperform complex ones when classical residual learning is removed; (4) Self-attention provides the largest neural-only contribution. GatedLinear+RF achieves best overall performance with 9x fewer neural parameters than AttnCorrect+RF. We report three complementary correlation metrics - mean per-day, cross-day cumulative, and pooled - to provide a complete picture of predictive quality. Our results provide practical guidance: effective foundation model adaptation requires careful integration of neural and classical components, with classical methods playing a crucial complementary role.
Peter Cotton
We consider a market maker who can only obtain and dispose of inventory by responding to a sequence of sealed-bid enquiries, and whose customers arrive with imbalanced intent: sellers more often than buyers, or the reverse. Under the assumption that the best competing response is exponentially distributed around a commonly discerned fair price, we observe a symmetry in the steady state solution that compresses the imbalanced problem onto the perfectly balanced one. Order imbalance is absorbed, exactly, by a translation of the market maker's skew, a widening of her quotes, and a multiplication of her effective cost of carry. The adjustment is simple even though the solution it adjusts is not, and it involves no free parameter beyond the observable market width. The exponential assumption is needed only locally, at the quotes actually made, and the width that enters is the locally observed one. Among the consequences: a market maker with zero inventory should still skew; skew responds to imbalance at first order whereas width responds only at second order; and the popular "constant width, linear skew" heuristic is recovered as the small-skew solution in the special case of balanced flow and quadratic holding cost.
Sara Chehab, Giorgos Iacovides, Parisa Yazdanparast +1
Current portfolio construction methods are either agnostic to the effects of idiosyncratic shocks (standard factor models) or to the latent data structure driving systematic returns (recent graph-based approaches). This presents an opportunity to combine the complementary market aspects captured by the factor and graph domains, allowing asset allocations to operate directly on the underlying market structure, rather than on its observed co-movement or its finite-sample artefacts. In this work, we introduce the Mutually-INformed Graph-Locality and Exposures framework (MINGLE), which mutually regularises the factor and graph domains by redefining graph locality through systematic factor exposure profiles, rather than via observed co-movements. This is formalised through a unified Alternating Direction Method of Multipliers (ADMM) framework that jointly learns a latent factor representation and its induced graph topology directly from market returns. The resulting exposure-similarity graph aligns more closely with established economic sectors than conventional correlation-based graphs. Portfolios constructed from this representation are shown to consistently outperform their correlation-based counterparts across a range of volatility regimes and transaction cost levels. For rigour, paired statistical testing confirms that these gains stem from the reconciliation of the graph and factor domains.
Klaus M. Frahm, Dima L. Shepelyansky
Using official government data sets of USA and France we analyze the occurrence/frequency/popularity distributions of given names on a time scale of more than 100 years. These distributions are characterized through the Lorenz and Pareto curves broadly used in the analysis of wealth inequality in the world. These curves remain stable during the considered time period with the Gini coefficient remaining in the narrow range 0.85-0.95. As for the case of wealth inequality, we show that the distributions of names are well described by the Rayleigh-Jeans (RJ) thermalization and condensation phenomenon well studied in various physical systems. The RJ thermalization results from two integrals of motion being analogous to energy and probability norm conservation in physical systems with energy states corresponding to popularity levels of names. Time correlations between names are also determined showing their stability until the middle of the twentieth century and a significant change after that.
Julius Döbelt
Predicting financial asset returns remains one of the most difficult challenges in empirical finance, driven by the low signal-to-noise ratio and the semi-strong form of market efficiency. While deep learning models, especially LSTM networks, have shown promise in capturing temporal dependencies, standard architectures often struggle to account for the cross-sectional heterogeneity of asset returns. This paper proposes a novel architectural extension to the basic LSTM model designed to improve both predictive accuracy and model interpretability. The framework integrates macro-financial covariates to capture broader economic signals and learnable sector embeddings to encompass heterogeneity by sector. The trading strategy involves constructing a long-short portfolio based on daily directional forecasts for each S&P 500 constituent, targeting stocks expected to under- or outperform the cross-sectional median return of the S&P 500. Model Performance is evaluated against three competitive benchmarks: a basic LSTM, a Random Forest model and a traditional market buy-and-hold strategy. The empirical results demonstrate that the LSTM with sector embeddings outperforms all benchmarks across key risk and return metrics. By utilizing sector embeddings, the model explicitly incorporates cross-sectional heterogeneity, allowing it to adapt to varying industry dynamics within the market. To address the black-box nature of deep learning, I use latent space visualizations to analyse how the model differentiates between sectors, providing insights into the internal representation of the sectors in the LSTM. The impact of the sector information can be quantified using a novel contribution metric by inspecting the weights of the LSTM. The predictive signal is driven by a short-term reversal factor and an industry momentum factor.
Alex Chen, Maria Hybinette
Intraday market manipulation is hard to detect because its footprint is brief, buried in millions of quotes, and statistically similar to ordinary volatility. Detectors reach high recall only by flagging so many other days that measured precision collapses, producing alerts no regulator can act on. We show that this manipulation leaves a distinctive dynamic signature: a pump-and-crash pattern visible in the velocity of market state, rather than its level. We build a minute-level detection pipeline, strictly partitioned in time, based on smoothed state velocity: option-Delta velocity for index options and price velocity for equities. We explain every alert with SHAP attribution. We hold the test period strictly out-of-sample and fix all thresholds before evaluation. On the locked Indian BANKNIFTY index-options test, the plain autoencoder recovers 10 of 10 regulator-identified manipulation days. Conditioning detection on market regimes inferred by a hidden Markov model yields an instructive negative result. The regimes are descriptively distinct, but using them trades recall for precision. Under the closed-world assumption that unlabeled days are normal, precision remains near 25%. The same dynamic appears in thinly traded U.S. equities (SEC v. Patel). The shape of the signature survives the transfer; its velocity magnitude does not. A pump-reversal shape score ranks the complaint's alleged manipulation days with AUC 0.91 (ARQQ) and 0.81 (ACY). On the ARQQ worked example, the score peaks inside the complaint's documented minute window. Finally, exact SHAP attribution over every alert shows that unconfirmed alerts share the regulator-identified days' attribution profile (cosine similarity 0.99). The precision ceiling is consistent with incomplete enforcement labels rather than detector failure. What transfers across markets and instrument types is the dynamic signature itself.
Ramon Marc Garcia Seuma
We study seven major crypto-perpetual liquidation cascades (2022-2025), and in the largest of them we can watch the mechanism directly. From the on-chain fill log of a fully transparent venue we measure the branching ratio of that event -- the October 2025 crash, the largest on record -- in flight, with both of its factors observed and no free constants. It ran deeply subcritical: the structural ratio and the amplification bookkeeping both place it at λ^≈0.1−0.2 throughout, while a third, flow-based estimator falls through the climax rather than rising. All three agree on subcriticality within the venue, not on a common numerical level. Alongside them, 88% of all post-onset forced selling landed within thirty minutes and 63% of it was absorbed off-book by the venue's backstop, which drives the branching ratio down precisely at the climax. Across the full set of seven, at onset -- the minute ending the steepest hour of each crash -- the order parameter (mean inter-asset coupling) jumps by between 1.6 and 4.4 baseline standard deviations into a near-fully-ordered phase, while the susceptibility proxy χ collapses in five of the seven events and diverges in none; the jump is invariant under subsampling. The transition is abrupt and scale-robust rather than critical, and its in-cascade signature lives in the liquidity sector: price impact spikes on two venues and two instruments while open interest clears by 25-70%. The natural mechanistic account, a Galton-Watson cascade with λ=kρ~, is then eliminated as a description of the pre-cascade state: both of its falsifiable predictions fail at simulated power >= 0.96, on proxied and on directly measured regressors alike. Severity is set by shock times map-in-path times liquidity withdrawal rather than by a diverging multiplier, which is why none of the scalar pre-state measures we can construct grades it.
Vance Martin, Yoshihiko Nishiyama, John Stachurski +1
We introduce a new density-based goodness of fit test for ergodic Markov processes. Our test compares the data against the class of models specified in the null hypothesis, and rejects if no model in the class yields a stationary density that matches with the data. No alternative needs to be specified in order to implement the test. Although our test compares densities, estimation of smoothing parameters is not required, and the test has nontrivial power against 1/n local alternatives. The test provides new perspectives on some existing problems in econometric and financial modeling.
Giulia Livieri, Gianluca Palmari
Observation-driven filters update a time-varying parameter with the likelihood score, linking the recursion to the logarithmic scoring rule. We replace this update with the negative parameter derivative of a differentiable proper scoring rule, within a declared working family and predictable scaling. For a general rule, the conditional mean update is a pre-conditioned stochastic-gradient of conditional scoring risk; when an autoregressive pull is included, the centre is the zero of a composite mean field. We derive local realised-loss descent and conditional-mean contraction results, and decompose the local dynamics into risk curvature and innovation variability. These two quantities coincide for the log score under the Bartlett identity but generally differ, which clarifies how bounded drivers can limit the transmission of extreme observations to the filtered path. We also establish consistency and dependent-data sandwich asymptotic normality for batch minimum-scoring-risk estimation of the static recursion parameters. For high-frequency scale models, centred updates yield a diffusion limit, non-centred updates yield a mean-flow limit, and a local transfer theorem gives an Ornstein-Uhlenbeck approximation around a moving scoring-rule risk projection. The working family, scoring rule, scaling, and autoregressive parameters jointly determine the filtered path, while the working family also determines predictive quantiles. Controlled experiments illustrate these channels. An empirical density-by-criterion factorial on international equity returns evaluates point-variance loss, value-at-risk coverage, and probability-integral-transform diagnostics without imposing a universal ranking. The results provide a criterion-based framework for observation-driven filtering under misspecification and identify the assumptions required for estimation, local tracking, and continuous-time approximation.