7,453 papers in this slice of arXiv.
C. P. Barrington-Leigh
The dispersion of self-reported life satisfaction has been proposed and used as a comprehensive measure of societal inequality. A negative cross-country association between mean life satisfaction and its standard deviation has been read as evidence that this inequality is itself welfare-relevant, but critics have pointed to the nonlinearity and boundedness of response scales. I describe a further problem: a substantial and predictable share of respondents simplify the 0--10 scale to the subset {0, 5, 10} --- "focal-value rounding" (FVR) --- a behaviour that breaks not only linearity of the scale but even its order. I derive a sharp bound on the bias that FVR can induce in the standard-deviation; at empirically-typical FVR fractions the possible bias is roughly half the cross-country range of dispersion. I estimate and correct for the bias by fitting a model of FVR behavior to the Gallup World Poll Cantril ladder and three other multi-country life evaluations. FVR inflates measured standard deviation in 133 of 135 countries, by a median of 0.090 points. However, because the corrections are nearly uniform in sign across countries, rankings of the standard deviation survive essentially intact. The ordinal inequality indices advocated as the theoretically appropriate alternative suffer even worse from FVR. The correlation between mean and standard deviation survives FVR correction with modest attenuation and some caveats. Thus FVR does not, by itself, overturn previously reported correlations; it does, however, add a further reason for caution in interpreting scalar measures of subjective wellbeing inequality: candidate statistics are differently contaminated, the ordinal repairs no less than the cardinal originals, and they disagree with one another. Lastly, I show that FVR-correction can improve the pairwise dominance-comparability of full distributions.
Zihao Zhang, Yuanbo Zhang, Xiaolei Ma +1
Urban decarbonization often raises the cost of travel, yet which neighbourhoods can adapt remains largely invisible under normal conditions. We leverage the 2026 US-Iran oil shock as a natural experiment, applying a hierarchical panel regression discontinuity design to 1.7 trillion point-of-interest visits across 122,000 neighbourhoods in China and the United States. Mobility range declined in nearly three-quarters of neighbourhoods, but responses varied systematically with pre-shock urban conditions. Exposure to energy-intensive travel explained the largest share of modelled heterogeneity in both countries, while adaptive capacity and activity composition further shaped how travel was reorganized. Longer baseline travel intensified contraction, whereas greater car dependence constrained adjustment. Crucially, similar mobility outcomes arose from different processes: some neighbourhoods maintained travel by absorbing higher costs, whereas others appeared structurally locked into travel they could not reorganize. Fuel-price shocks, therefore, act as urban stress tests, revealing otherwise hidden inequalities in mobility adaptation.
Aaron Chatterji, David Holtz, Neel Rakholia +2
We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026. These linked data enable a privacy-preserving analysis of adoption, worker roles, and message-level tasks at scale: for instance, the worker-level sample we analyze at the six-month adoption horizon includes over 1,500 organizations and over 17 million messages. We document four facts about enterprise AI adoption and use. First, ChatGPT Enterprise usage has grown rapidly due to a combination of new firm adoption and growing intensity among existing adopters. Second, U.S.-based public company adoption is concentrated among larger, more valuable, and more R&D- and SG&A-intensive firms. Third, active use within adopting firms spans job functions and seniority levels, with especially high usage intensity among early-career workers. Fourth, ChatGPT Enterprise usage encompasses a broad range of knowledge work tasks, including writing, technical work, communication, and information synthesis. In aggregate, these results suggest that firms differ widely in the speed, breadth and purpose of their enterprise AI adoption, and that they are still actively learning how to integrate AI into organizational workflows.
Juergen Renn
Designs for international climate cooperation face a trade-off between allocative efficiency and robustness to the erosion of institutions by defection, renegotiation, and political turnover. We formalize this trade-off in a stylized coalition-formation game in which membership is driven by two market-based channels, a membership premium and an outsider drain, and is stabilized against a bounded set of institutional perturbations using the apparatus of robust control. The free-rider gap that each member faces is derived from economic primitives, the carbon price differential and the emission intensity of exports, and separated from the architecture-specific channels; the separation makes the model solvable and comparative. The model is bistable: a small remnant club and a near-universal coalition are separated by a critical mass. The established coalition withstands roughly twelve times the perturbation admissible at ignition, and the tipping survives heterogeneity up to a critical disorder of about seven times the membership premium. Two robustness coordinates, an ignition cost and a collapse threshold, place each entry architecture in a map with three regimes: self-igniting, founding-dependent, and permanent-support-dependent. The carbon currency, a newly proposed quantity-based system whose emission rights are reissued each period, extinguished upon use, and enforced at the border, is self-igniting; the border adjustment is founding-dependent; the export rebate is permanent-support-dependent. Architectures without an outsider drain, such as clean-up certificates and emissions-trading-system linking, form a separate class that improves efficiency within a coalition but cannot drive its formation. Robustness governs whether cooperation forms and endures; efficiency decides only how much an established coalition delivers.
Gregor Schubert
This study proposes that firms move along an "organizational technology ladder": adopting one technology transforms hiring and work processes and builds skills and organizational capital that change the cost of adopting subsequent technologies. I study how firms' adoption of remote work technology during the COVID-19 period shaped later uptake of generative AI. Using U.S. job-posting data and an instrumental-variables strategy based on predicted differences in labor-market pressure to offer remote work, I estimate that a 10 percentage point increase in remote hiring in 2021-2022 increases the share of job postings mentioning generative AI in 2023-2024 by 0.4 percentage points across firms and 0.7 percentage points across occupations within firms. I provide evidence on mechanisms consistent with a technology-ladder channel: remote work adoption shifts hiring toward technical and managerial capabilities that predict faster conversion of generative AI exposure into adoption. Firms with return-to-office mandates---interpreted as revealing low remote productivity---exhibit a substantially larger response of generative AI adoption to remote work, consistent with an organizational frictions channel.
Yun-Long Zhang, Jia-Ning Kang, Xiaoming Kan +5
Decarbonizing existing coal-fired power plants can contribute to near-term climate mitigation, but identifying cost-effective retrofit strategies is complicated by interactions among mitigation technologies. Here we develop an interaction-aware optimization framework that jointly evaluates energy conservation, biomass co-firing, and carbon capture across 1,885 coal-fired power plants in China while accounting for plant heterogeneity and shared biomass and CO2 storage resources. We find that technology interactions alter both mitigation costs and the emission reductions attributable to individual measures, thereby changing cost-optimal technology portfolios and marginal abatement cost curve at the fleet level. Approximately 1.2 Gt CO2 yr-1 can be mitigated at negative marginal cost, while reaching carbon neutrality requires a marginal abatement cost of US56 t CO2-1. Progressively deeper mitigation shifts the cost-optimal portfolio from energy conservation toward biomass co-firing and ultimately carbon capture, with biomass combined with carbon capture enabling net-negative emissions. Explicitly accounting for interactions among mitigation technologies therefore provides a more consistent basis for evaluating coal-power decarbonization and coordinating retrofit investment, infrastructure development, and climate policy.
Hongseok Choi, Jeongbin Kim, Matthew Kovach +3
We study how differences in AI-generated financial recommendations are transmitted into individual portfolio choices. In an experiment with 400 employed adults enrolled in workplace defined contribution pension plans in South Korea, participants allocate a hypothetical pension balance across eleven products and may revise it after receiving one of two fixed AI-generated recommendations. A 2×2 design randomizes recommendation content and whether the recommendation includes a short rationale. Approximately 37% of the experimentally induced difference between the aggressive and conservative recommendations passes through to final portfolios. This causal contrast changes expected portfolio return, volatility, allocations across risk grades, and the number of products held, but produces no detectable difference in computed Sharpe ratios. 81% of participants revise. Among revisers, 95% move toward the assigned recommendation and implement about half of the suggested adjustment. Rationales do not detectably alter pass-through. These results show that users partially and selectively transmit recommendation content into economically meaningful differences in risk exposure while retaining substantial weight on their initial choices.
Yves Achdou, Johannes Brumm, Lukas Frank
We propose a comprehensive framework for solving overlapping-generations (OLG) models in continuous time with both idiosyncratic and aggregate risk. Our general characterization of equilibrium through the master equation operates on the joint distribution over the continuous idiosyncratic states, age and wealth. Our computational strategy is to take a finite-dimensional representation of this distribution as an input of a neural net which in turn outputs a finite-difference representation of the (conditional) value function. This idea can be applied generally to heterogeneous agent models with aggregate risk, and we call it finite-difference neural operator. Our method combines advantages from modern neural nets and traditional finite-difference methods: It is grid-free in the high-dimensional distribution, and retains control on boundary conditions in low-dimensional state variables. Moreover, our method is able to enforce shape constraints. We showcase its flexibility by solving a continuous-time OLG model with aggregate risk alone where we characterize the distribution by its supporting function; and to an OLG model with both types of risk.
Henri Keränen
The New Keynesian example of Beraja (2023) is not identified at its printed calibration: more than one structure satisfies every condition of its identification step. Two of the paper's six identifying restrictions coincide on equilibrium equations consistent with the printed reduced form, leaving eleven conditions to determine an equation's twelve coefficients. The note derives the rank condition the identification step requires, computable from the reduced form and the restrictions alone. The note also documents and corrects a misprint in the paper's counterfactual display. A one-clause amendment to the paper's Theorem 1 restores its conclusion.
Johanna Bolaños-Zuñiga, Alberto J. Lamadrid
In this study, we use electricity demand growth, cooling requirements, and backup system operation to evaluate the environmental and economic implications of artificial intelligence data centers in the United States. Our results indicate that impacts are not determined solely by facility design, but by the broader electricity, water, and land-use systems in which these facilities operate. Emissions are primarily driven by electricity consumption and therefore depend on marginal generation mixes, transmission constraints, and the spatial and temporal distribution of demand. Analysis further shows that local effects include pressures on water resources, increased noise exposure, and land-use changes, with outcomes varying across regions and infrastructure conditions. The assessment of technological and operational measures shows that improvements in energy efficiency, cooling configurations, and operational strategies can reduce these impacts, although their effectiveness depends on system-level conditions. Evaluation of regulatory and market structures suggests that existing frameworks may not fully account for location- and time-specific externalities. These findings support the need for integrated policy approaches that align data center deployment and operation with electricity system characteristics, water availability, and land-use planning to improve overall environmental and economic performance.
Kwan Soo Shin
National planning counts population, human capital, and artificial-intelligence preparedness in separate ledgers. Demographic accounting has advanced from headcount to skills-adjusted stocks and still debates how much age structure retains once skills are modeled, yet no existing unit carries the conditions under which preparedness becomes productive capacity. This study introduces the Effective Cognitive Population (ECP), a decomposable unit that weights population by capability and by the conditions under which capability is deployed, anchored to the World Bank Human Capital Index Plus (HCI+) and the non-overlapping dimensions of the IMF AI Preparedness Index. The architecture is portable in principle; the case tested here is artificial intelligence, which has a published preparedness index. For 144 countries, HCI+ becomes a productivity level, AI opportunity uses digital infrastructure and innovation integration, conversion governance uses regulation and ethics, and the benchmark is ECP = N H(1 + AC). Against 2024 total output on identical population bases, ECP raises criterion R-squared from 0.849 for the HCI+-adjusted stock to 0.882 and lowers leave-one-country-out RMSE from 0.723 to 0.641, with the working-age comparison identical and bootstrap intervals excluding zero. Eighty-nine of 144 countries move at least ten rank positions from headcount, mostly through the human-capital adjustment itself. Results are stable across denominators, vintages, aggregation forms, and a 27-rule multiverse. The direct A by C interaction is not statistically supported, so the conjunction is a planning rule rather than causal complementarity. ECP is a diagnostic ledger whose scope excludes forecasts of population decline and estimates of AI's causal productivity effect.
Harsha Dutta, Pulak Ghosh, Arkodipta Sarkar +1
We examine the effect of political power-sharing on local economic activity. This effect depends on the relative importance of the risks associated with unchecked power and the potential efficiency gains or losses arising from checks and balances. Our research design exploits a geographic discontinuity design due to the haphazard overlap of electoral and administrative boundaries that generates quasi-random variation in the number of politicians governing adjacent regions. We supplement this design using an episode of electoral delimitation that allows us to exploit within-region variation in the number of politicians. We find increasing the number of politicians governing an area can lead to new firm creation, lower unemployment, and greater real economic activity. Our results suggest that non-aligned multiple politicians enhance state efficiency by imposing checks and balances on each other, leading to lower regulatory obstacles, less cronyism, and improved provision of public infrastructure, creating an economically favorable environment for firm creation.
Akhil Rao
The near-term growth of the commercial space economy and sustainability of the low-Earth orbit (LEO) environment depends on the commercial prospects of large LEO satellite telecommunications constellations (``mega-constellations''). If successful, mega-constellations can spur competition between launch providers, which may lead to lower launch prices. They may also demand orbital sustainability services and/or generate additional collision risk, which may in turn affect space sustainability interests. Consumer demand for mega-constellations' internet services is critical to their commercial success. This analysis presents a method to estimate consumer demand for mega-constellations' internet services, using SpaceX's Starlink over December 2021--November 2023 as an example. This analysis finds that Starlink users are highly concentrated in the wealthiest countries. It also finds that consumer demand for Starlink is growing more slowly than orbital capacity is being added and more slowly than demand for non-Starlink internet overall. Continuation of these trends may indicate that the market for mega-constellation internet is less lucrative than previous forecasts have indicated.
Haruka Nagamori, Kazuhiko Nishimura
Section 1502 of the Dodd--Frank Act, enacted in 2010, requires U.S.-listed companies using tin, tantalum, tungsten, and gold (3TG) from the Democratic Republic of the Congo and adjoining countries to disclose information on the minerals' origins. Concerns have been raised that the regulation may have induced a de facto embargo through avoidance of sourcing from the covered region. However, how the price responsiveness of mineral exports evolved under changing institutional and market conditions remains insufficiently understood. Since tungsten production in the covered region is concentrated almost entirely in Rwanda, this study examines the price responsiveness of Rwandan tungsten exports from January 2009 to December 2023. Because missing export quantity data prevent continuous observation of export unit values, we apply the identification approach of Nakano and Nishimura (2025), combining monthly mirror trade data from UN Comtrade with exchange rates and a world average price. An importer fixed-effects model is estimated using export value as the dependent variable, with the sample divided into four periods according to changes in the institutional and market environment. The results reveal substantial temporal variation in price responsiveness. A statistically significant negative price response is observed in Period 1 (η=−20.814, p<0.01), disappears in Period 2 (η=1.814, p>0.10), reappears in Period 3 (η=−5.277, p<0.01), and disappears again in Period 4 (η=0.440, p>0.10). Coefficient-difference tests confirm significant changes between Periods 1 and 2 (p=0.0021) and between Periods 3 and 4 (p=0.0006). These findings suggest that the price responsiveness of Rwandan tungsten exports varied substantially over time rather than following a uniform trajectory after the regulation.
Johan Fourie
Artificial intelligence is associated with larger research teams, yet in mathematics, among the most codifiable fields, individual researchers working with AI now produce research-grade results. A span-of-control model reconciles these observations. AI lowers execution cost, which expands laboratory scale, and automates codifiable tasks, which lowers the member share of each unit. Team size is therefore quasi-concave in AI capability, with at most one peak. The model predicts that a fully codifiable team peaks when effective automation coverage reaches a closed-form threshold, typically near complete coverage, and, among fields with shared primitives that possess an interior peak, those with less irreducibly human task content peak first. Under explicit priors, the 90 percent forecast intervals for the fully codifiable peak span 2026 to 2030.
Gary Charness, Francesco Feri, Matthew O. Jackson +2
We provide a first causal analysis of the behavioral consequences of the friendship paradox-the fact that people's friends in a network have more connections than average. We find that people's behavior is biased by their network position: they do not best respond to what they should infer the average behavior of the population to be, but instead simply to the average behavior of their friends. Moreover, we find that they fail to learn to overcome such a bias when relocated within the network, varying their observational environment. In these games of complements, the friendship paradox generates a systematic upward distortion in actions, increases behavioral dispersion, and persists despite learning opportunities.
Sasan Mansouri, Daniel Saad, Mark Wahrenburg +2
Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answer can be grounded in the wrong evidence. Similar facts and disclosures recur across sections of a filing, across reporting periods of the same firm, and across comparable firms. FinRank targets this provenance-sensitive retrieval problem by requiring systems to identify evidence for the intended entity, reporting period, and disclosure context. The benchmark contains 1185 manually authored question-answer records over the 10-K and 10-Q filings of 22 companies. Each record includes a reference answer, gold supporting passages, and hand-curated hard negatives drawn from confusable passages within filings, across reporting periods, and across comparable firms. FinRank evaluates passage retrieval, reranking, and hard-negative discrimination as separately measured tasks. Baseline results demonstrate the difficulty of this setting: among the evaluated systems, even a 7B instruction-tuned embedder reaches only 44.8% Recall@10 on the pooled evidence corpus; sub-billion-parameter encoders gain at most 3.5 points over BM25, a finance-adapted embedder trails BM25 by 9.7 points, and pairwise accuracy falls by 13.0-20.5 percentage points when random negatives are replaced with the curated hard negatives. FinRank provides an evidence-first benchmark for developing financial question answering systems that are not only accurate but also grounded in the correct disclosure.
Gabriel de Macedo Santos
This paper documents an applied natural-language-processing framework for measuring the tone of Brazilian Monetary Policy Committee (Copom) statements. The project is explicitly inspired by iSent, Itaú's Central Bank sentiment classifier, particularly its sentence-level division of official communication into hawkish, dovish, neutral, and out-of-context classes. The implementation extends that idea in three directions. First, an LLM identifies short hawkish and dovish expressions and assigns each a 0-to-1 intensity weight. Second, the document index combines sentence counts with document-specific average signal intensities, producing a bounded score from -1 to 1. Third, a separate full-document layer measures forward-guidance direction, guidance explicitness, uncertainty level, and change in uncertainty. The empirical sample is restricted to communications dated August 2016 or later and contains 80 statements and 1,498 classified sentences from August 31, 2016 through August 5, 2026. Across this sample, 33.3% of sentences are hawkish, 18.0% dovish, 42.1% neutral, and 6.5% out of context. The average document score is +0.107, while the most hawkish reading is +0.570 in August 2021. The latest statement, dated August 5, 2026, scores +0.232, with eight hawkish, two dovish, and nine neutral sentences. Its structural overlay is more nuanced: guidance is directionally ambiguous but partly explicit, while uncertainty is classified as central and higher than at the prior meeting. Tone and the guidance-direction score have a contemporaneous Pearson correlation of 0.719. These are descriptive outputs, not a validated forecast of Selic decisions or DI returns. The main contribution is therefore methodological: a transparent, incremental, auditable system that separates rhetorical tone from policy guidance and uncertainty.
Luc Hazenoot, Zhaochun Ren, Amirhossein Zohrehvand
Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities. They score the words a text uses, not the judgment a reader forms about it. Recent work has shown that a gap exists in what Large Language Models (LLMs) know internally versus what they express in their response. This paper asks whether that internal knowledge, read by monitoring the activations of frozen, out-of-the-box LLMs, can stand in for task-specific fine-tuning when measuring concept content, and which extraction method reads it best. We extract such measures via the Recursive Feature Machine (RFM) algorithm and via linear probing, and compare these against an embedding baseline, surface baselines, and the same model's own answer to the question. We demonstrate the approach on financial text, a domain studied extensively and served by established annotated resources, using a human-annotated Environmental, Social and Governance (ESG) dataset. The best linear probe comes within 0.6 percentage points of a fine-tuned domain classifier's accuracy without any task-specific fine-tuning, and outscores the same model's own answer to the question in eleven of twelve comparisons, so the activations carry concept content the response does not report. The simple probe consistently beats the RFM concept vectors, which in turn provide what classification alone does not: a continuous score intended to reflect how strongly a concept is present in a text, whose validation awaits graded labels.
Jerzy Grobelny, Rafał Michalski
This paper extends prior work on linguistic pattern based facility layout optimization by enhancing the LP Alinks framework with an explicit spatial uniformity criterion. While earlier studies demonstrated that linguistic patterns can effectively encode expert knowledge and guide agent based layout emergence, their optimization scope remained limited to cost oriented objectives. To address this gap, we introduce the Normalized Coverage Score (NCS), a scale adjusted measure of spatial evenness that complements the classical flow distance economic objective and enables systematic exploration of cost uniformity trade offs. A full factorial experiment comprising 324 conditions evaluates the combined effects of problem size, link density, virtual force scaling, and two families of membership functions on both objectives. Five way and nested three way ANOVAs reveal that structural factors (number of objects and link density) exert a dominant baseline influence on both economic performance and spatial uniformity. However, while economic cost is overwhelmingly governed by these structural characteristics, uniformity is substantially more sensitive to the linguistic pattern parameters, which control the dispersion compaction dynamics of the emerging layouts within the structural constraints. Based on these interactions, we derive a parameter-selection matrix that prescribes settings for cost minimization, uniformity maximization, or balanced compromise. Comparative analyses with Drezner's method, MDS, and non metric MDS demonstrate that the extended LP Alinks framework consistently attains superior uniformity while maintaining competitive economic performance, making it a robust and interpretable decision support tool for early stage facility layout design.