23,760 papers in this slice of arXiv.
Orkun Irsoy, Leman Akoglu, Osman Yagan
Networked systems, from power grids to traffic networks and cloud clusters, carry loads across nodes with limited capacity. A node whose load exceeds its capacity fails and sheds its load onto its neighbors, which can trigger a system-wide cascade. We study how to allocate a fixed capacity budget across nodes to resist these cascades under local load redistribution. The problem is difficult because no optimal allocation is known, and the fail-or-survive objective is non-differentiable and piecewise constant, so exact and gradient-based optimization methods do not directly apply. We introduce TANGCO (Topology-Aware Neural Graph-Guided Capacity Optimization), which uses a graph neural network policy trained through the cascade simulator with policy-gradient learning and a heuristic anchor. We evaluate TANGCO on five synthetic graph families and five real networks spanning power, road, air, and Internet topologies. The learned policy improves on the best of four hand-designed heuristics in all 450 synthetic instances and in 40 of 45 real-network conditions, with robustness gains ranging from 1.6% to 246%. The learned policies transfer to unseen graphs within a family and partially across related topologies, and TANGCOpre
Davide Grossi, Andreas Nitsche, Georgios Papasotiropoulos +3
This paper develops a computationally tractable framework for measuring voters' power in collective decision-making platforms that support transitive and suspendable delegations---commonly known as liquid democracy. We employ Random Walk Decay centrality to capture how influence propagates through delegation networks, argue for its intuitive appeal in this context and its advantages over alternatives such as PageRank, and derive a natural axiomatic characterization within the class of power metrics. Moreover, we conceptualize how the framework can be extended into a practical tool for analyzing power distributions in such platforms, and propose and study, both axiomatically and algorithmically, methods for selecting representative slates of influential participants.
Nuno Crokidakis, Celia Anteneodo
Many traditional models of electoral dynamics emphasize opinion change, social influence and persuasion mechanisms. In contrast, contemporary polarized societies often exhibit relatively stable partisan blocs coexisting with an electorally relevant mobile population. We propose a minimal dynamical model in which electoral competition is governed not by direct persuasion between opposing blocs, but by the dynamics and allocation of a mobile electoral interface. The electorate is partitioned into two stable partisan blocs and a mobile fraction, while direct transitions between the polarized blocs are strongly suppressed. Mobile voters are allocated between the competing blocs through a Fermi-like probabilistic rule governed by rejection asymmetries. Analytical calculations show that the electoral susceptibility to political shocks is proportional to the stationary size of the mobile electoral interface, identifying electoral mobility as the key quantity controlling the macroscopic response of polarized systems to external perturbations. The model naturally predicts a continuum of mobility regimes, ranging from frozen polarization to highly responsive electoral states characterized by a broad mobile interface. These results suggest that, in strongly polarized elections, aggregate electoral changes may be governed primarily by fluctuations at the mobile electoral interface rather than by large-scale ideological conversion.
Jiamu Zheng, Xiaojun Shan
This paper reports two empirical findings on session-based recommendation (SBR), unified in a single model, DTAMLP. First, existing time-aware and GNN-based models (e.g., TiSASRec, SR-GNN) treat every click-time interval as equally informative, even though very short dwell times often reflect accidental clicks carrying little preference signal -- a phenomenon we call sporadic noise. We show that a lightweight, plug-and-play weight fusion module, blending a model's attention weight with a threshold-capped time-interval weight, can be inserted into such models with almost no architectural change and yields a consistent accuracy gain; we view this as the most directly verifiable contribution of this work. Second, we revisit an under-explained observation from FMLP-Rec, where a learnable frequency-domain filter on item embeddings improves accuracy, and offer a possible explanation: time-domain behavior mixes several entangled psychological preferences, and a frequency-domain view may let a model separate and down-weight such preference noise more naturally -- an interpretive conjecture rather than a proven mechanism. Building on both insights, DTAMLP, an all-MLP framework combining weight fusion and FFT-based filtering, is validated on Diginetica and RetailRocket. While this system-level design reflects the state of the field circa 2023 rather than a state-of-the-art claim, ablations confirm the two mechanisms contribute complementary, non-redundant improvements.
Eduardo Graells-Garrido, Daniela Opitz, Francisco Rowe +1
Understanding when migration generates social integration or exclusion is a central challenge for urban communities. Existing research has mostly relied on surveys, administrative data, or aggregate indicators that fail to capture expressions of exclusion at fine spatiotemporal scales. Here, we analyze over 550,000 geolocated reports from SOSAFE (Chile's largest citizen reporting platform) to examine the relationship between migration and hate speech in Santiago. We fine-tune a Spanish hate speech classifier and validate it against human labels. Reports that mention migrants are more likely to contain hate speech than other reports. Hate speech concentrates in areas with recent demographic change (post-2010 arrivals) rather than in established migrant communities. The spatial analysis shows that hate speech hotspots coincide with neighborhoods where recent migrants comprise over a third of the population. Coldspots appear in high-education sectors with minimal recent migration. Reports with hate speech and reports that mention migrants receive more engagement, although their combination is not amplified further. These results show digital bordering on a citizen reporting platform: exclusionary discourse concentrates, and receives more engagement, in neighborhoods undergoing recent demographic change.
Terese Mendiguren-Galdospin, Koldobika Meso-Ayerdi, Jesús Ángel Pérez-Dasilva +3
The dissemination and viralization of information on social media has been widely studied from various perspectives, including that of digital activism. On the other hand, disability-related activism has conquered the online environment, thus obtaining a reach that goes beyond the offline space and generating dialogue in the digital sphere. This article analyses the conversation generated on Twitter, taking as a sample all the tweets with the #disability hashtag before and after the International Day of Persons with Disabilities. More than 18,000 tweets, containing almost as many mentions, were analysed and interpreted as the weighted edges of a graph created using Gephi software and applying the Force Atlas 2 brute force algorithm. The focus was placed on the conversational communities generated around that hashtag, their main themes and the prominent participants in them. In conclusion, although the network of mentions is very dispersed, Twitter is the setting for assertions that receive certain institutional and political support in the Latin American environment, for organisations related to the cause (such as the ONCE Foundation) and, above all, the clear predominance of female activists.
Penny R. Atkins, Manish Parashar
Artificial intelligence is changing the scale and tempo of scientific inquiry. Models can now search, integrate, and reason over data far beyond data repositories familiar to any individual researcher. Yet this expansion creates a prior problem: before a model can produce a trustworthy scientific result, it must locate data that are appropriate for the question, sufficiently reliable for the intended analysis, and accompanied by enough context to support responsible interpretation. As data becomes increasingly abundant, the challenge of finding data has been overcome by the challenge of finding data that you can trust. This paper explores how the social and empirical evidence that accumulates when data are used in research can be used, analogous to social trust networks, to determine fit for purpose and trust. Specifically, the paper explores data-usage graphs as a new layer of scientific data infrastructure. A data-usage graph connects datasets to the publications, people, institutions, topics, software, models, workflows, and other datasets through which they are produced and used. These connections reveal the social life of data: who has relied on a source, for which questions, in what combinations, with which methods, and with what observable impact. They can turn scattered traces of practice into data-usage descriptors that complement conventional metadata and support judgments of trust and fitness for purpose. The central claim is not that popularity establishes trust, but that this can be grown with appropriate contextual history. Usage evidence must therefore be combined with production quality, provenance, governance, semantic clarity, and community validation. The feasibility and value of data usage graphs is demonstrated by implementing the prototype data insights discovery service within the National Data Platform (NDP).
Konstantin Avrachenkov, Lucas S. Sibemberg, Alexander Van Werde
We study spectral clustering in the presence of a confounding latent geometry. The leading eigenvectors may then be dominated by the latent geometry rather than by the communities. Nevertheless, we show in a block latent-space model that communities can be recovered from eigenvectors deeper in the spectrum. We analyze the spectral properties of the adjacency matrix through a limiting integral operator and use its structure to develop DBSPEC, a density-based spectral clustering algorithm that requires only approximate localization of the informative eigenvalue and is robust to poor eigenvalue separation. Crucially, this approach handles general latent geometries, overcoming restrictions to homogeneous toroidal models in prior works. Our theoretical predictions for the location of the informative eigenvalue notably align with observations in real-world experiments.
Katherine Hamilton, Irina Epure, Frank Takes
Register-based social networks have become of increasing interest in countries where formal government-curated microdata is available. Due to the non-trivial generative process of register-based networks, existing random graph models fail to facilitate effective structural analysis, hindering the discovery of meaningful insights in the underlying social system. In this paper we introduce the Multiplex Affiliation-based Random Spatially-embedded (MARS) graph framework, which replicates the construction method of register-based social networks. We derive fundamental statistical properties of MARS ensembles in general and special cases. To demonstrate the applicability of the framework, we implement a simple model under the MARS framework and show that it recovers similar properties to those exhibited by the population-scale register-based social network of the Netherlands. Furthermore, we analyse the effect of spatial tie strength on closure in the network and compare our results with existing empirical findings, showing that increased spatial freedom is correlated with decreased social cohesion.
Ainara Larrondo-Ureta, Simón Peña-Fernández, Jordi Morales-i-Gras
The aim of this paper is to (1) identify textual and visual themes and sub-themes associated with the #wellbeing hashtag on Instagram, (2) assess their varying levels of engagement and (3) investigate gender bias present in the analysed visual narratives. This study employs a range of big data analysis techniques to investigate various dimensions of wellbeing on Instagram. Initially, a sample of 9,844 posts was processed using data mining and text analysis methods to identify and categorise significant hashtags, which facilitated the classification of posts into thematic clusters through hierarchical clustering. Engagement was then assessed using non-parametric statistical tests. Additionally, computer vision models were used to analyse and classify visual narratives, grouping images into communities based on visual similarities. Finally, gender representation in the images was examined using object detection models. The study reveals that the discourse around #wellbeing on Instagram is predominantly feminised and focuses primarily on mental, psychological and spiritual aspects. Therapeutic and positive psychology narratives are the most prevalent and engaging, while physical activity and nutrition play a relatively secondary role. Two major macro-narratives emerge: firstly, mental health and emotional wellbeing, often featuring hashtags related to therapy, self-discovery, spirituality and motivational quotes; secondly, though less prominently, themes and visual narratives concerning physical activity and healthy habits, emphasising exercise and nutrition.
Robert Šamárek, Radek Martinek
The academic publishing ecosystem is a vast, heterogeneous network of works, authors, institutions, journals, and topics. Traditional scientometrics reduces it to isolated tabular indicators (h-index, Impact Factor) that ignore topological context and are not designed to capture coordinated illegitimate practices. Building on our companion review, which proposed graph analysis of publishing integrity, this paper implements that approach. We define a heterogeneous multivariate graph model over OpenAlex open data (seven node types, seven edge types) and a methodology based on projections (citation and co-authorship networks), interpretable structural metrics, community detection, and three screening detectors of anomalous publishing patterns. We deliberately avoid binary classification: detectors return ranked candidates with explicit structural evidence for human assessment. On the institutional corpus of VSB - Technical University of Ostrava (2020-2025) with its one-hop citation neighbourhood, community detection recovers real research groups, centralities identify cross-disciplinary bridges, and the screenings flag dense co-authorship cliques, locally closed citation loops, and thematically isolated venues. On a second, venue-centric corpus with external ground truth (journals delisted by Scopus and DOAJ) and size-matched controls, a naive case-control design yields seemingly strong but spurious detectors (a prominence confound), whereas after matching the only robust signal is the breadth of disciplinary scope (AUC 0.70); an open graph-based prestige measure (PageRank over the journal citation network) tracks a JIF proxy while being an order of magnitude more resistant to citation gaming than count-based indicators. We release the method as the open-source library apnet with a reproducible CLI workflow and a web interface; the analysis runs on commodity hardware in minutes.
Robert Šamárek, Radek Martinek
Predatory journals pose a significant challenge to the integrity of the Open Access (OA) publishing model by exploiting its framework for financial gain while bypassing essential editorial and peer-review standards. This study critically evaluates existing methodologies for identifying such journals, ranging from manual blacklist checks to advanced automated approaches utilizing machine learning. The analysis highlights critical limitations, including the lack of a universally accepted definition of predatory journals, over-reliance on binary classification systems (e.g., blacklists and whitelists), and issues with scalability, reliability and interpretability. To address these shortcomings, this paper introduces a novel methodology based on multivariate graph analysis. By modeling the academic publishing ecosystem as a network of interconnected entities (such as authors, articles, journals, and publishers), this approach provides broader insights into the dynamics of scholarly communication and could help identify illegitimate publishing practices by utilizing graph algorithms like centrality measures, community detection, and anomaly detection. The proposed framework aims to enhance the accuracy, scalability, and transparency of detection of illegitimate journals and publishing practices while fostering a more comprehensive understanding of the academic publishing landscape.
Lena Holzwarth, Rita González-Márquez, Dmitry Kobak
Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but even inside the Methods section, the overall prevalence of LLM usage is over 50%. We believe that our estimates are crucial to shape future guidelines and policies.
Guanqun Yang, Tong Qi, Xiaoxue Han
Real-world recommendation platforms routinely collect explicit negative feedback such as 1-star reviews, hate-button clicks, distrust between users, and very-low watch-ratio videos. Learned sign-aware recommenders exploit this signal for clear accuracy gains, but only at the cost of gradient-based training. In parallel, a line of training-free spectral collaborative filtering methods matches or beats learned graph recommenders at a fraction of the cost, yet operates on positive interactions alone. We bridge these two lines with DualSpectralCF, a training-free framework of two components that attach to any spectral backbone of the form r^u=F(M)ru: a signed input signal ru± that encodes the user's explicit dislikes, and a signed item-item operator M± that blends like-together and dislike-together similarity. The framework is backbone-agnostic and adds just two scalar hyperparameters. We instantiate DualSpectralCF on ChebyCF, GF-CF, and Turbo-CF, and evaluate on five sign-aware benchmarks: every instance matches or beats its unsigned backbone on all 5 datasets, with Recall@20 lifts up to +32.6% with backbone-specific (γ,κ) tuning and +1.9% to +16.0% for DualSpectralCF-Cheby at the fixed default (γ=−0.5,κ=0.1), and the family runs 7.7 to 155.3× faster than SIGformer while reaching 70.7% to 90.7% of its accuracy. Sign-awareness helps most for cold-start users, with up to +29.2% Recall@20 on Epinions users with 1 to 5 training items.
Michael Kreil, Tristan Manfred Stöber, Daniel Thilo Schroeder
The criminal case of Jeffrey Epstein has generated a complex, long-running global discourse on digital platforms, characterized by punctuated attention shocks, conspiracy theories, and blame attribution. To facilitate the computational study of these dynamics, we introduce the FUBU-EPSTEIN dataset, a large-scale, multi-dimensional research corpus of 54.38 million Twitter statuses authored by 7.11 million users and collected between August 2019 and April 2023. The source corpus was captured continuously in near-real time and enriched with a directed social contact graph of 37.03 million edge rows and annotations from Qwen2.5-7B-Instruct covering sentiment, conspiracy and misinformation stance, toxicity, moral emotion, and related dimensions. For public distribution, we created a textless, de-identified derivative that retains one row for every deduplicated status, categorical and numeric annotations, coarse temporal information, and 46.15 million internal status relationships. It excludes tweet text, original post and user identifiers, handles, profile fields, exact timestamps, and reverse mappings. The resulting release supports longitudinal content and diffusion analyses while reducing disclosure and platform-content redistribution risks. To request access to raw data for collaborative research under ethical and legal safeguards, contact fubu.dataset@gmail.com.
Haoyu Han, Yuming Liu, Lei Huang +3
Modern recommendation systems on social media platforms such as Meta must model complex social relationships, including friendships, group memberships, and creator interactions, alongside massive and heterogeneous content such as text and video. Traditional recommendation models, however, often omit these signals or treat them independently, lacking the reasoning capability to integrate multi-relational context for fine-grained personalization. We present ConnectionMind, a production-ready recommendation framework that tightly integrates the social network structure with large language models (LLMs) to enable scalable, interpretable, and reasoning-aware personalization in Meta. ConnectionMind constructs a heterogeneous graph connecting users, items, friends, groups, and creator pages, and formulates recommendation as a graph reasoning problem: discovering personalized paths from users to candidate items. An LLM-based policy is employed to reason over these graph structures and guide recommendation decisions. To train the system at scale, ConnectionMind adopts a two-stage learning strategy. We first perform supervised fine-tuning (SFT) on large-scale user-item interaction trajectories to initialize the reasoning policy, followed by end-to-end reinforcement learning (RL) to refine the model's ability to reason over social graphs for personalized recommendation. Extensive experiments on multiple real-world datasets demonstrate the effectiveness of ConnectionMind compared to representative baselines. More importantly, ConnectionMind has been deployed in Meta's large-scale recommendation pipeline and has been evaluated through online A/B tests, achieving a 0.43% improvement in video watch time. These results demonstrate measurable real-world impact in a production recommendation system.
Yu Tian, Eleanor Wiesler, Melanie Weber
Understanding the geometry of complex networks is critical for effective modeling and analysis across domains. While discrete notions of Ricci curvature have emerged as powerful tools for characterizing both local and global network structure, existing formulations are largely confined to undirected networks with real-valued weights. This limits the use of curvature-based analysis of directional and complex-weighted relations that arise naturally in many applications, from social and biological systems to quantum and signal-processing networks. In this work, we introduce a principled extension of Ollivier's Ricci curvature to complex-weighted graphs, which encompasses directed graphs as a special case. We establish fundamental theoretical properties of this new notion, including relations to the magnetic Laplacian and combinatorial upper and lower bounds that relate curvature to cycle structure in local neighborhoods. We further develop computational methods for curvature estimation and demonstrate their utility in community detection on directed networks.
Chenchen Mao, Hanjing Shi, Haiyan Jia +3
Independent researchers often lack access to intervention capabilities for controlled experiments on live social media platforms. We present Outer Limits, a browser-based system for controlled content experiments within the existing Old Reddit interface, rather than in a reconstructed simulation. The system renders content locally, records study events, and contains configured voting and commenting actions so that neither constructed content nor experimental write interactions reach Reddit. In a 219-participant perceptual-fidelity study, ART ANOVAs found no significant Post Type, Participant Awareness, or interaction effects. Exploratory TOSTs met the d = plus-minus 0.50 equivalence criterion for the marginal contrasts and for Post Type within the forewarned subgroup. We also illustrate the system with a factorial study varying post frame, comment frame, and comment stance. Outer Limits combines three properties that the approaches considered here provide separately: precise control over experimental content, an existing platform interface, and containment of experimental content and interactions from the host community.
Gayoung Jeon, Cameron Moy, Silvia Teliz +3
TikTok's global growth has made it a prime platform for both entertainment and political discourse, prompting increased social science research. However, this rapidly evolving research field faces a fundamental reproducibility crisis. TikTok's opaque algorithmic systems hinder researchers from drawing meaningful empirical inferences, while the lack of standardized data collection methods compounds these challenges. This study addresses these methodological gaps by systematically comparing three data collection tools - the official TikTok Research API, Pyktok, and Apify. We evaluated five endpoints: User, Hashtag, Keyword, Comment, and Related Video. Results show substantial cross-tool differences, especially for hashtag and keyword searches. The Research API uses back-end API calls, whereas Apify and Pyktok rely on front-end web scraping, producing systematic differences in the time periods and popularity levels represented in retrieved content. The three tools yielded comprehensive and consistent results only for the user endpoint. Our results question whether these tools can acquire truly random[-ized] samples, as they introduce methodological confounds that may compromise research validity in ways not yet fully understood. Based on these results, we offer methodological, transparency, and ethical recommendations and guidelines to increase TikTok research quality.
Valentijn Oldenburg, Floris de Kam, Stef de Wildt +1
In fair ranked link prediction, demographic parity (ΔDP) is a common fairness metric. Yet, Mattos et al. (2025) argue that it fails to detect exposure bias because it ignores where links appear in the ranking. In this study, we reproduce this claim by showing that ΔDP can indicate aggregate parity even when some subgroup-pair links are systematically ranked lower than others. The proposed rank-aware Normalized Discounted KL-divergence (NDKL), however, does detect such disparities. We also reproduce the effectiveness of MORAL, a post-processing method that improves exposure-based fairness while maintaining competitive utility. Beyond reproduction, we assess robustness using synthetic homophily settings, categorical sensitive attributes, and additional fairness and utility metrics, including subgroup-pair-adapted Attention-Weighted Rank Fairness (AWRF). Overall, our results show that exposure-based metrics uncover biases hidden by ΔDP and that MORAL reduces these biases with minimal utility loss across diverse settings and datasets. We release a corrected, reproducible implementation at https://github.com/Floris93100/reproducing-MORAL.