Paolo Maranzano, Raffaele Mattera, Shonosuke Sugasawa
Abstract
Area-level small area estimation (SAE) models, such as the Fay--Herriot (FH) model, borrow strength across domains through covariates and random effects, but they can struggle when the relationship between the covariates and the outcome is spatially heterogeneous, that is, when it changes across the spatial domain of interest. We propose a spatially-clustered FH (SC-FH) framework that simultaneously (i) estimates cluster-specific regression coefficients and random effects variances and (ii) generates spatially coherent partitions of the geographical domain. Estimation maximizes a penalized likelihood that augments the FH likelihood with a Potts-type spatial cohesion term over the areal adjacency graph, through an efficient strategy that alternates between sequential label updates and closed-form FH updates within clusters. In simulation experiments run on the real geography of the application, the method recovers the latent regimes almost exactly whenever they are separated in the covariate--response space and improves prediction accuracy over the standard FH benchmark, with the spatial penalty acting as a stabilizer of both classification and estimation. An empirical application to the average standard output of farms in the Po Valley (Northern Italy) identifies two spatially compact production regimes with significantly different cluster-wise coefficients, and shows that the clusterwise predictor improves on the direct estimates while avoiding the over-shrinkage of the pooled model.