Marco Zanni, Mohamad Assaad, Touraj Soleymani
Abstract
We study pull-based remote state estimation of an arbitrary, multi-state Markov source while accounting for both freshness and correctness attributes of information. To that end, we formulate a discounted optimization problem in terms of the age of incorrect information (AoII), and express it as a joint source-AoII belief Markov decision process (MDP) under maximum a posteriori (MAP) estimation. We then exploit the information structure of the model and prove that every reachable belief is represented by the last successfully observed source state and the number of time slots elapsed since that observation. For numerical computation, we truncate the elapsed no-success duration at a finite level and derive an explicit error bound and a criterion for selecting the truncation parameter. For reliable links, we show that an optimal policy can be represented by a look-up table of waiting times. For unreliable links, we propose a persistent policy and derive computable performance bounds. We also show that the MAP estimate stabilizes after a finite number of time slots. To further reduce memory requirements, we introduce a hybrid estimator with an early stationary switch and derive a computable bound on the resulting difference in performance. Finally, we extend the framework to multiple sources, formulate the scheduling problem as a restless multi-armed bandit, establish a sufficient condition for indexability, and develop an approximate Whittle index policy based on interpolation. Our numerical results illustrate the structure of the optimal single-source policy, evaluate the performance of the multi-source policies, and verify that the proposed heuristic policies closely approach the optimal solution while substantially reducing computational efforts.