Nearby in the stack

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World · arXivDesk